Pipeline for Translating Books with Claude Code and Codex
This skill is for translating books from English into Bulgarian in Claude Code or Codex, with the same stages and the same instructions as the translation tool wenyi. The difference is who does the work. wenyi calls a model through an API; here the agent in Claude Code or Codex runs each stage by hand. It can also review an existing translation and leave every paragraph without a provable error untouched. The skill is free under the MIT license, and you can download it from GitHub: github.com/valcheffnet/wenyi-pipeline-skill.
I wrote the first version during my comparison of AI literary translators. One of the contestants couldn’t go through wenyi, and I wanted its result to be comparable with the rest. Later, when I checked the skill against wenyi’s code, I found it had been skipping four stages. This post describes the new version, with every stage. It also has one rule I added afterwards: before it starts, the skill always asks which mode to translate the book in.

When this skill makes sense
wenyi through an API is faster and cheaper. In the comparison it was also the less dependable route. Through a reseller like OpenRouter, the same model behaves differently depending on which provider serves the call. deepseek-v4-pro through one provider was cut off at every output limit, and translated the whole book with no cut-offs when called directly at DeepSeek. gemini-3.8-flash through Google AI Studio returned an empty response nineteen times and stalled at 20% of the book until we moved it to Vertex. wenyi itself crashed on those empty responses until we patched it. In Claude Code or Codex it costs more, but there is one model behind every call and nobody switches the provider under you.
So for translating books, the skill makes sense in three cases. The first is when you need a translation that reaches the end without surprises and you’re willing to pay for that. The second is a model you can’t reach through an API but use in Claude Code or Codex. The third is a comparison: when the instructions and the batch boundaries are the same, the difference in the result comes from the model.
The review-only mode is useful on its own. Give it a finished translation, by a person or by a machine. You get a list of errors against the original, each with its evidence. You also get a corrected text in which only the corrected paragraphs differ.
What a skill for Claude Code and Codex is
A skill is a folder of instructions that the agent in Claude Code or Codex loads when a task matches them. This one contains:
wenyi-pipeline/
SKILL.md the stages, the modes and the rules
references/ wenyi's 20 prompts, verbatim
scripts/make_batches.py cuts an EPUB into chapters and batches
agents/openai.yaml name and description for the Codex interface
The prompts in references/ are not paraphrased. They were rendered by wenyi’s own rendering function with a Bulgarian language profile. wenyi gives every model the same instructions, so the agent has to get exactly those. Paraphrase them and you have a different pipeline with the same name, and the comparison quietly breaks.
Where to download it
The skill lives in the valcheffnet/wenyi-pipeline-skill repository on GitHub. Without Git, open the page, click the green Code button and choose Download ZIP. With Git it’s one command:
git clone https://github.com/valcheffnet/wenyi-pipeline-skill.git
Then copy the skill folder to where your agent looks for skills. Claude Code and Codex use the same open skill format, so the folder is the same and only the location differs. In your home folder it’s available in every project.
For Claude Code:
cp -r wenyi-pipeline-skill/skills/wenyi-pipeline ~/.claude/skills/
For Codex:
mkdir -p ~/.agents/skills
cp -r wenyi-pipeline-skill/skills/wenyi-pipeline ~/.agents/skills/
In both cases, also install the tiktoken library:
pip install tiktoken
On Windows, in PowerShell, the copy commands are:
# Claude Code
Copy-Item -Recurse wenyi-pipeline-skill\skills\wenyi-pipeline "$env:USERPROFILE\.claude\skills\"
# Codex
New-Item -ItemType Directory -Force "$env:USERPROFILE\.agents\skills" | Out-Null
Copy-Item -Recurse wenyi-pipeline-skill\skills\wenyi-pipeline "$env:USERPROFILE\.agents\skills\"
To use the skill in one project only, copy the folder to that project’s .claude/skills/ for Claude Code or .agents/skills/ for Codex. Claude Code loads it in the next session, and Codex picks it up on its own. In Codex you can also call it directly with $wenyi-pipeline. The repository also has a README with every stage and the measured token counts.
Before you start
- The skill installed as described above.
- Python with the
tiktokenlibrary. wenyi counts tokens with it, and without it the batch boundaries drift away from wenyi’s. - The book as an English EPUB. The prompts in the skill are for translating books from English to Bulgarian.
- An idea of how many tokens you can spend. More on that in step 2.
1. Cutting the book into batches
The first step is local and spends no model tokens at all. The script reads the EPUB in reading order and turns it into what wenyi would make of it. Every file with text is a chapter, headings go to a separate title translation, and every paragraph is a segment. A paragraph over 1,200 tokens is split at sentence ends. A chapter’s segments are packed into batches of up to 1,800 tokens.
python scripts/make_batches.py shadows.epub work/
counter: tiktoken cl100k_base
chapters 1 | paragraphs 296 | segments 296 (paragraphs split: 0) | titles 6 | batches 10
source tokens 16,060 | words 12,079
This is the story from the comparison, and the ten batches match the ten calls wenyi made for the same text. The script also finds headings in EPUBs converted from FB2, where a heading is an ordinary paragraph inside a separate block.
Before you go on, look at the first and last chapters in source.json. Converted books often carry a library card, a printed table of contents or ads as ordinary text. wenyi would translate them, so the skill points them out and asks whether to keep them.
2. The mandatory mode question
Here the agent stops and asks. It doesn’t guess the mode from the wording of the task and doesn’t pick one itself, even when the task mentions one. The reason is cost. The lightest and the heaviest mode differ by a factor of four. Which one is enough is for the person who will read the book to decide.
In Claude Code the question comes with ready answers to pick from. If the host has no such tool, the skill asks in plain text and waits for your answer.
| Mode | What it does | Tokens for a 12,000-word story |
|---|---|---|
| A, plain | style analysis, translation and glossary per batch, titles | 177,000 to 209,000 |
| B, translate + review | mode A plus the review loop | 415,000 to 490,000 |
| C, full | analysis, chapter and book synopses, translation, polish and glossary, titles, review | 640,000 to 830,000 |
| R, review only | review of a finished translation | about 290,000 (estimate) |
The figures for A, B and C were measured when wenyi translated the story through an API with two different models. R is an estimate from them. An agent in Claude Code or Codex spends more, because it carries the whole conversation along. That’s why the script from step 1 prints an estimate for each mode, and the estimate goes into the question.
3. Getting to know the book
The style analysis runs in every mode. The agent reads three samples, from the beginning, the middle and the end, and returns the genre, tone, narration, register and a list of characters and terms. One of the most important rules for Bulgarian sits here: a character’s gender is recorded only when the text proves it. A name, a form of address or a tone of voice is not proof.
In the full mode, a digest of every chapter follows, up to 150 words. Then comes a synopsis of the whole book, up to 350 words and including the ending. The translator knows what happens later and doesn’t render a scene in a way the plot will contradict.
4. Translating in batches
Every batch gets the style from the analysis, the synopses, the glossary for the current chapter, the last six translated paragraphs and the next paragraph of the original. That last one is only a reference. If the batch ends in the middle of a sentence, the translation has to stay unfinished too, instead of inventing an ending.
The strictest rule is the count. Every paragraph of the original gets exactly one translated paragraph. wenyi checks the count and rejects a batch where it doesn’t match. Nobody checks the agent, so the skill makes it count for itself after every batch.
5. Polish in the same conversation
This is where the old version was wrong. I had polish as a separate phase after the whole translation, while wenyi polishes right after each batch, as a continuation of the same conversation. The model still remembers what it translated and why, and smooths the word order without losing the meaning. The number and order of paragraphs stay the same.
6. A glossary after every batch
After every batch the agent pulls names, forms of address and fixed phrases from the original and the translation. These must stay the same until the end of the book. The following batches get the updated glossary. If a new rendering of a name disagrees with the old one, the first one stays and the disagreement is recorded. Nothing is overwritten quietly.
In the published Bulgarian translation of “Children of Ruin”, the names of characters and ships were spelled differently in different chapters, and I unified them by hand. The glossary exists to prevent exactly that.
7. Titles
Chapter titles and table of contents entries are translated separately, with the full glossary, so the names in them match the text.
8. The review loop
The review has three roles. The first reads the original and the translation in blocks of up to 5,400 tokens. It flags only certain errors, of five kinds: something missing, something added, wrong meaning, a glossary entry not followed, and a wrong pronoun. Style is not an error.
The second role checks every finding before accepting it. It can ask for evidence from the rest of the book: a glossary entry, other places with the same name, neighbouring paragraphs or the synopsis. It can ask for up to four things per round, for at most two rounds. The agent can see the whole book on disk and could easily read all of it, so the skill explicitly forbids reading outside those requests. Otherwise it would be a different review.
The third role fixes only the confirmed errors, with minimal edits, in a separate copy. The copy is reviewed again as if it were a new translation. The loop stops after two clean rounds in a row or after two rounds of fixes. At the end, every fix is checked once more, and only then does it go into the translation.
9. Self-check before finishing
- Every chapter has exactly as many translated paragraphs as the original, and none is empty.
- Latin letters appear only where the original has them.
- The text has no translator’s note and never mentions a model, a company or a tool.
- Glossary disagreements are listed.
- In mode R, paragraphs without a fix are identical to the last character.
Everything each stage produces is saved to disk. If the agent stops halfway, you lose one batch at most.
Translating books: what a whole novel costs
I ran the script on Adrian Tchaikovsky’s “Children of Ruin”, 148,000 words. It came to 93 chapters and 142 batches. The estimate is about 2.2 million tokens for the plain mode and 8.4 million for the full one, and that’s for wenyi through an API. An agent spends more.
On a subscription instead of an API key that isn’t practical. A story or a single chapter is fine for the skill in Claude Code or Codex, but a whole novel by hand is not realistic. For translating books the length of a novel, run wenyi with an API key, pin the provider explicitly and turn off fallbacks. Leave the review of selected chapters to the skill.
What I learned from the first version
I wrote the first version in a hurry, for one contestant in the comparison. It had no style analysis in the plain mode, and no chapter digests, glossary, titles or evidence-based checking of findings. At the time they looked like details you could skip for a short story.
It shows in the results. In the translation with review, the review found nothing. So the reviewed version came out byte for byte identical to the plain one, and two identical texts went into the blind comparison. I can’t prove that a fuller review would have found errors. I do know that this contestant went through a lighter pipeline than the models it was compared with. That’s one more reason not to count its result as a win.
Hence my rule. If a skill claims to reproduce a tool, check it against the tool’s code. That’s how I built the new version: its stages and thresholds come from wenyi’s code and settings.


