Ranking of eight AI translators by points and price, with the names hidden

In Search of the Best A.I. Literary Translator – Detailed

My wife reads faster than books get translated into Bulgarian. Some of the ones she wants don’t exist in Bulgarian and probably never will, because the market is small and the translation doesn’t pay for itself.

So I tried machine translation. It came out unreadable: sentences that are grammatically correct and mean nothing, and dialogue punctuated the English way. My wife got to page thirty and gave up.

I started with DeepL, which everyone calls the best. It turned out to be all talk, and later in the test it came last of ten.

The same happened with qwen3.8-max. I only included it because it is supposedly the best model for translation. It came second to last and was the only one that dropped parts of the text. All talk again.

So I decided to run my own study and find out which AI literary translator really does the job best. In Bulgarian specifically, because a model that translates beautifully into Spanish isn’t necessarily good at Bulgarian, and the English-language rankings say nothing about it. Bulgarian has things English doesn’t. Gender follows the word the translator chose, so the same ship is correctly “she” if you called it a galley and “he” if you later called it a ship. The definite article sticks to the end of the word, there is a separate mood for retelling what you didn’t witness, and dialogue opens with a dash. At every one of those points the models differed.

What I wanted to find out

Which model, with what processing around it, can translate a work of fiction into Bulgarian well enough that someone finishes the book, and what that costs. I also wanted to know whether it is possible to measure at all who translates better.

How I tested the AI translators

I took Robert E. Howard’s story “Shadows in the Moonlight” from 1934. Howard died in 1936, so in the EU the text has been free since 2007. The story is 296 paragraphs and 12,079 words. As a yardstick I used an old published human translation: 305 paragraphs and 12,276 words, with every paragraph of the original matched to one in it.

The translators were eight language models: gpt-6-astra, claude-opus-5, gemini-3.1-pro, gemini-3.8-flash, deepseek-v4-pro, grok-4.6, kimi-k3 and qwen3.8-max. Next to them I put DeepL, which is not a language model and serves as the floor. The models ran through wenyi, a tool for translating books, to which I added a Bulgarian profile. Every model got byte for byte the same task, and between any two setups only four lines differ: a comment, the model, the provider and the file path. In the first round they all translated with no extra steps, in chunks of 1,800 tokens.

I scored them three ways. First, mechanical checks over the whole text. Then a full read of all 296 paragraphs of every translation against the original, by the same reader, Fable 5.1, with serious, moderate and minor errors defined in advance. And finally a blind comparison: six passages from the ten translations, unlabelled and shuffled, in front of four judges from four different makers, Fable 5.1, Astra Pro, Gemini Pro and DeepSeek. Each of them has a relative among the translators, which is why there were four.

There were a few rules. The judges don’t know who translated what. They score all the translations at once, so they measure on the same scale. Their instructions also explain how gender works in Bulgarian, because without that almost every translation looked wrong. And I checked whether the models had seen the published translation during training.

The whole bill for model calls is about 46 dollars: 42.53 through OpenRouter and about 3.50 on DeepSeek’s own account.

What came out

Whether a model works depends on who answers the call

Most models are bought through a middleman like OpenRouter, and behind one model name a different company can answer each call. For DeepSeek the catalogue lists 16 such companies, for Kimi 20. The price changes with whichever answers: one and the same probe cost 0.0227 dollars by the catalogue, and 0.01151 was charged. Behaviour changes more.

The same model, two serving paths

Through one of the providers, deepseek-v4-pro cut the translation off at every ceiling I tried: 6,000, 16,000 and 48,000 tokens. Directly through DeepSeek’s own API it translated the whole story without a single break. gemini-3.8-flash through Google AI Studio returned 19 empty responses and stopped at 20% of the book, while through Google Vertex it went all the way. My first thought was censorship, because blood gets spilled in the story. I repeated the same call nine times and the empty response never came back. It was a temporary fault at the provider.

The empty response also exposed a bug in wenyi itself, which crashed on it. The fix is three lines and all 790 tests pass. One of the failures was mine: out of habit I set a ceiling of 32,000 tokens on DeepSeek, and the translation stopped exactly there. Since then I pin the provider explicitly and check it with a live call before I start the big book.

Counting can’t tell a good translation from a bad one

The mechanical checks count words, dialogue lines, dashes, quotation marks, very short paragraphs, and missing or repeated paragraphs. One solid thing came out of it: none of the nine dropped a whole paragraph. After that, counting stops helping.

Identical on the numbers, four tiers apart on reading

Six of the models come out practically identical: length between 0.948 and 0.975 of the original and 150 or 151 dialogue lines with a dash. kimi-k3 has excellent numbers and a text full of words that don’t exist in Bulgarian, and in the blind comparison it is among the last. Reading split those same six models into four clearly different levels. Counting only catches the obvious failures.

The full read produced these findings:

Translation Findings, count Severe, count Moderate, count Minor, count
gemini-3.1-pro 15 0 4 10
gpt-6-astra 22 0 6 15
grok-4.6 92 0 24 68
gemini-3.8-flash 57 1 14 42
deepseek-v4-pro 77 1 30 45
human translation 125 2 13 105
claude-opus-5 89 4 30 55
kimi-k3 70 8 53 9
qwen3.8-max 108 15 25 68
DeepL 146 16 79 46

The totals aren’t quite comparable, because each reader decides for itself what to note. The serious column is comparable, because serious was defined the same way for everyone. DeepL is last with 16, and almost all of them are words with two meanings. “Quarter!”, which in a fight is a plea for mercy, becomes “Четвърт!”, the fraction. “Painter”, the rope on a boat, becomes “Художникът”, the artist. In one paragraph the word is wrong, in the next it is right, because every sentence is translated on its own. qwen3.8-max drops text and the reader feels the gap. DeepL writes smooth Bulgarian that says something else, and an editor without the original in front of them won’t catch it. According to the reader, bringing the DeepL translation up to print costs 60 to 70% of a new translation.

Which AI literary translator came out best in the first round

Round one: ten candidates, three judges

One judge gave only an order without points, so the chart shows three. The first three places are the same for every judge: gpt-6-astra, gemini-3.8-flash and deepseek-v4-pro. Further down, opinions differ by up to three places.

I also checked how much each judge goes easy on its relative. Fable 5.1 doesn’t move claude-opus-5 at all, Astra Pro lifts gpt-6-astra by 0.33 places, DeepSeek lifts its own model by one place, and Gemini Pro lifts gemini-3.8-flash by 1.33 places and gives it 10 out of 10 on all six criteria, the only perfect score in the first round. The bias is real but small: if I remove each judge from the scoring of its own family, the top doesn’t move.

What it costs

What one 12,000-word story costs to translate

Expensive and slow doesn’t buy proportional quality. gemini-3.8-flash is the cheapest and the fastest, 33 cents and seven and a half minutes for the story, and it is second in the ranking. gpt-6-astra costs 4.20 dollars, 12.7 times more, for a difference the judges describe as one level. grok-4.6 works 10.5 times longer than flash and comes in sixth or seventh.

There is also a hidden bill. Some models think out loud before they answer, and that thinking is billed as output text, even though nobody reads it.

Output tokens billed per translation

The Bulgarian translation itself is about 65,000 tokens. grok-4.6 returned 284,459, deepseek-v4-pro 300,398, and claude-opus-5 managed with 41,254. The setting for how much a model should think doesn’t mean the same thing at different makers: at the lowest level three models don’t think at all, while two spend about 6,000 tokens each and return a broken answer. The only reliable way to budget is to translate one chapter and multiply.

What wenyi’s pipeline adds

wenyi can do more than a plain translation. Before translating, the model reads the whole book and makes a summary and a glossary of names. During translation it polishes the text. After translation it reviews its own work and fixes what it finds. I tried these steps on the four best models from the first round. The other four stopped here, because the same bill on a model four levels down answers nothing new.

The review on its own

First I ran only the review, on the finished translation. That way the translation stays word for word the same, and everything that changes afterwards is the review’s doing.

What the review pass actually moves

Nobody rewrote the text. deepseek-v4-pro changed the most, 11 paragraphs of 296, and found three real errors in its own translation, among them a line where the man had been put in the feminine. gpt-6-astra changed two paragraphs and got both right, and its review cost 3.07 dollars, three quarters of the price of the whole translation. gemini-3.1-pro changed eight and made one worse: of “В името на Ищар!”, “In Ishtar’s name!”, it left only “Ищар!”. That is the only one of the 22 changes that makes the text worse.

One and the same error, “педя” (a span) instead of “длан” (a palm), appears in six of the eight translations. gpt-6-astra’s review found it. The reviews of gemini-3.1-pro and deepseek-v4-pro went over the same error and said nothing. Same pipeline, different result, so the difference is in the reviewer. When the three versions are scored, the review mainly lifts the errors criterion, by one point in three of the four models, while language and completeness don’t move.

In the whole material there was one single serious error, the kind that turns the meaning upside down. At gemini-3.8-flash, “too drunk to vote either way” had become “напи се до козирката, за да може изобщо да гласува”, drunk so that he could vote at all. flash’s review changed exactly one paragraph of 296, and it was that one.

The full pipeline

Then I ran the four models again, this time through the full pipeline. In the chart each of them has three versions: the plain translation, the same translation with a review added, and a new translation through the full pipeline. Each set of three was scored separately, by one blind judge who sees it all at once, so compare the points only within one model.

What the pipeline does to the score

gpt-6-astra goes up from 6 to 8 points, gemini-3.1-pro from 7 to 8. gemini-3.8-flash drops from 8 to 7, deepseek-v4-pro from 7 to 6. For both drops the judges say the same thing: the translation becomes more precise in detail and in gender, but pays for it with literalisms and words that don’t exist. So there is no general answer to whether you should switch the pipeline on. It depends on the model, and the only way to find out is to try both. It costs 1.9 times the plain translation for gemini-3.1-pro and 3.4 times for flash.

Same texts, counted by severity

If you just add up the findings, the pipeline looks harmful: 90, then 93, then 106. The breakdown shows the opposite. Moderate errors fall from 29 to 19, while minor remarks rise from 60 to 87, and that holds for each of the four models on its own. The pipeline trades errors of meaning for style remarks. Four out of four in the same direction happens by chance once in sixteen times, so it is a signal, but not yet proof.

The human translation came in below the best machines

The human translation, criterion by criterion

Across the different scorings the published translation sits between sixth and eighth place out of ten. Its biggest loss is completeness, 4 out of 10, the lowest score after qwen3.8-max, which really does drop text. The translator condenses on purpose. Of 125 departures from the original, 66 are deliberate choices, 35 are errors of meaning and 23 are language or typography. The criterion, though, counts every cut as a loss, even when it works. The human translation has both errors and style, and this scale only sees the errors.

Completeness isn’t the only reason. For Bulgarian the translation gets 6 out of 10, level with claude-opus-5 and below four of the machines: gpt-6-astra gets 9, gemini-3.8-flash gets 8. It keeps using the short definite article where Bulgarian grammar needs the full one. There is wrong verb government too, one case of it twice. About ten technical errors remain in the text, such as a line-break hyphen in the middle of a word, doubled words and Latin letters inside Cyrillic words. On gender it scores 8 out of 10, and the review finds one such slip: once Olivia turns masculine. These most likely come from scanning the text, not from the translation. The articles and the verb government, though, are in the translation itself.

The models hadn’t seen the translation before

The published translation may have ended up in the data the models were trained on. Then they would simply be remembering it. I checked two ways. First, the longest common run of words: between a model and the human translation it averages 5.2 to 5.7 words, and between the models themselves 8.8. Then I gave the models a passage of the human translation and asked them to continue: the continuation resembles the human as much as it resembles their own translation. Nothing stands out. The stronger check, the same test on a book with no published Bulgarian translation, I skipped on purpose.

Longest wording a candidate shares with another text

The second round: the judge sits down to translate

After the first round, Fable 5.1 did all the reading and scoring. For two days it checked other people’s translations without handing in one of its own. So I had it translate the story too, with wenyi’s same instructions, run by hand. That is possible because wenyi’s instructions don’t depend on the model. I split the book into the same chunks of 1,800 tokens and got 10, exactly as many requests as gpt-6-astra made through the tool. It made the three versions the others got: a plain translation, one with a review, and one through the full pipeline.

Fable couldn’t judge its own translation, so two outside judges scored it, gemini-3.8-flash and deepseek-v4-pro. And all the translations came back into play. The ten from the first round plus Fable’s three, thirteen texts at once in front of each judge, on the same six passages. It doesn’t work any other way: an eight given without knowing there is a nine in the same room isn’t the same eight.

Round two: thirteen candidates, two outside judges

gpt-6-astra stays first. Fable is right behind it. But I don’t count it as a win, for four reasons, all in its favour. It knew what it would be judged on, because its rules were written after the first round and answer exactly the errors the others were penalised for. After every chunk, it checked its own work. And it had the whole book in front of it, while the others see a little at a time. And once it caught and fixed its own error by itself, a Latin “p” inside a Bulgarian word. This round shows what a well-tuned process adds. Which model translates better is a different question.

Fable’s review of its own translation found nothing, so the reviewed version is byte for byte the same as the plain one. That is how two identical texts ended up in the blind comparison without anyone planning it.

What I learned about measuring itself

The same reader doesn’t repeat its findings

The translations with a review added differ from the plain ones by a few paragraphs, and the same reader read both. In practice that is the same text read twice.

One reviewer, the same text, read twice

Only 68% of the paragraphs with a remark on the first read have a remark on the second. On average 14 findings per text appear or vanish in paragraphs the review never touched. So on 296 paragraphs the finding count has an error of about 10 to 15. The difference between 15 and 22 findings doesn’t exist. The difference between 15 and 92 does.

The second round gave a cleaner experiment. Fable’s two identical texts went into the blind comparison under different letters. The judges saw them as two different candidates and gave them identical scores on all six criteria. So the disagreement comes from splitting the reading into several sittings, and when a judge sees everything at once, it is precise. Hence the rule: score everything in one call.

How I almost published an error that doesn’t exist

What one rule in the instruction was worth

The first scoring instruction checked gender the way an English speaker would. In English the ship and the galley are both “it”, so a translation that says “ще го огледаме” (masculine) about the ship and “ще я потопя” (feminine) about the galley looked confused. By that rule almost every translation had a gender error, and for a while I had it written down as a finding. There is no such error. “Галера” is feminine, “кораб” is masculine, so both are right. Seven of the ten readers noticed it on their own, and that is how I noticed it too. After I rewrote the rule, “almost all” shrank to one real gender error in the blind comparison. The instruction has to carry the rules of the language even when most readers know them, because the rest will fill the table with errors that aren’t there.

The hidden bill

The dollars cover only the calls to the models. All the reading, the scoring, Fable’s translations and this article ran on a monthly subscription, which isn’t paid per call and has its own ceiling.

The bill that never reaches an invoice

The recorded usage is over 3 million tokens across 19 tasks. One read of the story against the original is about 137,000 tokens, one translation about 190,000. The subscription ceiling isn’t shown in tokens, only in percentages. I hit it twice, and twice the work stopped halfway. Nothing was lost, once because the files had been saved in time and once because the task could be resumed. So now every task saves its result as soon as it has one.

What this study doesn’t prove

  • Everything rests on one story of 12,079 words. The next step is a longer book with the top three models.
  • Most numbers after the first round come from a single reader, and that reader repeats only 68% of itself.
  • The pipeline result is four out of four in the same direction. It is a signal until a second book says otherwise.
  • Thirteen candidates on a ten-point scale squash the top. At one judge three of them share the ten.
  • The stronger check of whether the models know the human translation, with a book that has no published translation, hasn’t been done yet.

So which model translates fiction best

Overall, gpt-6-astra showed the best results in every round, but at a price. It is the most expensive model, more than 12 times the price of the cheapest one.

AI literary translator ranking: points and price

gemini-3.8-flash gave the best value for money: it works fast, cheap and well enough.

I also saw that using a pipeline during translation improves the results significantly, and if I set out to translate fiction myself, I would definitely do it with something like wenyi, even if only written up as a skill or a custom agent in the tool I use.

The short version of this article is here: In Search of the Best A.I. Literary Translator.

Leave a Reply

Your email address will not be published. Required fields are marked *.

*
*