Benchmark · October 2026

Live translation, five languages at once.

By Saamer Mansoor, ConferenceCaptioning · Updated · Method · Data · Changelog

Captions are only useful to a multilingual audience if the translations keep up with the speaker. We fed an English transcript at conversational speed (3 words per second) into ConferenceCaptioning’s fastest translation mode, Blazing, with one to five target languages running together, and timed how long each word took to appear in every language on screen. Translation runs on the device; nothing is sent to a translation service.

10-core Apple silicon Mac16 GB memoryEnglish to Spanish, French, German, Japanese, Portuguese (Brazil) 3 words per second1–5 languages at once30 timed runs
Translation step, 5 languages
0.27 s
Blazing · per translated word, 95th percentile 0.54 s
Speech to OBS overlay
0.9 s
average, 5 languages, fastest refresh (100 ms), speech recognition included
Words timed
8,840
across 30 timed runs of 40 s each
Privacy
On-device
translation runs locally; text is not sent to a translation service

The short version

With five languages translating together, the average word appeared on screen 0.27 s after it was available, 0.54 s for 95 out of 100 words, and the slowest word in any run took 0.61 s. Adding languages from one to five changed the average delay only slightly (see the table), because each translation takes a few tens of milliseconds on the device. That is well under a second, so for this part of the pipeline “five languages in under a second” holds. With the refresh interval turned down to 50 ms, five languages averaged 0.08 s (see faster refresh), at the cost of more processor time. This is the translation step only; the delay from speech to the screen, with speech recognition included, is in the next section.

From speech to the OBS overlay

The test that matters most for an event: how long after a word is spoken does it show on the page you put into OBS, a vMix browser source or a stage display? We replayed a recorded news conference through the same local pages the app serves, with the on-device speech recognizer in the loop, and timed every word from the moment it was spoken to the moment it was visible in the overlay (including the scroll animation). Translations were timed the same way, in the last of the translated languages, which is the slowest one.

Refresh settingCaptions only1 language3 languages5 languages
Default (500 ms)1.3 s (95%: 1.9 s)1.5 s (95%: 2.3 s)1.6 s (95%: 2.3 s)1.8 s (95%: 2.4 s)
Faster (250 ms)0.9 s (95%: 1.5 s)1.1 s (95%: 1.6 s)1.1 s (95%: 1.7 s)1.2 s (95%: 1.7 s)
Fastest (100 ms)0.7 s (95%: 1.3 s)0.8 s (95%: 1.3 s)0.8 s (95%: 1.4 s)0.9 s (95%: 1.4 s)

Average time from a word being spoken to it being visible, with the 95th percentile in brackets. “Refresh setting” is how often both the translation page and the overlay check for new text (the pollInterval option in the Caption URL Generator); 60 s per run, 3 runs per setting, first 10 s discarded.

What this means

  • At the fastest setting (100 ms), words appear on average within 0.7 s of being spoken for captions and 0.9 s with five translated languages, and 95% of words within 1.4 s.
  • At the default setting the average is 1.3 s for captions and 1.8 s for five languages, so we do not claim “under a second” for the default.
  • Roughly 0.5 s of every figure is the speech recognizer itself (the Rapid engine); the rest is the pipeline, which the refresh setting controls. Faster refresh uses more processor time (browser CPU rose from about 21% to 31% of one core for one language).
  • Languages add little: going from one to five languages adds about 0.1 s at the fastest setting and 0.2 s at the default.

What this does not include

  • The microphone, the room and the sound mixer: the audio is a clean recording played at real-time speed, and a recorded recognizer run is replayed so every setting sees the same speech. Real rooms add noise that makes recognition less accurate, not necessarily slower.
  • OBS itself: it adds about one video frame (17–33 ms) on top. Attendee phones and the cloud stream are separate paths and are not measured here.
  • One clip (a clean press conference), one Mac (Apple M4), English into Spanish, French, German, Japanese and Portuguese. Only words with a known spoken time (about 60% of the clip) are timed.

Translation delay

Time from the moment a word became available in the transcript to the moment its translation was on screen, averaged over every word, with the line marking the 95th percentile. Lower is better. Blazing is the sky-blue bars; grey bars are our other on-device translation mode (Native Translation) for scale.

1 languageBlazing · 95th percentile 0.40 s
2 languagesBlazing · 95th percentile 0.47 s
3 languagesBlazing · 95th percentile 0.48 s
4 languagesBlazing · 95th percentile 0.54 s
5 languagesBlazing · 95th percentile 0.54 s
Native Translation, 1 languagesame Mac, different harness, about
Native Translation, 4 languagessame Mac, different harness, about

Bars run from 0 to 4 seconds. The grey Native Translation bars come from an earlier test on the same Mac with a different harness (translation_benchmark_report.md in our repository), so compare the shape, not the last digit.

All the numbers

Refreshing every 500 ms (the default). Each row pools 6 runs of 40 measured seconds (the first 10 s of each 50 s run are warm-up and discarded), across continuous speech and speech with a 2 s pause after every sentence. “Call time” is how long one translation request took.

LanguagesMean delayMedian95th pctSlowestCall time (mean)Call time (95th)Avg CPUPeak memoryRuns
1 language0.25 s0.20 s0.40 s0.54 s22 ms34 ms15%1.3 GB6
2 languages0.25 s0.21 s0.47 s0.56 s28 ms50 ms14%1.2 GB6
3 languages0.25 s0.22 s0.48 s0.58 s38 ms65 ms14%1.3 GB6
4 languages0.25 s0.23 s0.54 s0.59 s45 ms78 ms14%1.3 GB6
5 languages0.27 s0.24 s0.54 s0.61 s51 ms90 ms14%1.4 GB6

CPU is the translation process as a share of one core. Memory is the whole translation process including its language models. Continuous and paused speech gave the same result at 5 languages (mean 0.27 s and 0.25 s). Errors across all timed runs: 0.

With a faster refresh

The refresh rate (how often the translation page checks for new words) is a setting. Faster refresh shows words sooner but uses more processor time. Continuous speech, one and five languages:

Refresh1 language: mean95th pctSlowest5 languages: mean95th pctSlowestCPU (1 language)
50 ms0.06 s0.07 s0.09 s0.08 s0.12 s0.15 s28%
75 ms0.07 s0.10 s0.11 s0.10 s0.15 s0.17 s25%
100 ms0.08 s0.12 s0.12 s0.11 s0.16 s0.19 s23%
125 ms0.09 s0.14 s0.16 s0.12 s0.19 s0.22 s20%
150 ms0.10 s0.17 s0.18 s0.13 s0.21 s0.25 s19%
200 ms0.14 s0.22 s0.24 s0.17 s0.28 s0.36 s11%
500 ms0.25 s0.40 s0.54 s0.27 s0.56 s0.60 s15%

Supported languages and expected speed

67 languages can be translated on the device, in three speed groups. Each language is listed once, under the fastest mode that supports it. Only the five Blazing languages named in the table were timed directly in this test; the other times are estimates from the matching measurements and could differ for individual languages.

ModeLanguagesExpected translation delayTested withBasis
Blazing44about 0.3–0.5 sup to 5 languagesMeasured for Spanish, French, German, Japanese and Portuguese (Brazil); estimated for the rest of the group
Native Translation1about 1.0–1.5 s for one language, 3–4 s for four1–4 languagesMeasured end to end on a similar Mac; the only on-device option for these languages
Language Pack22about 1.0–1.5 s for one language, 3.4–3.8 s for four1–4 languagesEstimated from our test of the same pipeline on Spanish, French, German and Japanese; one-time 620 MB download

Blazing (44 languages, fastest)

ArabicBengaliBulgarianChineseCroatianCzechDanishDutchEnglishEstonianFinnishFrenchGermanGreekGujaratiHebrewHindiHungarianIndonesianItalianJapaneseKannadaKoreanLatvianLithuanianMalayalamMarathiNorwegianPolishPortugueseRomanianRussianSlovakSlovenianSpanishSwahiliSwedishTamilTeluguThaiTurkishUkrainianUrduVietnamese

Native Translation (1 language, not in Blazing)

Malay

Language Pack (22 languages that no other on-device mode covers)

AlbanianAmharicAzerbaijaniBosnianBurmeseDariFilipinoHausaKazakhKhmerKurdish (Kurmanji)Kurdish (Sorani)LaoNepaliPashtoPersianPunjabiSinhalaSomaliTigrinyaUzbekYoruba

Regional variants (for example Spanish for Mexico or Chile, or English from Australia or India) are counted under their language. Languages are added through a one-time on-device download. The delay is the translation step only, as described above.

What is not measured

Not in these numbers

  • Speech recognition (in this section). The transcript here is scripted. In a live event the words first have to be recognised, which takes about 0.5 s on average for our fastest engine (Rapid) and several seconds for the more accurate Mix engine on continuous speech (see the captioning benchmark). The speech-to-screen section includes it.
  • Delivery to the audience. Sending results to attendees’ phones and the attendee refresh interval come on top.
  • Real audio, other languages and other machines. One transcript, five target languages, one Mac. Quality of the translations was not scored here.

How to read the delay

  • The delay is measured per word: each word’s age when it first appears in the translated text. Words that arrive early in a refresh wait longer than the last one, which is included.
  • Runs where something else heavy was using the Mac are excluded (3 of 39 runs) so the figures reflect the engine, not background load.

Methodology

The input

  • A fixed English transcript (a public speech) released word by word at 3 words per second, about 180 words per minute, a brisk conference pace.
  • Two patterns: continuous speech, and a 2 s pause after every sentence.

The translation

  • Blazing, ConferenceCaptioning’s fastest on-device translation mode, on the same translation page attendees and stage screens use.
  • Target languages added in the order Spanish, French, German, Japanese, Portuguese (Brazil).

The timing

  • Each word’s creation time is known from the script. Its delay is the time until it is visible in the translated text on screen.
  • 50 s per run, first 10 s discarded; 3 repeats per setting; refresh interval 500 ms by default, with 50–200 ms in the refresh table.

The machine

  • 10-core Apple silicon Mac, 16 GB, macOS 27.
  • Nothing else was run on purpose during the final runs; runs that were disturbed are excluded and counted above.

Data and files

License

The data is licensed CC BY 4.0: share and adapt it, including commercially, with credit to ConferenceCaptioning and a link to https://conferencecaptioning.com/multilingual-translation-benchmark/.

Cite this page

Saamer Mansoor, ConferenceCaptioning. “Live translation speed benchmark: five languages at once.” October 2026. https://conferencecaptioning.com/multilingual-translation-benchmark/

Changelog

  • October 8, 2026: added “From speech to the OBS overlay”, timing words from speech to the overlay with speech recognition included.
  • October 5, 2026: first published.

Need several languages in the room?

Tell us your languages, your venue and how attendees will follow along and we’ll help you plan the setup.

Captioning benchmark → Start free →