Simultaneous interpretation works by having one interpreter listen, convert and speak at the same time, running a few seconds behind the speaker — a deliberate lag the profession calls décalage, or ear-voice span. The interpreter sits in a soundproof booth, hears the source language through headphones, and delivers the target language into a microphone while the speaker is still talking. The audience selects a channel and hears their own language live.

That's the mechanism. The interesting part is why it's so hard.

If you've watched a UN session and wondered how the voice in your ear keeps pace with a speaker it couldn't possibly have heard yet, your instinct is right: it hasn't. It's trailing by a few seconds, and every one of those seconds is doing work. The technique was first deployed at scale at the Nuremberg Trials, which opened in 1945, and the profession has spent the decades since building an entire apparatus — booths, consoles, relay chains, two-person teams — around a task that from the outside looks like fast talking.

This guide walks through what actually happens between the speaker's mouth and the listener's ear: the cognitive chain, the hardware, the reason interpreters tap out every half hour, and how streaming AI translation handles the same problem with a different set of trade-offs. If you're deciding whether your next multilingual meeting needs a booked interpreter or a browser tab, the last two sections are the ones that matter.

Key Takeaways

What Simultaneous Interpretation Actually Is

Simultaneous interpretation is the live conversion of spoken language into another spoken language while the original speaker continues talking. No pauses are built in. The event runs at its natural pace, and the interpretation rides alongside it.

Two distinctions clear up most of the confusion around the term.

Simultaneous vs consecutive interpretation

Consecutive interpretation works in turns. The speaker says a few sentences, stops, and the interpreter delivers them — often working from a specialised note-taking shorthand. It's more accurate sentence by sentence, because the interpreter has heard the complete thought before rendering it. The cost is time: a consecutive session runs roughly twice as long as the same content in one language.

Simultaneous interpretation buys that time back. Nobody waits. A three-hour conference stays a three-hour conference in six languages. The cost is cognitive: the interpreter is committing to a rendering before the sentence has finished arriving.

Interpretation vs translation

Translation is written and revisable. You draft, reconsider, look something up, and revise. Interpretation is spoken and irreversible — once it's out of the microphone, it's in the record. This is the single most useful thing to understand about the profession. Interpreters aren't translators who work faster; they're doing a structurally different job under conditions that remove every safety net translation relies on.

How Simultaneous Interpretation Works, Step by Step

Break the process into its overlapping stages and the difficulty becomes obvious. All four of these are running at once, on different parts of the same sentence.

1. Listening and prediction

The interpreter takes in the source audio and starts predicting where the sentence is heading — from syntax, from the speaker's argument structure, from the subject matter briefing they read beforehand. Prediction is not optional. In language pairs where the verb arrives late, an interpreter who waits for full grammatical certainty will never catch up.

This is also why interpreters demand preparation materials. An interpreter who has read the agenda, the slide deck and the glossary predicts better, and prediction is most of the job.

2. Holding the sentence — décalage

The interpreter deliberately lags the speaker by a few seconds. Too short a lag and they commit to a reading of the sentence that the next clause contradicts. Too long and working memory overflows: they're now holding two unresolved sentences while a third arrives.

Skilled interpreters vary this gap consciously. They stretch it when a speaker's structure is convoluted and compress it when the content is predictable — numbers, names and lists get delivered almost immediately, because those are the items memory drops first.

3. Reformulating

The interpreter converts meaning, not words. Idiom, register and cultural framing all get re-cast for the target audience. A Japanese client saying ちょっと難しいです is literally "that's a little difficult." A good interpreter conveys what it functions as in the room, which is usually a polite no. Word-for-word would be accurate and useless.

4. Speaking — and monitoring

The interpreter delivers into the microphone while still listening to incoming audio, and simultaneously monitors their own output for errors. That last loop is the one that fails first under fatigue. It's also the reason booth acoustics and clean source audio matter so much: an interpreter fighting a bad feed loses the capacity they need for self-correction.

Need the same outcome without booking a booth? MirrorCaption streams transcription and translation next to your meeting in 50+ selectable languages. Open it in your browser →

The Equipment Behind the Booth

Simultaneous interpretation at conference scale is an infrastructure problem as much as a linguistic one.

The booth and the console

Interpreters work from sound-isolated booths with a clear view of the room and the screen — seeing the speaker's gestures and slides is part of the input. The interpreter console handles channel selection, a cough button for muting, volume and a relay selector. Booth and equipment specifications are covered by published ISO standards, which exist precisely because poor acoustics and cramped booths measurably degrade output over a working day.

Relay interpreting

When no interpreter available covers a given language pair directly — Finnish to Korean, say — the booth works from a colleague's output instead of the original. Finnish goes into English, and the Korean booth interprets the English.

Relay solves an otherwise impossible staffing problem, and it has an obvious cost: each hop adds delay and one more opportunity for meaning to shift. Anyone who has played telephone knows the failure mode. Professional bodies such as AIIC, the international association of conference interpreters, publish working standards covering exactly these conditions.

Remote simultaneous interpretation

Remote simultaneous interpretation (RSI) moves the booth into software. Interpreters work from home or a hub, audio streams to them over the internet, and listeners pick a channel in the meeting client rather than on a receiver. Zoom, for instance, offers a language interpretation feature that lets a host assign interpreters to their own audio channels — you can check the current requirements in Zoom's support documentation, since availability depends on the host's account.

RSI made interpretation dramatically more available. It also made audio quality everyone's problem: a participant on a bad laptop microphone in a reverberant kitchen is now degrading the input for every interpreter on the call.

Illustrative scenario

Priya is coordinating a two-day regulatory workshop with participants in three languages. She books a team of interpreters and confirms the booth setup weeks ahead. On day one, a presenter dials in from an airport lounge on a phone headset.

The interpreters can still work, but they're now spending capacity on decoding a noisy feed instead of on nuance. By mid-afternoon Priya's team is asking presenters to repeat figures. Nothing was misconfigured; the source audio was simply worse than the workflow assumed. Preparation materials and clean microphones do more for interpretation quality than almost anything downstream of them.

Why Simultaneous Interpreters Work in Pairs

Book a full-day event with one interpreter per language and you have made a scheduling error, not a saving.

The task loads listening, comprehension, working memory, reformulation and speech production onto one person concurrently, with no recovery gaps. Output quality declines with time on microphone, and — this is the dangerous part — it declines before the interpreter notices. Self-monitoring degrades along with everything else.

So interpreters work in teams and swap at regular intervals. The resting interpreter doesn't leave; they stay in the booth writing down numbers and proper nouns, looking up terminology, and watching for trouble. The booth is a two-person system where only one person is audible at a time.

This is also the honest answer to "why does interpretation cost what it costs." You're not booking one professional for a day. You're booking a team, their preparation time, and usually the equipment too.

Human Interpretation vs Live AI Translation

These get compared constantly, usually badly, because the two things produce different outputs. Here's the comparison stated in the terms that actually differ.

DimensionHuman simultaneous interpretationLive AI translation (MirrorCaption)
OutputFinished spoken language, register-adjustedStreaming text on screen, plus optional spoken output
LagA few seconds of décalage, by designSub-second, word by word, self-correcting as context arrives
BookingDays to weeks ahead, plus prep materialsOpen a browser tab when the meeting starts
CoverageLimited by which interpreters you can staff50+ selectable languages, switchable mid-session
RecordEphemeral unless separately recordedSearchable transcript, exportable, side-by-side with the original
CostTeam rates plus equipment, per event€99 one-time Premium with 200h of hosted transcription credit included
HandlesAmbiguity, humour, politics, deliberate vaguenessVolume, cost, availability, and meetings nobody would staff

Note the row that matters most: lag measures different things in each column. An interpreter's few seconds buy you a complete, culturally re-cast spoken sentence. A live transcription layer's sub-second lag puts partial text on screen that firms up as more context arrives. Comparing the two numbers directly is a category error — they're not the same unit of work.

Where interpreters remain unmatched is exactly where language is doing something other than transmitting information. Diplomatic hedging, a joke that lands on a cultural reference, a witness choosing their words carefully — these need a human who understands the stakes in the room. We cover the measurable side of this in more depth in our guide to real-time translation accuracy.

How Live AI Translation Works Differently

The AI pipeline decomposes the interpreter's single overlapped task into stages that run concurrently across separate systems.

Streaming speech-to-text. Audio streams continuously to a transcription engine that emits partial results word by word rather than waiting for the sentence to end. Those partials revise themselves as more context arrives — the same self-correction an interpreter does internally, made visible on screen.

Context-aware translation. Each segment is translated with the previous few segments as context, so pronouns and terminology stay consistent instead of resetting every sentence.

Optional spoken output. With Speak Translations enabled, MirrorCaption can read your translated speech aloud in the target language while the exchange is still live. You speak Mandarin, the other side hears English, and they answer in English while you read Mandarin. Playback can run through the laptop speaker, a paired phone, or — on the Mac client — a virtual microphone so Zoom, Meet or Teams receives the translated audio as microphone input.

Capture without a bot. Meet mode captures the meeting tab's audio in desktop Chrome or Microsoft Edge, so nothing joins the call as a participant. Talk mode uses the phone microphone for face-to-face conversation and runs as one continuous session — you start it once and both people take turns, rather than tapping a button for every sentence.

Illustrative scenario

Daniel runs partnerships from Berlin and takes a Thursday call with a supplier in Osaka. It's a routine check-in — nobody was ever going to book an interpreter for it, and historically he just accepted that half the nuance was lost.

He opens MirrorCaption in a second tab. The Japanese side of the conversation appears next to its English translation while his counterpart is still speaking, so when they say something that reads as hesitation rather than agreement, he can ask about it in the same call instead of discovering it in a transcript on Friday. The point isn't that he replaced an interpreter. It's that this meeting never had one.

Ready to test the difference on a real call? Every account starts with 1 free hour — no credit card, no monthly reset. Start free →

Choosing Between an Interpreter and Live AI Translation

The decision is usually straightforward once you name what's at stake.

Book human simultaneous interpretation when the words carry legal or clinical consequence, when a misreading is expensive or unrecoverable, when the setting is formal enough that a machine reading would be inappropriate, or when someone must be accountable for the rendering. Diplomacy, court proceedings, regulated medical consent, board-level negotiation. Our page on medical interpreting in the browser is explicit about where the professional line sits.

Use live AI translation when the meeting is one of the many that would otherwise happen in a language somebody in the room only half-follows: recurring standups, supplier calls, customer support, lectures, interviews, in-person conversations while travelling. Also when you need a searchable record afterwards, which ephemeral interpretation doesn't give you — see live captions vs transcripts for that distinction.

Use both when the event is interpreted but participants still want text they can search, quote and share. Interpretation and live captions aren't mutually exclusive, and for distributed multilingual teams the combination often works better than either alone.

Frequently Asked Questions

What's the difference between simultaneous and consecutive interpretation?

Simultaneous interpretation happens while the speaker is still talking, so the event runs at normal speed. Consecutive interpretation happens in the gaps: the speaker pauses, the interpreter delivers, the speaker resumes. Consecutive is more accurate per sentence but roughly doubles the length of the session.

How far behind the speaker is a simultaneous interpreter?

A few seconds, and the gap is deliberate. It's called décalage, or ear-voice span. The interpreter needs enough of the sentence to know where it's going before committing to a rendering — too short and they guess wrong, too long and working memory overflows.

Why do simultaneous interpreters work in pairs?

Because the task loads listening, comprehension, memory and speech production onto one person at the same time, and quality degrades with sustained time on microphone. Interpreters work in teams and swap at regular intervals, with the resting partner staying in the booth to support the active one.

Can AI replace simultaneous interpreters?

Not for high-stakes settings like diplomacy, courts or regulated medical consent, where accountability and liability matter. AI does replace the interpreter for the far larger set of everyday meetings that never had one because booking a human was too slow or too expensive.

Do I need special equipment for simultaneous interpretation?

For on-site human interpretation, yes: a booth, an interpreter console, and receivers for the audience. Remote simultaneous interpretation replaces the hardware with a software platform. Live AI translation replaces it with a browser tab.

How do I get live translation in a Zoom or Teams meeting without hiring an interpreter?

Open MirrorCaption in a second browser tab alongside the call. Meet mode captures the meeting tab's audio in desktop Chrome or Microsoft Edge and streams transcription and translation next to it. No bot joins the meeting and nobody else has to install anything.

The Bottom Line

Simultaneous interpretation works because a trained professional deliberately falls a few seconds behind a speaker and uses that gap to listen, predict, reformulate and deliver — all at once, in a booth, with a partner ready to take over. The décalage isn't lag to be engineered away. It's the working space the whole technique depends on.

Understanding that changes how you evaluate the alternatives. A live transcription layer isn't a cheaper interpreter; it's a different instrument that solves a different problem — the enormous number of cross-language meetings that were never going to be staffed at all. For those, MirrorCaption puts transcription and translation next to your call in 50+ selectable languages, with no bot in the meeting and nothing for other participants to install. If you're weighing the options across tools, our 2026 meeting translator comparison lays them out side by side.

Book the interpreter when the stakes demand one. For every other meeting, open a tab.

Try MirrorCaption Free

1 free hour to try. No credit card. No monthly reset. Nothing to install to get started.

Get Started Free