What a human voice still does better
A skilled narrator makes choices a synthetic voice does not. They know which word in a sentence carries the meaning, when to slow down for a difficult name, and how long a pause should sit before the next idea. Across a full guide those choices compound into something that feels authored rather than generated.
Human voice matters most where the material is emotionally weighted. Memorial content, personal testimony, difficult history, and anything connected to a living community are places where synthetic delivery can sound indifferent at exactly the moment it should not.
There is also a provenance argument. A guide narrated by a curator, a local historian, or a community member is not just narration, it is part of the interpretation. That cannot be synthesised.
Where synthetic narration earns its place
The economics change sharply once you go multilingual. Booking a native-speaking voice actor in each of eight languages means eight bookings, eight studio sessions, eight rounds of review, and eight schedules to coordinate. Synthetic narration collapses that into a rendering step.
Update cost matters just as much as launch cost. Museums change labels, correct facts, and rotate exhibitions. With a booked voice actor, a single corrected sentence can mean rebooking the original speaker to match the take, and if that person is unavailable the fix may not be possible at all. With synthetic narration, a corrected line is regenerated in minutes.
That difference tends to decide the question for temporary exhibitions, for content that is expected to change, and for languages where a museum could not realistically fund a studio session at all.
- Eight languages means eight bookings with a human voice, one rendering step with synthetic
- Corrections do not require the original speaker to be available again
- Makes languages viable that a small museum could not otherwise afford
- Temporary exhibitions can launch and change without studio time
The mixed approach most museums land on
In practice the split is rarely all or nothing. A common pattern is a human voice for the main visitor language, where most listening happens and where the institution's character matters most, and synthetic narration for the translated versions that make the guide accessible to everyone else.
That approach puts the budget where the listening is, without letting cost decide that visitors in other languages get nothing at all. It also reflects an uncomfortable truth: the realistic alternative to a synthetic translation is usually not a human translation, it is no translation.
The same logic applies within a guide. Signature stops and oral history can carry a human voice while routine wayfinding and room introductions are synthesised.
Getting good results from synthetic narration
Synthetic voices fail in predictable ways, and most of those failures are fixable in the script rather than the settings. Place names, personal names, and technical terms are the usual problem, particularly Māori and other Indigenous names, which generic voices routinely mangle.
Build a pronunciation list alongside the script. Any name that matters should be checked in the rendered audio before the guide goes live, not after a visitor points it out. Where a name is culturally significant, getting it wrong is a real failure rather than a cosmetic one, and it is worth using a human recording for those specific lines if the synthesis cannot be corrected.
Punctuation also does more work than people expect. Commas and full stops shape pacing, and splitting a long sentence often fixes a delivery problem faster than adjusting voice parameters.
- Keep a pronunciation list for names, places, and technical terms
- Listen to every name in the rendered audio before launch
- Use punctuation to control pacing rather than fighting the voice settings
- Consider a human recording for culturally significant names
Whichever you choose, record the source properly
Audio quality is judged unconsciously and quickly. Visitors will not identify a room reflection or a noise floor, but they will decide the guide feels amateur and turn it off.
If you are recording a human voice, a treated room matters more than an expensive microphone. Soft furnishings, a small space, and a consistent distance from the microphone will beat a good microphone in a hard-walled gallery. Record everything in one session where possible, because rooms and voices do not match perfectly across days.
Keep the unprocessed source files. Formats change, delivery platforms change, and a museum that kept its masters can re-export a guide years later without recording anything again.
Frequently asked questions
Is AI narration good enough for a museum audio guide?
For most informational content, yes, and it is now the only realistic way many smaller museums can offer several languages. It is weaker for emotionally weighted material and for names it has not been corrected on.
How much does a voice actor cost for an audio guide?
It varies widely by market, experience, and usage rights, and the figure that matters is total cost across every language and every future correction rather than a single session rate. Ask for pricing that includes revisions.
Can we mix human and AI voices in the same guide?
Yes, and many museums do. A common split is a human voice for the primary language and synthetic narration for translated versions.
What audio format should the finished guide use?
Deliver compressed audio for visitors, since it loads faster on mobile data, but keep uncompressed masters so the guide can be re-exported later without re-recording.
Make this easier for your visitors
Journey Pal helps museums, heritage sites, zoos, and attractions deliver audio guides and translations through QR codes, with no app download required.

