A voice-over for your podcast, no studio required
Intro, jingle, narrated episode, audio version of an article: a synthetic voice covers everything in a podcast that doesn't need to be improvised. Here is how to write it, produce it and fit it into your episodes.
Intros, outros and jingles
The opening is the most rewritten part of a podcast and the least pleasant to re-record. With an AI voice you lock a clean intro in one generation, adapt it per season by changing one sentence, and produce the outro, the subscribe reminder or the sponsor read in the same voice. A twenty-second intro is about 300 characters; the counting rule is simple, spaces and punctuation included.
Narrated episodes and audio articles
Some formats are fully scripted: columns, stories, news digests, audio versions of a newsletter or a blog. They suit a regular synthetic narrator published on schedule with no recording constraints. How the engine works and where it falls short is described on the AI voice generator page; for reading long texts in general, see text to speech with natural voices.
Mixing a synthetic narrator with human hosts
The mix works when roles are clear: the AI voice carries the frame (opening, transitions, sidebars, listener questions), the hosts carry the conversation. Pick a voice whose timbre contrasts with the presenters so the ear knows instantly who is speaking. Paul Neutral and Jane, in English, and Marie Neutre, in French, hold up over long durations without tiring the listener; the samples are on the voices page, all reading the same text so you can compare.
Writing for the ear
Text to be heard is not text to be read. Short sentences, one subject each, present tense, concrete words. Read your script aloud before generating it: if you run out of breath, so will the voice. Punctuation is your mixing desk, a full stop is a real pause, a comma a breath, an ellipsis a slowdown. Common acronyms are read the way people say them; test rare terms with the free trial, two sentences with no account.
Episode length and the per-generation limit
One generation covers 1,500 characters, about a minute and a half. A ten-minute narrated episode is therefore seven or eight chained generations in your studio, with the same voice: split by sequence, number the files in order, then assemble them in your editing software. Your history keeps each file for 90 days (see how long audio is kept), long enough to finish the episode.
Loudness and export
Export MP3 for distribution and WAV when you mix with human voices or music: WAV takes processing better. Normalise the whole episode to the usual target of listening platforms, keep the theme music several decibels under the voice and leave half a second of silence before the first sentence. The same habits apply to video, as detailed in the YouTube voice-over guide.
Telling your listeners
Say it plainly, in the show notes or in the intro: "narration generated by a synthetic voice". The European AI Act requires this when synthetic content could be mistaken for the real thing, and listeners appreciate the honesty. Every MP3 also carries a technical marker in its metadata.
What it costs
A ten-minute narrated segment, about 10,000 characters, costs €0.90 with the €9 pack of 100,000 characters, valid for twelve months. A fully narrated weekly podcast fits in the Medium subscription at €14.99 a month for 200,000 characters; intros and outros alone are covered by the pack. Everything is on the pricing page, and commercial use, sponsor reads included, comes with paid credits as the FAQ on usage rights explains. For short clips cut from your episodes, see AI voice for TikTok and Reels; for teaching content, voice-over for e-learning.