Make an audiobook with an AI voice: a method that holds for five hours
The publisher behind Aievalu produced hundreds of hours of French audiobooks with synthetic voices before opening the tool to everyone. Here is what that production taught us: where an AI voice shines, where it stumbles, and how to prepare a manuscript so it can be listened to in one sitting.
Prepare the manuscript first
The engine reads what you give it, faithfully, including what was never meant to be read. Before the first generation, remove page numbers, running headers, footnote markers and the footnotes themselves, or move them to the end of the chapter if they must be heard. Check the punctuation of dialogue: a well-placed dash gives a breath, a stray quotation mark gives an odd silence. Write numbers and dates the way you want to hear them, "eighteen forty-eight" rather than "1848" when the reading should feel literary, and expand rare abbreviations. How the engine works, its strengths and its limits are detailed on the AI voice generator page.
Split into 1,500-character segments
One generation covers 1,500 characters, about a minute and a half of listening. A 12,000-character chapter is therefore eight chained generations. Cut at the end of a paragraph, never mid-sentence: the voice shapes its intonation on punctuation, and a hard cut is audible in the edit. Number your segments in order (01-03-intro, 01-03-next…), keep the same voice and the same tone from the first page to the last, then assemble chapter by chapter in your audio software. The counting rule is simple, spaces and punctuation included; it is explained in the FAQ on character counting.
Choose a voice that lasts
A voice that charms for ten seconds can tire over five hours. For long narration, the neutral tone is almost always the right call: Paul Neutral or Jane Neutral in English, Marie Neutre in French. Save the cheerful or confident tones for passages that truly ask for them, or for children's books. Listen to the samples on the voices page, which all read the same text, then test one page of your own manuscript with the free trial, two sentences with no account, before committing a whole chapter.
Pace is set with punctuation
There is no "slower" button: punctuation is your only control, and it is enough. A full stop closes the sentence and marks a real pause, a semicolon or colon a shorter one, an ellipsis a slowdown. For a scene change, leave a blank line or end the segment there. Very long sentences, common in nineteenth-century prose, often benefit from extra punctuation in the read version without touching the published text. For reading long texts in general, see the guide on text to speech with natural voices.
Export, assemble, level
Export WAV during production: the file takes processing better, and you convert to MP3 once, at the end. Assemble each chapter in a separate file, named in order, then normalise the whole book to the same loudness so the listener never touches the volume between chapters. Leave a second of silence at the start and end of each chapter, long enough for the listening app to move on. Files stay in your history for 90 days, as the FAQ on retention explains; download them as you go.
What a novel costs
A 300,000-character novel gives about five hours of listening. With the 300,000-character pack at €24.99, valid for twelve months, the whole book fits in a single purchase; three €9 packs of 100,000 characters do the same job. A series or a steady output is better served by the Pro subscription at €49.99 a month for 800,000 characters. All prices are on the pricing page, in euros with no VAT applicable. Compare with a human narrator: several hundred euros per finished hour, for a different result, and that is exactly the trade-off to weigh book by book.
Rights: the text, then the audio
The generated audio is yours and commercial use comes with paid credits, as the FAQ on usage rights states. The text, however, must be yours or free of rights: your own manuscript, a public-domain work, or written permission from the rights holder. Before publishing, check each distribution platform's policy: some restrict or label AI-narrated audiobooks, and those rules change. Every MP3 carries a technical marker in its metadata, and the European AI Act asks you to tell listeners when content could be mistaken for a human reading: a short notice at the start of the book is enough.
And for other long formats
The same method works for a narrated podcast, a multi-hour online course described on the e-learning voice-over page, or a documentary video for YouTube. At the other end of the scale, the very short messages of a phone system call for a different kind of writing, detailed on their own page.