
P2P | September 2026 | 5.35 MB
安装方法:安装预配置版本,所有功能均已激活。
将任何文本转换为自然语音——59 种语音,9 种语言,完全离线,无需注册,无需订阅,无需上传。导出带字幕的 MP3 文件。
内置 59 种语音,9 种语言。
包含 54 种 24 kHz 自然录音室语音——英语(美式和英式)、西班牙语、法语、意大利语、葡萄牙语、印地语、日语和普通话——并列出 Windows 系统已安装的所有语音。无需下载,无需注册。所有功能均在处理器上运行,无需独立显卡。
点击“浏览所有语音”按钮,即可按名称或语言搜索,按男声或女声筛选,收藏您喜欢的语音,并立即收听其对应的语音。
为每个角色赋予专属声音。
“角色配音”界面可将故事转化为表演。粘贴包含引用对话的散文,它会自动将旁白与语音分离,并根据“玛伦说”或“艾莉丝摇了摇头”等短语判断说话者。粘贴包含“NAME: lines”的脚本,它就会朗读出来。为每个角色分配一个声音,按下“录制”按钮,每行台词都会以各自的声音朗读,并合并成一个录音。
混合独一无二的声音:
使用滑块混合任意两种同语种的录音室声音。这会在模型定义声音的层面上将它们混合,因此最终得到的是一个真正全新的声音,而不是经过过滤的副本。
编写和旁白:
输入或粘贴最多 20 万个字符。长段落会自动分割,并显示进度条和剩余时间的实时预估。使用时间标尺聆听实际波形,点击任意位置即可跳过,并在听起来合适时导出。
您的作品会在输入时自动保存。按名称保存文档,导入文本文件,并在长篇稿件中使用查找和替换功能。
导出 MP3、WAV、FLAC 或 OGG 格式文件 – 带字幕:
MP3 文件大小约为 WAV 的五分之一,并且可以在所有设备上播放。FLAC 是无损格式。勾选复选框,即可在音频旁边生成 .srt 字幕文件、.vtt 网络字幕或带时间码的文本稿。时间码以实际生成的音频为准,因此字幕与音频完全同步——可直接导入 Premiere、Resolve、CapCut 或 YouTube。
一次性完成整本书的旁白:批量处理
功能可处理多个文本文件或一个长脚本,并在您进行其他操作的同时自动完成。您可以将所有内容合并到一个文件中,也可以将它们分开保存。一个错误的文本块不会影响其他部分,任何单个文本块都可以单独重做,而无需重新运行整个队列。
一次性修正发音:
发音错误的姓名、品牌、缩写和技术术语会被添加到发音列表中——例如,Nguyen 会变成“win”,SQL 会变成“es cue el”。规则适用于所有地方,因此角色名称只需在整本书中修正一次即可。
控制文本内的朗读方式。
输入[pause 800]即可获得时长为800的静音。使用[slow]、[loud]、[high]或[strong]将短语括起来。使用[spell]逐字母朗读缩写词。
隐私至上。
完全没有网络功能。该应用在进程级别拒绝自身的出站连接,因此您的文字和音频不会传输到任何地方。无需账户、无需登录、无需遥测数据。您输入的任何内容都不会用于训练任何程序。
坦诚相待。
它不会克隆真人录音中的声音,也无法做到这一点。它播放合成语音,并允许您调整和混合它们。每个导出的文件在其属性中都会包含一条注释,将其标记为AI生成的合成语音,并包含日期和使用的语音。
– 安装预先设置好的程序,全部已激活。
Turn any text into natural spoken audio – 59 voices, 9 languages, entirely offline on your PC. No account, no subscription, no upload. Export MP3 with subtitles.
59 VOICES IN 9 LANGUAGES, INCLUDED
Ships with 54 natural studio voices at 24 kHz – English (US and UK), Spanish, French, Italian, Portuguese, Hindi, Japanese and Mandarin Chinese – and also lists every speech voice Windows has installed. Nothing to download, nothing to sign up for. Everything runs on the processor; no graphics card needed.
Press Browse all voices to search by name or language, filter by female or male, star your favourites, and hear any voice instantly from its card.
GIVE EVERY CHARACTER THEIR OWN VOICE
The Cast screen turns a story into a performance. Paste prose with quoted dialogue and it separates narration from speech and works out who is talking, from phrases like “said Maren” or “Ellis shook her head”. Paste a script with NAME: lines and it reads that instead. Assign a voice to each character, press Record, and every line is spoken in its own voice and joined into one recording.
BLEND A VOICE NOBODY ELSE HAS
Mix any two studio voices of the same language with a slider. This combines them at the level the model uses to define what a voice is, so the result is a genuinely new voice, not a filtered copy.
WRITE AND NARRATE
Type or paste up to 200,000 characters. Long passages split automatically, with a progress bar and a live estimate of time remaining. Listen to the real waveform with a time ruler, click anywhere to skip, and export when it sounds right.
Your work saves itself as you type. Keep documents by name, import text files, and use Find and Replace on long manuscripts.
EXPORT MP3, WAV, FLAC OR OGG – WITH SUBTITLES
MP3 is about a fifth the size of WAV and plays on everything. FLAC is lossless. Tick a box and it also writes an .srt subtitle file, .vtt web captions, or a timecoded transcript beside the audio. Timings are measured from the audio actually produced, so captions line up exactly – drop them straight into Premiere, Resolve, CapCut or YouTube.
NARRATE A WHOLE BOOK IN ONE RUN
Batch takes a stack of text files, or one long script, and works through it while you do something else. Join everything into one file, or keep them separate. One bad block never abandons the rest, and any single block can be redone without re-running the queue.
FIX PRONUNCIATION ONCE
Names, brands, acronyms and technical terms that come out wrong go in a pronunciation list – Nguyen becomes “win”, SQL becomes “es cue el”. Rules apply everywhere, so a character’s name is fixed once for the entire book.
CONTROL DELIVERY INSIDE THE TEXT
Type[pause 800] for a silence of exactly that length. Wrap a phrase in[slow],[loud],[high] or[strong]. Use[spell] to read an acronym letter by letter.
PRIVACY IS THE POINT
No network features at all. The app refuses its own outbound connections at the process level, so your writing and your audio cannot be transmitted anywhere. No account, no sign-in, no telemetry. Nothing you write is used to train anything.
HONEST ABOUT WHAT IT IS
This does not clone a real person’s voice from a recording, and it cannot. It plays synthetic voices and lets you tune and blend them. Every exported file carries a note in its properties marking it as AI-generated synthetic speech, with the date and the voice used.
– Install pre-done setup, all activated.


评论0