Clone voices. Write scripts. Generate podcasts.
Record or upload voice samples → assign paragraphs to each person → generate multi-voice audio. 4 providers · 29 languages · Audio enhancement · Mobile + PC. Free: 2/day. Premium: 10 for $5.
🎤 Voices
📝 Script Editor
🎛️ Audio Settings
Assign voices to Left/Right/Center for stereo separation (podcast feel). Set in each voice's Settings panel.
Add [CHAPTER: Chapter Name] on its own line in the script to create chapter breaks in the audio.
🔧 Advanced
- 10 multi-voice audio generations per key
- 4 providers: ElevenLabs, Azure, Google Cloud, Hugging Face
- Voice cloning (record/upload → instant clone)
- Unlimited voices (Person A, B, C, D...)
- 29 languages with multilingual TTS
- Voice settings (stability, similarity, style, speed)
- Audio enhancement (normalize, crossfade, noise reduction)
- Background music mixing
- Script templates (podcast, interview, debate, audiobook)
- Export MP3 or WAV
- 1-PC device lock · Auto-expires after 10 uses
Free: 2 generations per day · Premium: 10 for $5 · 4 providers · 29 languages · BYOK
VoxKit — Multi-Voice Audio Studio
Clone voices. Write scripts. Generate podcasts, dialogues, and audiobooks with multiple AI voices. 4 providers · 29 languages · Voice cloning · Mobile + PC. Free: 2/day. Premium: 10 for $5.
What is VoxKit?
Voice Cloning Text-to-Speech Podcast Generator BYOK
VoxKit is a multi-voice audio studio that lets you create professional podcasts, interviews, audiobooks, and dialogues using AI-generated voices. Record or upload voice samples, clone them into AI voices, write a script with different speakers, and generate a complete audio production — all from your browser, on mobile or PC.
Unlike other voice tools that lock you into one provider, VoxKit supports 4 different AI voice providers — bring your own API key from ElevenLabs, Azure Speech, Google Cloud TTS, or Hugging Face, and use whichever provider offers the best quality for your needs. This BYOK (Bring Your Own Key) model means zero markup on AI costs — you pay the provider directly.
How It Works
- Add Voices: Create voice profiles for each person (Person A, Person B, Person C, etc.). Record from your microphone or upload an audio sample.
- Clone Voices: Click "Clone Voice" to send the sample to ElevenLabs — it creates an AI voice that sounds exactly like the recording. Get a Voice ID in seconds.
- Write Script: Add script lines, assign each line to a specific voice, and type what that person says. Reorder lines, use templates, or import from a file.
- Generate Audio: Click "Generate Audio" and VoxKit sends each line to the AI provider with the assigned voice. Audio clips are returned and played sequentially.
- Play & Download: Play all clips in sequence, or download individual clips as MP3 files.
4 Voice Cloning Providers (BYOK)
VoxKit supports 4 AI voice providers plus a free browser fallback. Add your API key from any provider — keys are stored in your browser only, never on our server.
🎙️ ElevenLabs (Best Quality)
Industry-leading voice cloning. Upload 10-60 seconds → instant clone. 29 languages, stability/similarity/style controls. Free tier: 10K chars/month. Best for professional podcasts.
☁️ Azure Speech (400+ Voices)
Microsoft's neural TTS with 400+ voices in 50+ languages. SSML support for emotion, emphasis, and prosody. Free tier: 500K chars/month. Best for enterprise use.
🌐 Google Cloud TTS (380+ Voices)
WaveNet voices that sound incredibly natural. 380+ voices, prosody control (rate, pitch, volume). Free tier: 1M chars/month. Best for multilingual content.
🤗 Hugging Face (Free)
Open-source TTS models: VITS, SpeechT5, Bark (with sound effects), MMS. Free without signup (rate-limited). Best for testing and budget projects.
Browser TTS (Free Fallback)
No API key? No problem. VoxKit includes a free Browser TTS mode that uses your device's built-in text-to-speech engine. Quality varies by browser/device, but it works on Chrome, Edge, Safari, and most mobile browsers. Perfect for testing the tool before getting a provider key.
Voice Cloning — How It Works
Voice cloning creates an AI voice that sounds exactly like a specific person. Here's the process:
- Record or Upload: Record 10-60 seconds of someone speaking (clear audio, no background noise). Or upload an existing audio file (MP3, WAV, M4A).
- Clone: Click "Clone Voice" — VoxKit sends the sample to ElevenLabs, which analyzes the voice characteristics (pitch, tone, accent, pacing, breathing patterns).
- Get Voice ID: ElevenLabs returns a unique Voice ID. This ID represents the cloned voice in all future generations.
- Generate: Type any text, select the cloned voice, and the AI speaks in that person's voice — indistinguishable from the original.
The cloned voice can say anything, in 29 languages, with adjustable stability (conservative vs. expressive), similarity (how close to original), and style (how stylized).
Script Editor Features
📝 Multi-Voice Script
Write dialogues with unlimited speakers. Assign each line to a different voice. Perfect for podcasts, interviews, debates, and audiobooks.
📋 Script Templates
Pre-built templates for Podcast, Interview, Debate, and Audiobook. Load with one click, then customize the text.
🔄 Reorder Lines
Move lines up or down with arrow buttons. Duplicate, delete, or insert lines anywhere in the script.
📊 Live Stats
See character count, line count, and estimated audio duration in real-time as you type.
💾 Auto-Save
Your script and voices are automatically saved in your browser. Close the tab, come back later, everything is still there.
🌍 29 Languages
Each voice can speak in a different language. English, Spanish, French, German, Arabic, Chinese, Hindi, and 22 more.
Per-Voice Settings
Each voice has independent settings that control how the AI speaks:
- Stability (0-100%): Low = more expressive and varied. High = more consistent and conservative.
- Similarity Boost (0-100%): How closely the AI voice matches the original sample. Higher = closer match.
- Style Exaggeration (0-100%): How much the AI exaggerates the voice's style. Higher = more stylized.
- Speed (0.5x - 2x): Playback speed. 1x = normal, 0.5x = slow, 2x = fast.
- Language: 29 languages for multilingual TTS.
Audio Enhancement
VoxKit applies enhancement algorithms to improve audio quality:
🔊 Volume Normalization
All voice clips are normalized to the same loudness — no more quiet voices and loud voices in the same podcast.
🎵 Crossfade
Smooth transitions between paragraphs. No abrupt cuts or awkward silences between speakers.
⏸️ Pause Control
Configurable pause duration between lines (0-5 seconds). Add dramatic pauses or keep it fast-paced.
🎶 Background Music
Upload background music and mix it with the generated audio. Control music volume separately.
Mobile + PC Compatibility
VoxKit works on both mobile and desktop:
- Microphone Recording: Uses
navigator.mediaDevices.getUserMedia()— works on mobile Chrome, Safari, and desktop browsers - Audio Playback: HTML5
<audio>elements — works on every device - Responsive Design: Mobile-first layout with touch-friendly controls
- Browser TTS: Uses device's built-in speech synthesis — works offline on most mobile devices
Use Cases
🎙️ Podcasts
Create multi-host podcasts without recording. Write the script, assign voices, generate audio. Perfect for testing formats before recording for real.
📚 Audiobooks
Convert written stories into audiobooks with different voices for each character. Narrator voice for descriptions, character voices for dialogue.
🎓 E-Learning
Create training videos with instructor voice and student voices. Multi-language courses without re-recording.
📺 Video Voiceovers
Generate voiceovers for videos in multiple languages using cloned voices. Same narrator across all languages.
💬 Dialogue Demos
Test dialogue scripts for films, games, or plays. Hear how conversations sound before hiring voice actors.
📢 Marketing
Create audio ads, product demos, and social media content with professional voices. No recording studio needed.
Auto-Update System (Deploy Once, Updates Forever)
VoxKit uses an auto-update architecture — you deploy the HTML block ONCE, and all future features are automatically loaded from the server. No need to re-paste the HTML block when new features are added.
How It Works
- The HTML block contains a bootstrap loader that runs on every page load.
- The loader checks the server for the latest voxkit-remote-features.js file via the
get_features_manifestandget_remote_featuresAPI endpoints. - If a newer version exists, the new JavaScript code is automatically injected and executed.
- The version is cached in
localStoragefor instant loading on subsequent visits. - The check runs on page load, then every 24 hours automatically.
What This Means For You
- Deploy once: Paste the HTML block into WordPress one time. That's it.
- Automatic updates: When the developer adds new features to
voxkit-remote-features.json the server, all deployed HTML blocks auto-load them within 24 hours (or on next page load). - No re-deployment: You never need to re-paste the HTML block again.
- Manual check: Click "Check for Updates" in the Advanced section to fetch updates immediately.
- What's New modal: Click "What's New" to see the changelog and current version.
- Version display: Footer shows current auto-update version with a green "Live" indicator.
- Update notifications: "VoxKit updated!" badge appears when new features are loaded.
Currently Auto-Loaded Features
The voxkit-remote-features.js file (v1.0.0) currently auto-loads 15 features:
- Auto-update notification badge — shows when new features load
- Version display in footer — current version + live indicator
- "What's New" button + modal — full changelog
- "Check for Updates" button — manual update check
- Live system status indicator — provider, voices, version
- Auto-backup every 5 minutes — last 10 backups saved
- Restore from auto-backup — recover unsaved work
- Random tip rotation — helpful hints on load
- Editing indicator — "● Editing..." when typing
- Crash recovery prompt — restore work after browser crash
- Performance monitor — logs slow requests to console
- Welcome back message — greets returning users
- Session time tracking — total time in VoxKit
- Feature usage tracking — anonymous, local only
- First-visit demo — shows welcome tip for new users
For Developers: Adding New Features
- Edit
voxkit-remote-features.json the server — add your new feature code (usewindow.VoxKitAPI,window.VoxKitState,window.VoxKitAddToolbarButton,window.VoxKitCreateModal). - Bump the version in
voxkit-features-manifest.json(e.g. "1.0.0" → "1.1.0"). - Add the feature name to the
featuresarray in the manifest. - Add a changelog entry in the manifest.
- Save both files — all deployed HTML blocks auto-update within 24 hours.
That's the entire workflow. No HTML editing, no re-deployment, no user action needed.
Auto-Updateable AI (Like Plagiarism Checker)
All AI features use a 4-key fallback system (Gemini → GLM → Groq2 → Groq1) — exactly like the Plagiarism Checker. No user API key needed for any AI feature. The system auto-routes through whichever server-side key is working, with 2 retries per provider.
Auto-Update Mechanisms
- License Auto-Refresh — Premium license status is refreshed every 5 minutes. Auto-detects when license is exhausted and switches to Free tier.
- Voice Catalog Auto-Update — Azure and Google voice catalogs are cached locally for 24 hours, then auto-refreshed from the server.
- Music/SFX Library Auto-Update — Background music and SFX libraries auto-refresh every 24 hours.
- Version Check — VoxKit checks for new versions and features on load, then daily. Shows update badge when new features are available.
- Provider Health Monitor — Provider endpoint reachability checked every 10 minutes. ElevenLabs health verified live.
- Auto-Save — Script and voices auto-save every 30 seconds with visual indicator showing last saved time.
AI-Powered Features (11 New — Server-side, no API key needed)
- AI Translate Script — Translate entire script to 13 languages (Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Russian, Dutch, English). Preserves speaker names. One click applies translation + updates voice language setting.
- AI Show Notes Generator — Creates podcast show notes from your script: 2-3 sentence summary, key points bulleted, timestamps with topics, tags, and call-to-action. Copy to clipboard for podcast hosts.
- AI Social Media Posts — Generate platform-optimized posts for Twitter/X, LinkedIn, Facebook, and Instagram. Includes hashtags. One-click copy per post.
- AI Episode Titles — Generate 5 SEO-friendly, catchy episode titles from your script content. Mix of styles: question, list, bold claim, mystery, how-to. Click to set as episode title.
- AI Voice Casting — AI acts as casting director, recommending which voice should say each line based on tone, emotion, and content. Includes reasoning and suggested emotion per line. One click applies all recommendations.
- AI Long Text → Script — Paste any long-form text (article, blog post, transcript up to 8000 chars) and AI converts it into a multi-voice conversational script. Choose 2-4 speakers and 6-16 lines.
- AI Chapter Auto-Tagger — AI analyzes your script and identifies topic shifts where chapter breaks should occur. Inserts [CHAPTER: Title] markers automatically.
- AI Pronunciation Suggestions — AI identifies difficult-to-pronounce words (technical terms, names, foreign words) in your script and suggests phonetic respellings. One click adds all to your Pronunciation Dictionary.
- AI Emotion Matching — AI recommends the best emotion tag (happy, sad, angry, calm, whisper, excited, serious, neutral) for each line based on content. Includes intensity score. One click applies all.
- AI Smart Auto-Complete — "AI Next" button on each script line. AI suggests the next line of dialogue based on the conversation context.
- Voice Quality Scorer — When you upload or record a voice sample, AI automatically analyzes it and gives a quality score (0-100) with recommendations. Badge displayed on each voice card.
Additional Advanced Features (14 New)
- VU Meter During Recording — Real-time audio level meter (green→orange→red) shown during microphone recording. Helps you maintain consistent recording levels.
- Audio Spectrum Analyzer — Real-time frequency spectrum visualization during audio playback. 32-band FFT display.
- Live Caption Preview — When audio plays, the current line's text appears as a floating caption at the bottom of the screen. Great for accessibility.
- Episode Auto-Numbering — Remembers your last episode and season number. Auto-suggests the next episode (S1E1, S1E2, ...). One click sets it as current. Reset counter anytime.
- Cross-Episode Voice Memory — Save your voice configurations as named sets. Load them in future episodes to maintain consistent voices across your podcast series. Save up to 20 voice sets.
- Auto-Trim Silence — Automatically detects and removes silence at the beginning and end of voice recordings. One click per voice, or enable auto-trim in preferences.
- Auto-Normalize Volume — Peak-normalizes voice recordings to -3dB. Ensures consistent volume across all voices. Shows the gain applied in dB.
- Episode Trailer Generator — Automatically picks the 4 most spread-out clips from your generated audio and combines them into a 15-second trailer preview. Download as single WAV.
- Script Difficulty Analysis — Analyzes reading level of each line (easy/medium/hard) based on word count, sentence length, and long-word ratio. Helps identify lines that may be hard for TTS to pronounce clearly.
- Audio Watermarking — Adds a barely-audible (15kHz at -40dB) marker to the end of generated audio. Detectable but inaudible — useful for tracking unauthorized use.
- Recording Quality Presets — Choose recording quality: Studio (48kHz/256kbps), Podcast (44.1kHz/128kbps), Phone (16kHz/64kbps), Voice Memo (22kHz/96kbps).
- Persistent Preferences — Auto-save, auto-translate, auto-score-quality, auto-trim-silence, and watermark settings saved across sessions.
- Auto-Retry Failed Generations — When a line fails to generate, the system automatically retries with browser TTS as fallback.
- Smart Auto-Save Indicator — Floating "Saved [time]" indicator appears top-right whenever the script or voices are saved. Persists for 2 seconds.
Pro Features (40+ Advanced Tools)
All features below are unlocked in VoxKit. Click the corresponding button in the Advanced section of the tool to use them.
AI-Powered Writing & Analysis
- AI Script Writer — Generate a complete podcast, interview, debate, audiobook, news, comedy, educational, or storytelling script from a topic. Powered by GLM-4. Choose 2-4 speakers, 6-16 lines, 6 languages.
- Sentiment & Mood Analysis — AI analyzes each line's emotion (happy/sad/angry/calm/excited/serious/whisper/neutral), sentiment score (-1 to +1), and reading level (easy/medium/hard). One click auto-applies the suggested emotion per line.
- Spell Check — GLM-powered grammar and spelling check across your entire script. Lists errors with suggestions.
- Voice Mood Matching — AI suggests the best emotion tag for each line based on its content.
Audio Production
- Audio Effects Chain — EQ presets (Podcast, Audiobook, Broadcast, Warm, Bright, Phone), Reverb (Small Room, Hall, Stadium, Cave), Compression (Light/Medium/Heavy), LUFS Loudness Normalization (-23, -16, -14, -19 LUFS), Noise Gate threshold.
- Post-Processing — Adjust overall speed (0.5x-2x), pitch (-12 to +12 semitones), and volume (-12 to +12 dB) of generated audio.
- Multi-Track Export — Export each voice as a separate WAV stem. Perfect for mixing in DAWs like Audacity, Reaper, or Pro Tools.
- Single WAV Export — Combine all generated clips into one continuous WAV file with proper encoding.
- MP3 Bitrate Selection — 64 / 128 / 192 / 320 kbps for size vs. quality tradeoff.
- Real-time Waveform Visualization — Live audio waveform during microphone recording.
- Audio Trimming & Noise Reduction — Cut silence from beginning/end, apply noise gate filter.
Voice Management
- Voice Library Browser — Browse 400+ Azure voices and 380+ Google voices by name, language, gender, or style. Searchable catalog with one-click apply.
- SSML Editor — Edit Speech Synthesis Markup Language for fine control. Templates for <break>, <prosody>, <emphasis>, <phoneme>, <say-as>, <sub>.
- A/B Voice Comparison — Generate the same text with two different voices side-by-side. Play both to pick the best.
- Voice Avatars — Upload an image for each voice (auto-resized to 100x100). Visual identification in long projects.
- Voice Cloning History — Track all cloned voices with original Voice IDs, provider, and timestamp.
Publishing & Distribution
- Subtitle Generation — Generate SRT (SubRip) or VTT (WebVTT) subtitles synced to your script timing. Download as a file.
- RSS Feed Generator — Create a podcast RSS XML feed with iTunes tags, season/episode numbers, categories, cover art URL, audio enclosure.
- Episode Metadata — Title, description, author, season, episode, language, categories, cover art URL, audio URL, duration. Saved locally and reusable.
- Cover Art Generator — Generate 1400×1400 podcast cover with title, subtitle, color gradient (linear/solid/diagonal/radial), adjustable font size. Meets Apple Podcasts spec.
Script Management
- Project Templates — Save the current voices + script as a reusable template. Load or delete saved templates anytime.
- Script Import — Import from .txt, .srt, .vtt, or .json. Auto-detects "Name: text" format for plain text.
- Script Export — Export as TXT (with speaker labels), JSON (full project), or SRT (subtitle format).
- Find & Replace — Search across all script lines. Replace one occurrence or all at once.
- Multi-Select Lines — Checkbox-select multiple lines for bulk operations: delete, duplicate, change voice, bookmark.
- Bookmarks — Mark important lines with a star. Quickly find them later.
- Collaborative Notes — Add a per-line note for team collaboration. Notes are saved with the project.
- Line Color Coding — Each voice gets a unique color dot. Quick visual identification of who speaks where.
Music & Sound Effects
- Music Library — 5 built-in royalty-free background tracks (Ambient Calm, Upbeat Corporate, Tech Pulse, Soft Piano, Corporate Warm) from Pixabay.
- SFX Library — 8 sound effects (Applause, Laughter, Whoosh, Ding, Phone Ring, Door Close, Rain, Typewriter). One-click set as transition sound.
Visualization & Stats
- Speaker Timeline View — Visual timeline showing when each speaker talks, with color-coded segments and a time ruler.
- Statistics Dashboard — Total generations, total lines, total audio minutes, voice count, script line count, pron entries, breakdown by provider, recent generation history.
- Audio Preview Cache — Cache up to 20 generated audio clips locally. Re-generating the same line is instant.
- Generation History — Last 5 generations with timestamp, line count, duration, and provider.
UX & Accessibility
- Onboarding Tour — 5-step interactive walkthrough for first-time users. Auto-shows on first visit. Replayable anytime.
- Hotkeys Reference Card — Press ? anytime to see all keyboard shortcuts (Ctrl+Enter/S/D/F/B/Shift+N).
- PWA Installable — Install VoxKit as a standalone app on desktop or mobile. Works offline for browser TTS.
- Dark Mode — Toggle dark theme. Persists across sessions.
- Auto-Save — Script and voices auto-save every 30 seconds. Visual indicator shows when saving.
Pricing
- 2 audio generations per day
- Resets at midnight
- All 4 providers (BYOK)
- Browser TTS (no API key)
- Unlimited voices
- Script editor with templates
- 29 languages
- Voice cloning (ElevenLabs)
- Per-voice settings
- Audio playback
- No signup required
- 10 audio generations per key
- 1-PC device lock
- Auto-expires after 10 uses
- All 4 providers (BYOK)
- Voice cloning (ElevenLabs)
- Unlimited voices
- 29 languages
- Per-voice settings (stability, etc.)
- Audio enhancement (normalize, crossfade)
- Background music mixing
- Script templates
- MP3 download per line
- Sequential playback
Frequently Asked Questions
Do I need an API key to use VoxKit?
No. Select "Browser TTS" as your provider — it uses your device's built-in text-to-speech, no API key needed. For professional quality (voice cloning, natural voices), add an ElevenLabs API key (free tier: 10K characters/month).
How does voice cloning work?
Record or upload 10-60 seconds of someone speaking. Click "Clone Voice" — the sample is sent to ElevenLabs, which analyzes the voice and creates an AI clone. You get a Voice ID that represents this voice. Any text you type will be spoken in that person's voice.
Can I use multiple voices in one script?
Yes. Add as many voices as you need (Person A, B, C, D...). Each script line can be assigned to a different voice. The generated audio plays each line in the assigned voice, creating a natural conversation.
Does it work on mobile?
Yes. VoxKit is fully mobile-compatible. Microphone recording works on mobile Chrome and Safari. Audio playback uses standard HTML5 audio. The interface is responsive and touch-friendly.
Which provider should I use?
ElevenLabs: Best quality and voice cloning. Free tier available. Azure: 400+ neural voices, best for enterprise. Google Cloud: 380+ WaveNet voices, great for multilingual. Hugging Face: Free open-source models. Browser TTS: Free, no key needed, basic quality.
Is my API key safe?
Yes. Your API key is stored only in your browser's localStorage — it is never sent to our server in plain text. It goes through the PHP proxy only when generating audio, and the proxy forwards it directly to the provider.
Can I download the generated audio?
Yes. Each generated audio clip has a download button. Download individual lines as MP3 files, or play all clips sequentially in the browser.
Does the premium key expire?
The key expires after 10 audio generations — not after a time period. You can use all 10 in one day or spread them across months. Once the 10th generation is complete, the key is permanently exhausted. Each key works on 1 PC only.
Ready to Create Your Podcast?
Clone voices, write scripts, and generate multi-voice audio. Free: 2/day. Premium: 10 for $5.
Open VoxKit → Get Premium — $5 →