Core modules
Speech-to-Text
Transcribe speech locally with configurable models and decoding settings.
Text-to-Speech
Generate speech from text with voice presets and offline synthesis.
VAD & Diarization
Detect speech segments and identify speaker turns for long recordings.
Enhancement & separation
Clean noisy audio and separate overlapping voices using local models.
Enhancement
Reduce background noise and improve speech clarity on-device.
Separation
Split mixed speakers into individual tracks for analysis.
Privacy-first
Voice Lab keeps processing local by design. You control which files and models are used.