VoiceStudio, formerly OmniVoice Studio, is an AGPL licensed desktop application for local speech generation and processing. It combines voice cloning and design, video dubbing, dictation, transcription, multi voice stories, and audiobook production. The project presents itself as a self hosted ElevenLabs alternative and is in active beta. It ships for Apple Silicon macOS, Windows, and Linux, with Docker and remote compute options. Product Repository Latest release
The product is an integration layer over a broad model catalogue rather than one speech model. Its current documentation lists sixteen TTS engines and eleven ASR engines, including OmniVoice, CosyVoice 3, GPT SoVITS, VoxCPM2, WhisperX, Faster Whisper, and several platform specific options. The advertised 646 language catalogue is therefore not uniform coverage: actual language support, cloning, quality, hardware requirements, and commercial terms depend on the selected engine. The application is AGPL 3.0, while optional models retain their own licenses. Engine matrix License
The architecture is unusually useful for agent workflows: a Tauri and React desktop shell runs a FastAPI backend on localhost, with SQLite persistence, REST, server sent events, WebSockets, an MCP server, and OpenAI compatible speech and transcription endpoints. Existing OpenAI audio clients can point at http://localhost:3900/v1, which makes local substitution comparatively low friction. Core work stays local by default, analytics is opt in, and remote workers or remote transcription are explicit choices. Remote access still needs careful configuration and authentication. Architecture API quickstart API authentication
The project has meaningful momentum and a packaged v0.5.0 release, but its own documentation labels it active beta and the issue tracker still shows backend crashes, remote UI authentication problems, and dubbing defects. VoiceStudio is therefore more credible today as a local experimentation and high volume production toolkit than as a drop in replacement for a managed service with predictable latency, voice consistency, and support. Release notes Open issues