Voice & Audio AI covers 48 AI tools, with the top 10 averaging 75,747 community votes. 34 of the tools here are open source or have significant community traction. This page ranks them by all-time community votes, not by paid placement.
The current top three: yt-dlp, transformers, whisper. Each entry below shows the tool, its open-source stars or community size, and a short description from the project's own README. Click through to a full review for pricing, alternatives, and what it's actually good at.
Looking for something specific? Try the AI tool search engine β it indexes every tool on saas.pet and will surface what fits your workflow, not just what has the most votes.
A feature-rich command-line audio/video downloader
π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Robust Speech Recognition via Large-Scale Weak Supervision
Clone a voice in 5 seconds to generate arbitrary speech in real-time
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
πΈπ¬ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
π Text-Prompted Generative Audio Model
Easily train a good VC model with voice data <= 10 mins!
π€ Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
π€ The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
π¬ Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.
Open Source framework for voice and multimodal conversational AI
A framework for building realtime voice AI agents π€ποΈπΉ
Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (...
Code for the paper "Jukebox: A Generative Model for Music"
100+ AI Agent & RAG apps you can actually run β clone, customize, ship.
Silero Models: pre-trained text-to-speech models made embarrassingly simple
Clone any website with one command using AI coding agents
The open-source AI voice studio. Clone, dictate, create.
πΈSTT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
Build local voice agents with open-source models
Become a cracked AI/ML Research Engineer
Free voice coding for Claude Code, Codex and more
Voice-first local agent orchestration runtime for auditable DAG workflows.
$100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol
Agent behavior clone for browser using, targeting general GUI using and distributed trajectory collecting.
A TTS that fits in your CPU (and pocket)