Voice & Audio AI

TTS, speech, music, voice cloning

48 tools Β· All-time leaderboard

← All Categories

What is Voice & Audio AI?

Voice & Audio AI covers 48 AI tools, with the top 10 averaging 75,747 community votes. 34 of the tools here are open source or have significant community traction. This page ranks them by all-time community votes, not by paid placement.

The current top three: yt-dlp, transformers, whisper. Each entry below shows the tool, its open-source stars or community size, and a short description from the project's own README. Click through to a full review for pricing, alternatives, and what it's actually good at.

Looking for something specific? Try the AI tool search engine β€” it indexes every tool on saas.pet and will surface what fits your workflow, not just what has the most votes.

πŸ† Top 30 in Voice & Audio AI

πŸ₯‡
GitHub

yt-dlp

A feature-rich command-line audio/video downloader

β˜… 171,312 votes
πŸ’¬ 0
πŸ₯ˆ
GitHub

transformers

πŸ€— Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

β˜… 161,696 votes
πŸ’¬ 0
πŸ₯‰
GitHub

whisper

Robust Speech Recognition via Large-Scale Weak Supervision

β˜… 102,982 votes
πŸ’¬ 0
#4
GitHub

Real-Time-Voice-Cloning

Clone a voice in 5 seconds to generate arbitrary speech in real-time

β˜… 59,924 votes
πŸ’¬ 0
#5
GitHub

GPT-SoVITS

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

β˜… 58,804 votes
πŸ’¬ 0
#6
GitHub

LocalAI

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

β˜… 48,082 votes
πŸ’¬ 0
πŸ“… 4d
#7
GitHub

TTS

πŸΈπŸ’¬ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

β˜… 45,582 votes
πŸ’¬ 0
#8
GitHub

bark

πŸ”Š Text-Prompted Generative Audio Model

β˜… 39,161 votes
πŸ’¬ 0
#9
GitHub

Retrieval-based-Voice-Conversion-WebUI

Easily train a good VC model with voice data <= 10 mins!

β˜… 36,052 votes
πŸ’¬ 0
#10
GitHub

diffusers

πŸ€— Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.

β˜… 33,878 votes
πŸ’¬ 0
#11
GitHub

datasets

πŸ€— The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools

β˜… 21,792 votes
πŸ’¬ 0
πŸ“… 4d
#12
GitHub

rasa

πŸ’¬ Open source machine learning framework to automate text- and voice-based conversations: NLU, dialogue management, connect to Slack, Facebook, and more - Create chatbots and voice assistants

β˜… 21,218 votes
πŸ’¬ 0
#13
GitHub

NeMo

A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)

β˜… 17,400 votes
πŸ’¬ 0
#14
GitHub

OpenMontage

World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.

β˜… 15,353 votes
πŸ’¬ 143
πŸ“… 3d
#15
GitHub

pipecat

Open Source framework for voice and multimodal conversational AI

β˜… 12,878 votes
πŸ’¬ 0
#16
GitHub

agents

A framework for building realtime voice AI agents πŸ€–πŸŽ™οΈπŸ“Ή

β˜… 11,033 votes
πŸ’¬ 0
#17
GitHub

meetily

Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (...

β˜… 8,885 votes
πŸ’¬ 290
πŸ“… 6d
#18
GitHub

jukebox

Code for the paper "Jukebox: A Generative Model for Music"

β˜… 8,037 votes
πŸ’¬ 0
#19
GitHub

awesome-llm-apps

100+ AI Agent &amp; RAG apps you can actually run β€” clone, customize, ship.

β˜… 6,252 votes
πŸ’¬ 29
πŸ“… 7d
#20
GitHub

silero-models

Silero Models: pre-trained text-to-speech models made embarrassingly simple

β˜… 5,970 votes
πŸ’¬ 0
#21
GitHub

ai-website-cloner-template

Clone any website with one command using AI coding agents

β˜… 5,624 votes
πŸ’¬ 26
πŸ“… 3d
#22
GitHub

voicebox

The open-source AI voice studio. Clone, dictate, create.

β˜… 3,336 votes
πŸ’¬ 576
πŸ“… 2d
#23
GitHub

STT

🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.

β˜… 2,588 votes
πŸ’¬ 0
#24
GitHub

speech-to-speech

Build local voice agents with open-source models

β˜… 788 votes
πŸ’¬ 97
πŸ“… 2d
#25
GitHub

maths-cs-ai-compendium

Become a cracked AI/ML Research Engineer

β˜… 725 votes
πŸ’¬ 10
πŸ“… 3d
#26
Product Hunt

SKI

Free voice coding for Claude Code, Codex and more

β˜… 595 votes
πŸ’¬ 0
πŸ“… 1d
#27
GitHub

homerail

Voice-first local agent orchestration runtime for auditable DAG workflows.

β˜… 408 votes
πŸ’¬ 4
πŸ“… 2d
#28
Hacker News

$100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol

$100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol

β˜… 396 votes
πŸ’¬ 535
πŸ“… 4d
#29
GitHub

Browser-BC

Agent behavior clone for browser using, targeting general GUI using and distributed trajectory collecting.

β˜… 354 votes
πŸ’¬ 1
πŸ“… 1d
#30
GitHub

pocket-tts

A TTS that fits in your CPU (and pocket)

β˜… 235 votes
πŸ’¬ 55
πŸ“… 1d