Image, Video & Design covers 96 AI tools, with the top 10 averaging 65,475 community votes. 58 of the tools here are open source or have significant community traction. This page ranks them by all-time community votes, not by paid placement.
The current top three: yt-dlp, transformers, stable-diffusion. Each entry below shows the tool, its open-source stars or community size, and a short description from the project's own README. Click through to a full review for pricing, alternatives, and what it's actually good at.
Looking for something specific? Try the AI tool search engine β it indexes every tool on saas.pet and will surface what fits your workflow, not just what has the most votes.
A feature-rich command-line audio/video downloader
π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
A latent text-to-image diffusion model
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
π Text-Prompted Generative Audio Model
Learn System Design concepts and prepare for interviews using free resources.
π€ Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
π¦ Repomix is a powerful tool that packs your entire repository into a single, AI-friendly file. Perfect for when you need to feed your codebase to Large Language Models (LLMs) or other AI tools like ...
A one stop repository for generative AI research updates, interview resources, notebooks and much more!
Generative Models by Stability AI
π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.
High-Resolution Image Synthesis with Latent Diffusion Models
A curated list of modern Generative Artificial Intelligence projects and services
PyTorch package for the discrete VAE used for DALLΒ·E.
Anti-AI-slop design skill for Claude Code, Cursor, and Codex.
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
Code for the paper "Jukebox: A Generative Model for Music"
A format specification for describing a visual identity to coding agents. DESIGN.md gives agents a persistent, structured understanding of a design system.
Extracted system prompts from Anthropic - Claude Fable 5, Opus 4.8, Claude Code, Claude Design. OpenAI - ChatGPT 5.5 Thinking, GPT 5.5 Instant, Codex. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. ...
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 Β΅s overhead at 5k RPS.
Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
SimpleX - the first messaging network operating without user identifiers of any kind - 100% private by design! iOS, Android and desktop apps π±!
Edit videos with coding agents
Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
An open source design system that's fully customizable and agent ready
Ο RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection β all without a single pixel of video.
A curated list of Generative AI tools, works, models, and references