Explore
Featured models
black-forest-labs / flux-1.1-pro
Faster, better FLUX Pro. Text-to-image model with excellent image quality, prompt adherence, and output diversity.
black-forest-labs / flux-schnell
The fastest image generation model tailored for local development and personal use
black-forest-labs / flux-dev
A 12 billion parameter rectified flow transformer capable of generating images from text descriptions
okaris / omni-zero-couples
Omni-Zero Couples: A diffusion pipeline for zero-shot stylized couples portrait creation.
levelsio / counter-strike
Take pics in the style of Counter-Strike 1.6's custom map fy_resort
meta / meta-llama-3.1-405b-instruct
Meta's flagship 405 billion parameter language model, fine-tuned for chat completions
I want to…
Generate images
Models that generate images from text prompts
Use a language model
Models that can understand and generate text
Caption images
Models that generate text from images
Edit images
Tools for manipulating images.
Restore images
Models that improve or restore images by deblurring, colorization, and removing noise
Upscale images
Upscaling models that create high-quality images from low-quality images
The FLUX.1 family of models
The FLUX.1 family of text-to-image models from Black Forest Labs
Get embeddings
Models that generate embeddings from inputs
Extract text from images
Optical character recognition (OCR) and text extraction
Transcribe speech
Models that convert speech to text
Chat with images
Ask language models about images
Use handy tools
Toolbelt-type models for videos and images.
Use a face to make images
Make realistic images of people instantly
Generate music
Models to generate and modify music
Generate videos
Models that create and edit videos
Generate speech
Convert text to speech
Fine-tune Flux
Create a fine-tuned Flux model using your own training images.
Make 3D stuff
Models that generate 3D objects, scenes, radiance fields, textures and multi-views.
Get structured data
Language models that support grammar-based decoding as well as jsonschema constraints.
Popular models
SDXL-Lightning by ByteDance: a fast text-to-image model that makes high-quality images in 4 steps
Fine-Tuned Vision Transformer (ViT) for NSFW Image Classification
Practical face restoration algorithm for *old photos* or *AI-generated faces*
Visual instruction tuning towards large language and vision models with GPT-4 level capabilities
A text-to-image generative AI model that creates beautiful images
Latest models
A text-to-image model with greatly improved performance in image quality, typography, complex prompt understanding, and resource-efficiency
MimicMotion: High-quality human motion video generation with pose-guided control
Make realistic images of real people instantly (w/ ip-adapter-plus-face_sdxl_vit-h)
PixArt Sigma 900M is a text-to-image generation model based on the PixArt Sigma architecture
NuminaMath is a series of language models that are trained to solve math problems using tool-integrated reasoning (TIR)
MARS5, a fully open-source (commercially usable) voice-cloning/TTS with break-through prosody and realism.
Cog wrapper for Ollama deepseek-coder-v2:236b
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Take audio from one video and add it to a second video. Good for adding back audio to liveportrait.
Change the fps of a video without changing its length or speed
Efficient Portrait Animation with Stitching and Retargeting Control
Kolors is a SOTA base image model for high quality image generation
The API automatically detects objects in an input image and returns their positional and mask information.
InternLM2.5 has open-sourced a 7 billion parameter base model and a chat model tailored for practical scenarios.
Phi-3-Mini-4K-Instruct is a 3.8B parameters, lightweight, state-of-the-art open model trained with the Phi-3 datasets
Qwen2 57 billion parameter language model from Alibaba Cloud, fine tuned for chat completions
SDXL fine-tune based on images of birds primarily from the British Library free archive
GLM-4V is a multimodal model released by Tsinghua University that is competitive with GPT-4o and establishes a new SOTA on several benchmarks, including OCR.
Convert speech in audio to text w/ `tiny`, `small`, `base`, and `large-v3` models
Dolphin-2.9 has a variety of instruction, conversational, and coding skills. It also has initial agentic abilities and supports function calling
Image generation, Inpaint Strength, loras custom_urls and enhancer.
Depth estimation with faster inference speed, fewer parameters, and higher depth accuracy.