Explore
Featured models
black-forest-labs / flux-1.1-pro
Faster, better FLUX Pro. Text-to-image model with excellent image quality, prompt adherence, and output diversity.
black-forest-labs / flux-schnell
The fastest image generation model tailored for local development and personal use
black-forest-labs / flux-dev
A 12 billion parameter rectified flow transformer capable of generating images from text descriptions
okaris / omni-zero-couples
Omni-Zero Couples: A diffusion pipeline for zero-shot stylized couples portrait creation.
levelsio / counter-strike
Take pics in the style of Counter-Strike 1.6's custom map fy_resort
meta / meta-llama-3.1-405b-instruct
Meta's flagship 405 billion parameter language model, fine-tuned for chat completions
I want to…
Generate images
Models that generate images from text prompts
Use a language model
Models that can understand and generate text
Caption images
Models that generate text from images
Edit images
Tools for manipulating images.
Restore images
Models that improve or restore images by deblurring, colorization, and removing noise
Upscale images
Upscaling models that create high-quality images from low-quality images
The FLUX.1 family of models
The FLUX.1 family of text-to-image models from Black Forest Labs
Get embeddings
Models that generate embeddings from inputs
Extract text from images
Optical character recognition (OCR) and text extraction
Transcribe speech
Models that convert speech to text
Chat with images
Ask language models about images
Use handy tools
Toolbelt-type models for videos and images.
Use a face to make images
Make realistic images of people instantly
Generate music
Models to generate and modify music
Generate videos
Models that create and edit videos
Fine-tune Flux
Create a fine-tuned Flux model using your own training images.
Generate speech
Convert text to speech
Make 3D stuff
Models that generate 3D objects, scenes, radiance fields, textures and multi-views.
Get structured data
Language models that support grammar-based decoding as well as jsonschema constraints.
Popular models
SDXL-Lightning by ByteDance: a fast text-to-image model that makes high-quality images in 4 steps
Fine-Tuned Vision Transformer (ViT) for NSFW Image Classification
A text-to-image generative AI model that creates beautiful images
Practical face restoration algorithm for *old photos* or *AI-generated faces*
Visual instruction tuning towards large language and vision models with GPT-4 level capabilities
Latest models
Phi-3-Mini-128K-Instruct is a 3.8 billion-parameter, lightweight, state-of-the-art open model trained using the Phi-3 datasets
This is wizard-vicuna-13b trained with a subset of the dataset - responses that contained alignment / moralizing were removed
Newest reranker model from BAAI (https://huggingface.co/BAAI/bge-reranker-v2-m3). FP16 inference enabled. Normalize param available
Generate a video that morphs between subjects, with an optional style
An efficient, intelligent, and truly open-source language model
Make stickers with AI. Generates graphics with transparent backgrounds.
yuan2.0-2b-mars是源2.0-2B模型的2024年3月版本,源2.0 是浪潮信息发布的新一代基础语言大模型。我们开源了全部的3个模型源2.0-102B,源2.0-51B和源2.0-2B。并且我们提供了预训练,微调,推理服务的相关脚本,以供研发人员做进一步的开发。源2.0是在源1.0的基础上,利用更多样的高质量预训练数据和指令微调数据集,令模型在语义、数学、推理、代码、知识等不同方面具备更强的理解能力。
Idefics2 is an open multimodal model that accepts arbitrary sequences of image and text inputs and produces text outputs
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
FlashFace: Human Image Personalization with High-fidelity Identity Preservation
text2img model trained on LAION HighRes and fine-tuned on internal datasets
snowflake-arctic-embed is a suite of text embedding models that focuses on creating high-quality retrieval models optimized for performance
input your name, and this model will print the most handsome man
Base version of Llama 3, a 70 billion parameter language model from Meta.
A 70 billion parameter language model from Meta, fine tuned for chat completions
An 8 billion parameter language model from Meta, fine tuned for chat completions
Base version of Llama 3, an 8 billion parameter language model from Meta.
a powerful and competitive model like Midjourney v6 and DALL-E 3 but Open and Decentralized
InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models
lightweight text-to-speech (TTS) model, trained on 10.5K hours of audio data
Midjourney v6 text-to-image quality model but Open and Decentralized
Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
A large, stereo MusicGen that acts as a useful tool for music producers
Nous Hermes 2 Mixtral 8x7B DPO is a Nous Research model trained over the Mixtral 8x7B MoE LLM
Use a subset of https://github.com/barun-saha/slide-deck-ai to create powerpoint slides from a json description - using python-pptx (https://github.com/scanny/python-pptx)
StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
Turn a face into 3D, emoji, pixel art, video game, claymation or toy