
Lyria 3 Pro
Full-length text-to-music songs with vocals and lyrics.

Lyria 3
Generate 30-second songs with vocals from text or images.

Gemini 3.7 Flash
Fast multimodal LLM for coding, agents, and long-document analysis.

Kokoro 82M
Text-to-speech with 54 multilingual voices.

Whisper Large V3
Transcribe speech-to-text in 99 languages with timestamps.

Wan 3.0 Video
Generate 30-second 1080p video with native audio.

Wan 2.6 Image to Video Flash
Animate photos into 15-second 1080p video with native audio.

Grok Imagine Image 2
Text-to-image and image editing with crisp, legible text.

LTX 2.5 Pro
Generate 1080p video with native audio and multi-shot scenes.

LTX 2.5 Fast
Text-to-video and image-to-video with native audio, up to 4K.

Qwen Image 3.0
Generate and edit legible in-image text, up to 2K.

Seedream 5.0 Pro Layer Decomposition
Split any image into editable transparent PNG layers.

Bria Extract Object
Extract any named object into a transparent PNG cutout.

Seedance 2.5
Generate cinematic multi-shot AI videos up to 30 seconds with synchronized native audio from text, images, or references.

FLUX 3 Draft Enhance
Upscale AI video drafts to Full-HD with native audio.

FLUX 3 Extend Video
Extend clips into seamless video continuations with synchronized audio.

FLUX 3 Image to Video
Animate images into 20-second clips with synchronized native audio.

FLUX 3 Text to Video
Cinematic text-to-video with native lip-synced audio, up to 20s.

Qwen3.8 Max
Multimodal reasoning and agentic coding with 1M-token context.

Grok Imagine Video 1.5 Reference to Video
Character-consistent video from up to 7 reference images.

Grok Imagine Video 1.5 Image to Video
Animate a still image into 1080p video with synced audio.

Grok Imagine Video 1.5 Text to Video
Text-to-video clips up to 1080p with native synchronized audio.

Sonilo Video to Video
Add frame-synced AI music and sound effects to video.

Sonilo Text to Audio
Commercial-safe music and sound effects from text prompts.

Sonilo Video to Audio
Generate video-synced music and sound effects from footage.

MiniMax Hailuo H3 Reference to Video
Keep characters and products consistent in 2K reference-to-video.

MiniMax Hailuo H3 Image to Video
Animate a still image into 2K video up to 15s.

MiniMax Hailuo H3 Text to Video
Text-to-video: cinematic 2K clips with native audio.

Pruna P Image Ideogram
Sub-second text-to-image with legible in-image text.

Video Editor Agent
General-purpose AI media agent: describe the transformation in natural language and it runs ffmpeg in a sandbox (transcode, resize, extract frames, GIFs, speed changes and more).

Reve 2
Generate and edit 4K images with sharp in-image text.

VEED Lipsync v2
Dub talking-head videos with emotion-matched lip-sync.

Ideogram V4 Remix
Restyle any image into posters with legible in-image text.

MiniMax M3
Reason over 1M-token context for coding and agents.

Nemotron 3 Ultra
1M-token reasoning for coding agents and deep research.

GLM 5.2
1M-token open-weight LLM for long-horizon coding.

HeyGen Generate Look
Change avatar outfits and backgrounds while keeping the same face.

Timeline
Utility node: Timeline. declarative multi-track video compositor

Video Split
Utility node: Video Split. 1->N split; returns videos[] array

Ideogram V4 Fast
Generate posters and logos with accurate in-image text.

Seedream 5.0 Pro
Region-precise image editing with native multilingual text.

Higgsfield Soul 2.0
Generate fashion-editorial photorealistic photos from text or reference.

VEED Subtitles
Automatically transcribes and burns styled, translated subtitles into any video with 30 presets and a single API call.

VEED Video Background Removal
Remove any video's background with no green screen, or cleanly key chroma footage, using AI matting.
VEED Avatars
Generate UGC-style talking avatar videos from text or audio using 28 stock presenters with realistic lip-sync.

VEED Lipsync
Re-syncs the lips of any talking-head video to a new speech audio track for realistic dubbing and localization.

VEED Fabric 1.0
Animate any image into a realistic talking video, lip-synced to your audio or generated from a text script.

OpusClip - Clips From Video
Turn long videos into captioned vertical shorts.

Pruna P Video Replace
Swap on-screen video characters while preserving motion and audio.

Pruna P Video Animate
Transfer video motion and audio onto any still image.

Nano Banana 2 Lite
Generate and edit 1K images in about four seconds.

Gemini Omni Flash
Text-to-video and image-to-video with synchronized native audio.

Seed Audio 1.0
Generate full audio scenes: dialogue, music, effects, voice cloning.
Pruna P Video Avatar
Animate any portrait into a lip-synced talking avatar.

Pruna P Image Try-On
Dress photos in multiple garments with photorealistic virtual try-on.

Seedance 2.0 Mini
Fast text-to-video and image-to-video with synchronized audio.

HappyHorse 1.1
Generate cinematic video with synchronized native audio and multilingual lip-sync from text, an image, or reference images.

Luma Ray 3.2
Cinematic text-to-video and image-to-video clips up to 1080p.

Luma Uni-1 Max
Generate and edit images from plain-text instructions.

Luma Uni-1
Reasoning-first text-to-image and natural-language image editing.

Grok Text-to-Speech
Convert text to speech in 20 languages with five voices.

Grok Imagine Video 1.5 (Preview)
Image-to-video with native synchronized audio, up to 720p.

Grok Imagine Video
Text-to-video and image-to-video with native synchronized audio.

Ideogram 4.0
Generate 2K posters and logos with accurate text rendering.

Grok Imagine Image
Text-to-image generation and editing, up to 2K resolution.
HeyGen Avatar V — Create Avatar
Train a Digital Twin avatar from reference video.
HeyGen Avatar V
Studio-quality talking-avatar videos from text or audio.

NSFW Checker
Detect NSFW and other inappropriate content in images. Returns a boolean has_nsfw_concepts flag, an overall NSFW score (0-1), and the full label tree with confidences.

Pixverse Mimic
Transfer motion from reference videos onto still images.

Gemini 3.1 Flash TTS
Expressive, controllable TTS with 70+ language support.

Gemini Embedding 2
Natively multimodal embeddings — text, image, audio, video and PDF mapped into one vector space, with 8 task-specific modes.

Gemini Embedding 001
MTEB #1 text embeddings for RAG, search, and clustering.

Imagen 4 Fast
Fast photorealistic image generation for bulk and iteration.

Imagen 4 Ultra
Photorealistic images with native 2K resolution and precise text.

Gemini 2.5 Flash Lite
Fastest Gemini 2.5 model for high-volume text and vision tasks.

Gemini 3.1 Flash Lite
Ultra-fast, affordable LLM for high-volume AI pipelines.

Gemini 3 Flash
Frontier-class reasoning and multimodal AI at scale.

Gemini 3.1 Pro
Frontier reasoning across text, images, video, and code.

GPT 5.5
Frontier reasoning and coding with 1M-token context window.

Smart Banner Resizer
Recompose one image into multiple ad and banner sizes.

HappyHorse 1.0
Cinematic 1080p text-to-video with native audio and lip-sync.

Image Metadata
Read image metadata: dimensions, format, EXIF (with GPS decoded to decimal), ICC profile name, raw XMP. Returns JSON, not an image.

Image Mask
Build / refine binary masks from JSON-described shapes (rect/polygon/ellipse). Optional dilate/erode/feather/invert; multi-shape merge.

Image Transform Pipeline
Apply an ordered pipeline of resize / crop / rotate / flip in a single call. Replaces the four separate tools.

GPT Image 2
Generate photorealistic images with legible multilingual text and 2K output.

Claude Opus 4.7
Anthropic's most capable AI model excelling at agentic coding, complex reasoning, and high-resolution vision with a 1M-token context window.

Seedance 2.0 Fast
Professional-grade video creation model with native audio, similar to SeeDance 2.0 but faster and cheaper.

Seedance 2.0
Cinematic AI videos with native audio and multi-shot narratives.

Wan 2.7 Video Editing
Edit existing videos precisely using natural language text instructions.

Wan 2.7 Reference to Video
Character-consistent multi-subject videos from reference images.

Wan 2.7 Image to Video
Animate any image into cinematic 1080P video with audio.

Wan 2.7 Text to Video
1080P cinematic videos with audio sync and multi-shot control.

Wan 2.7 Image Generation Pro
4K images with chain-of-thought reasoning and multilingual text.

Wan 2.7 Image Generation
2K image generation with precise multilingual text rendering.

Pixverse V6
15-second AI videos with native audio and cinematic controls.

Qwen Flash
Fastest low-cost LLM with 1M context for high-volume tasks.

Qwen Plus
Mid-tier 1M context LLM for summarization and content tasks.

QVQ Max
Chain-of-thought visual reasoning for math, charts, and diagrams.

Qwen 3 VL Flash
Fast, affordable vision-language model with 262K context OCR.

Qwen 3 VL Plus
Powerful visual QA and document analysis from images.

Qwen 3 Coder Flash
High-volume code generation with 1M token context window.

Qwen 3 Coder Plus
Generates, debugs, and refactors entire codebases efficiently.

QwQ Plus
Deep chain-of-thought reasoning for math, code, and logic.

Qwen 3 Max
1T-parameter LLM with hybrid reasoning and 262K context.

Qwen 3.5 Plus
Multimodal 1M context AI for image, video, and text.

Qwen 3.5 Flash
Fast multimodal AI processing text, images, and video affordably.

OpenAI o3 Mini
Cost-efficient reasoning model for coding, math, and science.

OpenAI o3
Frontier reasoning model for complex coding, math, and science.

GPT 5.4 Nano
Flagship-class AI for classification and extraction tasks.

GPT 5.4 Mini
Fastest efficient model for coding and computer-use tasks.

GPT 5.4
Most powerful GPT for frontier reasoning and multimodal tasks.

HyperSwap Image Faceswap by FaceFusion Labs
High-quality face swapping built for real production workflows.

Kling O3 Image To Video
Images to cinematic videos with precise motion control.

Kling O3 Video To Video Edit
Text-based video editor — swap backgrounds, characters, restyle scenes.

Kling V3 Image 2 Image
Transform images into photorealistic, production-ready visuals.

Kling V3 Text to Image
Photorealistic, print-ready images from text prompts.

Kling O3 Video To Video Reference
Swap characters and restyle videos using reference images.

Kling O3 Text-to-Video
15-second cinematic AI videos with native audio.

HyperSwap: Video Faceswap by FaceFusion Labs
Realistic face swapping in videos from a single image.

Wan 2.2 Image to Video Flash
Convert a single image into a coherent dynamic video.

Sam Audio Large
Isolate any described sound from mixed audio tracks.

Seedream 5.0 Lite: Image-to-Image
Transform images intelligently with detailed text prompts.

Seedream 5.0 Lite: Text-to-Image
Fast, affordable instruction-following image generation.

Nano Banana 2
Fast photorealistic images — ideal for marketing and ads.

Kling Create Voice
Clone any voice from a single audio sample.

Kling 3.0 Pro Image-to-Video
Animated 1080p videos from images with dynamic motion.

Kling 3.0 Standard Image-to-Video
Controlled cinematic 1080p videos from starting images.

Kling 3.0 Pro Text-to-Video
Cinematic 1080p videos with realistic audio from text.

Kling 3.0 Standard Text-to-Video
Stunning 1080p cinematic videos from simple text prompts.

Segmind Faceswap v5
Ultra-fast face and head swapping in images.

Flux-2 Klein-4b
Sub-second photorealistic image generation and editing.

Flux-2 Klein-9b
Ultra-fast photorealistic image generation on consumer GPUs.

LTX-2-19B I2V
Synchronized 4K audio-video generation from images, fast.

LTX-2-19B T2V
Synchronized video and audio from text, multiple input types.

Kling O1 Reference Image 2 Video
Identity-preserving videos from static images with character reference.

Kling O1 Video 2 Video Reference
Video style transfer using reference character images.

Kling O1 Image 2 Video
Physics-driven animations from images for creative storytelling.

Kling O1 Video 2 Video Edit
Edit any video with precise natural language commands.

Kling O1
Text-to-video creation with precise AI-driven motion control.

Qwen Image 2512
Photorealistic image generation with precise text description following.
Kling V2 Pro Avatar
Talking avatar videos from image and audio, high quality.
Kling Avatar V2 Standard
Lifelike video avatars with precise lip synchronization.

Kling 2.6 Pro Motion Control
Transfer motion from videos to animate custom characters.

Kling 2.6 Standard Motion Control
Precise motion transfer from reference videos to characters.

Kling 2.6
Still images into immersive cinematic videos with synchronized audio.

Heygen Avatar IV
Single photo into a lifelike talking avatar video.

GPT 5.1
Precise code review and developer workflow assistant.

GPT 5.2
Advanced reasoning with multimodal input for precise tasks.

Gemini TTS 2.5 Flash
Fast, lifelike text-to-speech with expressive emotional tones.

Gemini TTS 2.5 Pro
Human-like speech synthesis with rich expressive emotional depth.

Seedance 1.5 Pro
Synchronized video and audio generation for dynamic storytelling.

LTX Retake Video
Precise segment-level video edits maintaining full scene continuity.

Bria Video Eraser
Remove unwanted objects from videos while preserving audio.

Video Concatenate
Merge videos with custom layouts, spacing, and audio.

Flux 2 Max
Photorealistic images with maximum consistency and fine detail.

GPT Image 1.5 Edit
Precise image editing via natural language instructions.

GPT Image 1.5
Stunning photorealistic images with exceptional instruction-following.

Wan Scail
Professional character animations from reference images.

Chatterbox Turbo TTS
Ultra-fast, human-quality TTS with emotional expression.

Wan 2.6 Image To Video
Transform images into high-quality videos with audio sync.

Sync.so React 1
Edit video actors' emotions with realistic re-expression.

Sam 3D Object
Single 2D image into detailed 3D object models.

Sam 3D Body
Reconstruct 3D human body meshes from a single photo.

Seedream 4.5
Photorealistic image generation with precise text understanding.

Z Image Turbo
Photorealistic images in under one second, bilingual text.

Sam3 Video
Real-time video segmentation and multi-object tracking.

Flux 2 Flex
Consistent-style photorealistic images using reference inputs.

Flux 2 Pro
High-quality photorealistic images with cross-output consistency.

Sam3 Image
Precise object segmentation and tracking in images.

Video Tryon V2
Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in high-quality, fully-preserved motion up to 50 seconds.

Gemini 3 Pro
Autonomous multimodal AI for complex reasoning and coding.

Nano Banana Pro
High-fidelity images with accurate multilingual text rendering.

Qwen Image Edit Plus Blend It
Product placement into backgrounds with precise lighting match.

Qwen Image Edit Plus Eigen Banana
Precise text-guided image transformation and creative editing.

Qwen Image Edit Plus Eraser
Remove unwanted objects while preserving realistic backgrounds.

Qwen Image Edit Plus Face To Portrait
Cropped face into full identity-preserving portrait photo.

Qwen Image Edit Plus Group Photo
Merge individual portraits into realistic group photos.

Qwen Image Edit Plus Multiple Angles
Transform image perspective with natural language prompts.

Qwen Image Edit Plus Next Scene
Create cinematic sequences with seamless visual continuity.

Qwen Image Edit Plus Product Photography
Transform white-background products into immersive lifestyle scenes.

Qwen Image Edit Plus Relight
Advanced image relighting using natural language prompts.

Qwen Image Edit Plus Remove Lighting
Remove artificial lighting effects and restore natural tones.

Qwen Image Edit Plus Texture Apply
Apply precise textures to images using natural language.

Qwen Image Edit Plus Texture Extract
Extract seamless, tileable textures from photographs.

Qwen Image Edit Plus Add People Lora
Generate realistic multi-character scenes with natural interactions.

Qwen Image Edit Plus Multi Lora
Multi-image editing with superior identity and style control.

Pruna P Image Edit
Multi-image editing with AI-guided precision and control.

Pruna P Image
p-image generates high-quality images from text prompts in seconds, optimizing for speed and fidelity.

TTS Elevenlabs With Timing
Emotionally expressive TTS with word-level timestamp output.

Elevenlabs Forced Alignment
Precise audio-text synchronization with word-level timestamps.

Elevenlabs Audio Isolation
Extract clear speech from noisy audio and video.

Elevenlabs Dialogue With Timing
Multi-speaker dialogue with expressive timestamps included.

Elevenlabs Voice Design
Generate unique synthetic voices without audio samples.

Elevenlabs Voice Cloning
Hyper-realistic voice cloning from short audio samples.

Elevenlabs Dialogue
Immersive, emotionally expressive multi-speaker audio dialogue.

Kimi K2 Instruct 0905
Deep contextual understanding and complex code generation.

Bria Fibo
Photorealistic images from structured prompts with brand control.

Bria Fibo Structured Prompt
Convert complex inputs into structured JSON prompts for generation.

ClarityAI Creative Upscaler
Creative image upscaling with fine detail enhancement.

ClarityAI Flux Upscaler
Transform low-resolution images into stunning high-quality visuals.

ClarityAI Crystal Upscaler
Upscale images up to 200x with enhanced detail and vibrancy.

LTX 2 Fast
Fast, high-quality text-to-video generation by Lightricks.

LTX 2 Pro
High-quality video generation with advanced motion control.

Hailuo 2.3 Fast
Professional-quality videos from text and images at speed.

Hailuo 2.3
Hyper-realistic videos from text with fluid character motion.

Seedance 1.0 Pro Fast
Cinematic videos from text and images at ultra speed.

Heygen Video Translate
Translate videos to multiple languages with natural lip-sync.

Multi Video Merge
Merge multiple videos into a single combined output.

Qwen Image Edit Plus
Multi-image editing with precise text-guided transformations.
Image resizer
Resize images to any dimension quickly and precisely.

Image Converter
Convert images between formats instantly.

Video Speed Change
Speed up or slow down any video precisely.

Start & End Frame Extractor
Extract first and last frames from any video.

Veo 3.1 Fast
Transforms static images into dynamic 1080p videos with synchronized audio and natural motion.

Veo 3.1
Static images into high-quality videos with synchronized audio.

InfiniteTalk
Full-body animation from images synchronized perfectly to audio.

Claude 4 Sonnet
Advanced coding and multi-step agentic reasoning model.

GPT 5 Nano
Ultra-fast LLM responses for real-time AI applications.

GPT 5 Mini
Rapid high-quality AI across text, images, and files.

Gemini 2.5 Flash
Multimodal AI with transparent reasoning, fast and affordable.

Gemini 2.5 PRO
Complex multimodal reasoning across diverse inputs and formats.

Pixverse 5 Extend
Seamlessly extend and continue AI-generated videos.

Pixverse 5 Transition
Seamless AI-generated video transitions between scenes.

Pixverse 5 Video
Cinematic videos from text and images with photorealism.

GPT Image 1 Edit Mini
Affordable text-driven image generation and editing.

GPT Image 1 Mini
High-quality image generation from text, fast and affordable.
Kling V1 Pro AI Avatar
Dynamic AI avatars with synchronized speech from image.
Kling V1 Standard AI Avatar
Lifelike AI avatars with precise lip-sync for presentations.

Sora 2 Pro
Cinematic-quality videos from text with temporal consistency.

Video Watermark Remover
Remove watermarks from any video instantly with AI.

Wan Animate
Animate characters and replace video subjects seamlessly.

Sora 2
Stunning dynamic videos from detailed text descriptions.

Sam V2.1 Hiera Large
Meta's next-gen segmentation model for images and video.

Video Frame Interpolation
FILM synthesizes smooth, high-quality intermediate frames for fluid motion in videos with significant movement.

Claude 4.5 Sonnet
Claude Sonnet 4.5 empowers developers with advanced coding and reasoning for complex software solutions.

Wan 2.5 Image to Video
Wan2.5-Preview creates stunning, high-resolution videos with flawless audio synchronization from multiple inputs.

Wan 2.5 Text to Video
Wan2.5-Preview generates synchronized multimedia content, merging text, image, video, and audio seamlessly.

Kling 2.5 Turbo
Kling AI 2.5 Turbo generates fluid, cinematic videos from text and images, enhancing content creation and storytelling.

VeenaMax TTS
VeenaMAX transforms text into expressive, real-time speech across multiple Indian languages for seamless communication.

Higgsfield Speech 2 Video
Transform images and audio into dynamic, lip-synced videos for engaging digital content.

Seedream 4.0 (4k)
Seedream 4.0 generates high-resolution, professional-grade visuals with superior text rendering for impactful design.

Bria Increase Video Resolution
Transform your videos with AI-powered upscaling and seamless background removal for professional quality.

Sync.so Lipsync 2 Pro
Lipsync-2-Pro seamlessly synchronizes lips in videos for instant, high-quality multilingual content creation.

Video Tryon
Video Tryon is Segmind’s next-generation AI video model for instant virtual try-on, allowing users to visualize any outfit on any person in high-quality, fully-preserved motion up to 50 seconds.

Bria Prompt Enhancer
Bria AI generates high-quality, commercially safe images tailored to diverse creative needs.

Bria Mask Generator
Bria AI Get Masks automatically generates accurate object masks for advanced image editing and enhancement.

Higgsfield Text 2 Image Soul
SOUL AI transforms text into stunning, customizable visuals with unparalleled style control and precision.

Higgsfield Image 2 Video
Transform static images into dynamic, motion-rich videos with unparalleled control and creative depth.

Nano Banana
Gemini Image Editor preserves authentic subject identity while enabling seamless image editing and manipulation.

Qwen Image Edit Fast
Qwen-Image-Edit enables precise bilingual image editing for seamless localization and professional content creation.

Qwen Image Fast
Qwen-Image expertly generates stunning images with complex text integration, especially for Chinese typography.

Ideogram Character
Achieve perfect character consistency across multiple generations from a single reference image.

Runway Gen4 Aleph
Runway Aleph revolutionizes video editing with intelligent automation for seamless object and environment manipulation.

Lifestyle Product Shot by Image
Transforms ordinary product images into stunning, marketing-ready visuals for eCommerce success.

Bria Lifestyle Product Shot by Text
Transform isolated product images into dynamic lifestyle scenes with AI-driven contextual realism.

Bria Product Shadow
Bria Product Shadow enhances product images with realistic shadows for professional eCommerce presentations.

Bria Product Packshot
Transform product photos into professional, market-ready images with intelligent enhancements and background removal.

Bria Product Cutout
Automates precise product cutouts and background removal for professional eCommerce imagery at scale.

Bria Increase Resolution
Seamlessly upscale and manipulate images while preserving the highest fidelity and safety standards.

Bria Enhance Image
Bria AI creates precise, high-quality image enhancements and manipulations for diverse creative applications.

Bria Expand Image
Bria Expand enables precise image manipulation and enhancement with generative AI, trained exclusively on licensed data for safe, risk-free commercial use.

Bria Blur Background
Bria AI Image Editing API v2 enables precise and context-aware image manipulation for stunning visual outcomes.

Bria Erase Foreground
Seamlessly removes foreground subjects and regenerates backgrounds for flawless image editing.

Bria Generate Background
Transform images through advanced background editing and generative content creation for diverse applications.

Bria RMBG 2.0
Effortlessly extract backgrounds with unmatched precision, powered by models trained exclusively on licensed data for safe and risk-free commercial use. Unlike traditional binary masking, Bria RMBG 2.0 delivers non-binary masks with 256 levels of transparency, ensuring seamless edges and natural blending for diverse creative workflows.

Bria Generative Fill
Bria AI enables precise generative image editing for seamless creative enhancements and transformations.

Bria Eraser
AI object removal with seamless context-aware inpainting.

Qwen Image Edit
Transform images effortlessly through semantic context and pixel-perfect appearance changes.

Bria 3.2 Text to Image
Bria 3.2 AI transforms natural language into stunning visuals for diverse creative applications — with Base, Fast, and HD modes to match your creative needs.

Bria Vector Graphics
Bria Vision enables high-quality text-to-image and text-to-vector graphic generation for versatile commercial use.

GPT 5
GPT-5 automates complex coding tasks with integrated tools for seamless software development and deployment.

Qwen Image
Qwen-Image revolutionizes image generation and editing with seamless multilingual text integration and photorealistic detail.

Hunyuan3d-2.1
Transform 2D images into photorealistic, high-fidelity 3D assets effortlessly.

Flux Krea Dev
FLUX.1 Krea generates stunning, photorealistic images with fine-tuned aesthetic control for diverse creative applications.

Wan 2.2 Text to Video Fast
Wan2.2 transforms text and images into high-quality video clips with cinematic flair.

Wan 2.2 Image to Video Fast
Transforms simple text prompts into breathtaking cinematic-quality videos in minutes.

Hailuo 02 Fast
Transform any static image into a captivating, high-quality video clip effortlessly.

Flux Dev Finetuned
Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions

Segmind SegFit v1.3
SegFit v1.3 enables hyper-realistic virtual try-ons, enhancing online fashion retail experiences without physical photoshoots.

Multi Image Kontext Pro
Transform text into stunning, professional-grade images with precise editing capabilities.

Infinite You
InfiniteYou generates high-fidelity portraits preserving identity while aligning with creative text prompts.

Faceswap V3 Multifaceswap
Faceswap V3 Multifaceswap enables realistic face swapping in images, preserving lighting and expressions for professional results.

Veo 3 Fast
Veo 3 Fast rapidly creates high-quality, 8-second videos with synchronized audio for diverse content needs.

Minimax Hailou 2
Generate breathtaking 1080P cinematic videos from text or images with ultra-realistic motion and physics.

O4 Mini
OpenAI o4-mini enhances decision-making by processing text and images with advanced reasoning capabilities.

Seedance 1.0 Pro
Seedance Pro transforms text and images into engaging 720p dynamic videos with cinematic storytelling.

Luma Modify Video
Transform videos seamlessly with high-fidelity generative edits while preserving original actor performances.

Veena TTS
Veena transforms text into high-fidelity, expressive speech in Hindi and English for real-time applications.

Pixverse Lipsync
PixVerse Lipsync expertly synchronizes lip movements to audio for flawless video content creation.

Chatterbox TTS
Chatterbox transforms text into rich, natural speech with adjustable emotional expressiveness for diverse applications.

FLUX.1 Kontext [dev]
FLUX.1 Kontext [dev] creates coherent and editable images by integrating text and visual cues for iterative design.

Segmind SegFit v1.2
SegFit v1.2 creates hyper-realistic virtual try-on images, transforming fashion retail engagement and conversion rates.

Multi Image Kontext Max
FLUX.1 Kontext [max] creates stunning, photorealistic images from text prompts and input images seamlessly.

Google Veo 3
Veo 3 revolutionizes video creation with advanced text-to-video generation and realistic audio synthesis for cinematic content.

Kling 2.1 AI Video Generator
Kling 2.1 offers hyper-realistic video generation with improved motion, sharper 1080p visuals, and instant restyling capabilities. Its cost-effective pricing and faster rendering times make it a game-changer for creators seeking cinematic-quality AI videos.

Flux Kontext Max
FLUX.1 Kontext [max] transforms textual descriptions into stunning, high-fidelity images with seamless typography integration.

Flux Kontext Pro
FLUX.1 Kontext Pro transforms text prompts into high-quality, customized images with remarkable efficiency and precision.

Lyria 2
Lyria 2 by Google DeepMind is an advanced model that generates high-fidelity 48kHz stereo instrumental music from text prompts or lyrics, offering precise control over tempo, key, mood, and structure.
Runway Gen 4 Image
Runway's Gen-4 Image API enables precise, multimodal image generation for innovative creative and technical applications.

Segmind FaceSwap Comic v1
FaceSwap Comic v1 is an AI-powered face swapping model designed to blend real faces into illustrated or cartoon-style images while preserving the target’s artistic look. Ideal for personalized children’s storybooks and stylized content, it offers fine control over facial expression, realism, and stylistic adaptation.
Pixverse 4.5 Video
Pixverse 4.5 transforms static images and text into dynamic, engaging videos for captivating social media content.

Imagen 4
Imagen 4 is Google’s most advanced AI image generation model, creating detailed, photorealistic or abstract images from text prompts. It excels at fine details and accurate text, perfect for professional visuals like posters and presentations.

Caricature Style
Transform everyday photos into lively, whimsical caricature illustrations that highlight individual features with playful exaggeration.
Pixverse 4.5 Effects
PixVerse 4.5 transforms photos and text into stunning animated videos for impactful storytelling and marketing.

Segmind Relighting V2
Transform images with customizable, photorealistic lighting for unparalleled visual creativity and authenticity.

Ace Step Music
ACE-Step generates high-quality music rapidly, enhancing the creative process for developers and artists worldwide.

Chroma
Chroma is an open-source, 8.9B parameter text-to-image model (based on FLUX.1-schnell) designed for diverse and uncensored content generation, including anime, furry art, and photography.

Ideogram 3 Reframe
Ideogram 3.0's Reframe effortlessly adapts images to diverse formats, enhancing visual content creation for any platform.

Ideogram 3 Replace Background
Effortlessly replace backgrounds in images, enhancing visual storytelling and creativity with precision and speed.

Ideogram 3 Remix
Ideogram 3 Remix enables versatile image transformation, enhancing creativity through customizable design iterations.

Ideogram 3.0
Ideogram 3.0 revolutionizes content creation with photorealistic text-to-image generation and diverse aesthetic styles.

Nomos Image Upscaler 4k
This upscaling model is ideal for enhancing amateur to professional photos, excelling with subjects like cats, hair, and party scenes. It handles both small (as low as 300px) and large images well, delivering sharp, clear results even when significantly resized.

Skin Contrast Upscaler
Enhances skin detail in images while preserving background quality for professional photography and art.
Kling dizzydizzy
Kling DizzyDizzy transforms static content into dynamic, high-resolution videos, enhancing engagement and storytelling for creators.
Kling bloombloom
Kling AI transforms text and images into dynamic, high-quality video content with realistic motion and sound.
Kling 2
Kling 2.0 is an advanced AI video generator (5 and 10 seconds) that creates cinematic, dynamic videos from text or images with lifelike motion and precise prompt control at 720p resolutions.

Supir Photo-Realistic Image Restoration
SUPIR restores and enhances images to stunning, photo-realistic quality with advanced AI techniques.

Topaz Labs Image Upscale
Topaz Labs image upscale is an industry-leading AI photo upscaler designed to increase the resolution of photos while preserving and enhancing fine details such as sharpness, and textures.

Topaz Labs Video Upscale
Topaz Video AI upscales, enhances, denoises, stabilizes, and increases frame rates in video footage, transforming low-quality or standard-definition videos into ultra-sharp, high-resolution cinematic outputs up to 4k and 120FPS.

Dia (Text to Speech)
Dia by Nari Labs is an advanced open-weights TTS model that brings scripts to life with natural speech, emotions, and nonverbal cues. Easily control tone, voice, and delivery. Great alternative to ElevenLabs.

GPT Image 1 Edit
Edit and compose images using natural language with GPT Image 1 Edit, OpenAI’s powerful inpainting and multi-reference editing model. Perfect for marketing visuals, product updates, and creative asset generation.

GPT Image 1
Create high-quality AI-generated images from text prompts using OpenAI's GPT Image 1 model. Ideal for product design, content creation, and rapid visual prototyping at scale.
Video Slicer
Video Slicer

HiDream-I1 (Fast)
HiDream-I1 is a next-generation, open-source image generative foundation model designed for text-to-image synthesis, especially for rendering text.

Segmind SceneCraft v0.1
SceneCraft transforms plain or existing product images into visually rich, photorealistic scenes. Whether starting from a white background or enhancing existing settings, it works seamlessly across furniture, home decor, and consumer goods.

Google Translate
Translate effortlessly with the powerful Google Translation AI model.

Segmind SegFit v1.1
Segmind's Fashion and Immersive Try-on model. SegFIT offers effortless AI virtual try-on from just a product image. No models needed! Boost engagement & conversions with this flexible and fast try-on model.

Runway Gen 4 Turbo
Generate videos faster and cheaper with Runway Gen-4 Turbo! Create high-quality text, image, and combined video generation for rapid content creation.
Kling fuzzyfuzzy
Transform your photos instantly into adorable, plush-toy-like visuals with Kling fuzzyfuzzy effect.

Kling Heart Gesture
Express affection visually with Kling AI's heart gesture effect! Input two portraits and instantly create heartwarming videos featuring a dynamic heart gesture.
Kling Expansion
Unleash dynamic visuals with Kling Expansion! Effortlessly inflate and stretch elements for surreal and captivating effects.
Kling Hug
Create heartwarming videos instantly with Kling hug effect! Generate tender embracing animations.
Kling Squish
Transform your visuals with Kling AI squish effect! Easily compress and distort images/videos for playful, exaggerated effects.

Kling Kiss
Create a heartfelt video in seconds with Kling kiss effect! Input two portraits and instantly generate a kissing animation.
Llama 4 Scout Instruct Basic
Unlock powerful multimodal AI with Llama 4 Scout basic, a 17 billion active parameters model offering leading text & image understanding.

Llama 4 Maverick Instruct Basic
Llama 4 Maverick Instruct Basic is a 400B parameter powerhouse with 128 experts for unparalleled text and image understanding.
Warmth of Jesus
Experience the viral "Warmth of Jesus" effect on PixVerse! Transform your images into heartwarming videos of Jesus embracing people.
Muscle Surge
Instantly add muscle and strength to your videos with Pixverse Muscle Surge effect!
Pixverse Image to Video
Animate your photos effortlessly with Pixverse Image to Video AI! Upload, add motion prompts and styles.
Pixverse Text to Video
Effortlessly create captivating videos from text with Pixverse text to video AI! Customize style, duration, and more.

Google Veo 2 Image To Video
Discover Google Veo 2, an AI-powered image-to-video model with 4K resolution, realistic motion, and cinematic effects for creators and developers.

Segmind Relighting
Prompts to auto-magically relight your images.
Juggernaut Lightning Flux
Juggernaut Lightning Flux: Blazing fast (<300ms!) & powerful inference with enhanced visuals.
Juggernaut Pro Flux
Juggernaut Pro FLUX: Create stunningly realistic AI images with unprecedented detail and sharpness.

Segmind SegSwap v0.1
Swap Objects Instantly. The Segmind SegSwap v0.1 model enables dynamic and precise image editing by allowing users to remove, replace, or add objects and transfer patterns seamlessly within images.

3B Orpheus TTS (0.1)
Orpheus TTS is an open-source text-to-speech (TTS) system powered by the Llama 3B language model, designed for high-quality and customizable speech synthesis.
Video Loop
Effortlessly loop videos for engaging social media & storytelling with our Video Loop.

Segmind Faceswap v4
Segmind FaceSwap v4 enables fast and precise face or head swapping between images with customizable options for style, output format, and image quality. Designed for creators and designers, it ensures natural-looking results with reproducibility through seed control for consistent outputs.

Hunyuan-3d 2mv
Hunyuan3D-2mv is finetuned from Hunyuan3D-2 to support multiview controlled shape generation.
Luma Ray flash 2 (720p)
Generate stunning 720p videos from text with the Luma ray-flash-2-720p model. Faster & cheaper than Ray 2, offering realistic motion & detail.
Wan Video Effects
Transform your videos with diverse video effects. Start creating captivating videos today.
Google Veo 2
Create stunning, realistic videos with Veo 2, Google's state-of-the-art AI video generation model. Experience enhanced quality & cinematic control.

Minimax-image-01
Generate high-fidelity images from text with precise control & stunning quality with Minimax Image-01.

Elevenlabs Transcript
Transcribe audio to accurate text in 99 languages with speaker diarization and word-level timestamps.
Ideogram Describe
Ideogram describe can effortlessly generate detailed prompts from images. Perfect for refining creations or replicating styles.

Ideogram Reframe
Transform your images with Ideogram Reframe! Easily reframe square images to your chosen resolution.

Ideogram 2a Image to Image
Ideogram Image to Image: Transform your images with ease! Enhance, modify, or create entirely new visuals using advanced AI. Perfect for artists, designers, and anyone seeking creative inspiration.

Ideogram 2a Text To Image
Create captivating designs, realistic images & innovative logos with Ideogram 2a text-to-image.

Ideogram Turbo Text To Image
Create stunning images in seconds with Ideogram Turbo Text to Image. Fast AI model for quick ideation & text rendering.

Ideogram Turbo Image To Image
Transform images instantly with Ideogram Turbo Image to Image! Fast AI for quick edits & creative remixes.
Wan_2.1 Text to Video
Create visually impressive and feature varied, lifelike motion videos with Wan2.1 using text prompts.
Wan 2.1 720p image to video
Create high-quality 720p videos with excellent visual quality and a broad spectrum of motion from static images.
Minimax AI Director
Minimax video-01-director: Create high-quality videos with control camera movements precisely using text prompts.

Imagen 3
Imagen 3 is Google DeepMind's highest quality text-to-image model. Generates detailed images with enhanced lighting, diverse styles, and improved text rendering.

Qwen2 VL 72B Instruct
Qwen2-VL-72B-Instruct is a state-of-the-art multimodal model excelling in image and video understanding, with advanced capabilities for text-based interaction.

Hunyuan3D-2
Hunyuan3D 2.0 enables the creation of high-quality 3D models with intricate details. Produce assets that are visually appealing and suitable for professional use.
AI Face Swap (image and video)
AI Face Swap: Effortlessly replace faces online. Fine-tune swaps with advanced controls for age, gender, and resolution.

DeepSeek Chat
DeepSeek V3 combines cutting-edge AI technology with practical usability. Featuring a 671B parameter architecture, enhanced reasoning capabilities, and lightning-fast processing, it sets new standards for open-source AI models.

DeepSeek R1
DeepSeek-R1 is a cutting-edge AI reasoning model that combines reinforcement learning with supervised fine-tuning. Excels in complex problem-solving, mathematics, and coding tasks.
Kling AI 1.6 Image to Video
Kling AI 1.6 Image-to-Video is a powerful AI tool that transforms static images into captivating, animated videos. Create high-quality content effortlessly with Kling AI's advanced capabilities.
Kling AI 1.6 Text to Video
Kling AI 1.6 Text-to-Video is a cutting-edge AI tool that transforms text into stunning, lifelike videos. Create professional-quality content effortlessly with Kling AI's advanced capabilities.

Ideogram Image To Image
Ideogram Image to Image: Transform your images with ease! Enhance, modify, or create entirely new visuals using advanced AI. Perfect for artists, designers, and anyone seeking creative inspiration.

LTX Video
LTX-Video is the first DiT-based video generation model capable of generating high-quality videos in real-time. It produces 24 FPS videos at a 768x512 resolution faster than they can be watched.
MiniMax AI (Hailuo)
With Video-01 by MiniMax, create high-definition videos at 720p resolution and 25fps, featuring cinematic camera movement effects based on text descriptions.
Omini Control
OminiControl is an innovative framework that optimizes Diffusion Transformer models for versatile image generation tasks.

AI Product Photography
Elevate your product imagery with our AI-powered photography model. Create stunning, professional-quality photos that boost engagement and sales. Perfect for e-commerce and digital marketing.

Flux Fill Pro
Professional inpainting and outpainting model with state-of-the-art performance. Edit or extend images with natural, seamless results.

Flux Depth Pro
Professional depth-aware image generation. Edit images while preserving spatial relationships.

Flux Canny Pro
Professional edge-guided image generation. Control structure and composition using Canny edge detection

Flux Depth Dev
Open-weight depth-aware image generation. Edit images while preserving spatial relationships.

Flux Canny Dev
Open-weight edge-guided image generation. Control structure and composition using Canny edge detection.

Flux Fill Dev
Open-weight inpainting model for editing and extending images. Guidance-distilled from FLUX.1 Fill Dev

Flux Redux Schnell
Fast, efficient image variation model for rapid iteration and experimentation.

Flux Redux Dev
Open-weight image variation model. Create new versions while preserving key elements of your original.

Transparent Background Maker
Transform your images with Transparent Background Maker. Quickly remove backgrounds using AI technology, supporting PNG and JPG formats. Ideal for enhancing product photos and creating eye-catching graphics for social media and marketing.

Flux-1.1 Pro Ultra
Create stunning visuals effortlessly with Flux 1.1 Pro Ultra. Experience unparalleled image quality and speed.

Recraft V3
Recraft V3, the latest iteration of Recraft AI, offers a significant advancement in AI-driven image generation. This state-of-the-art model is designed to produce high-quality, detailed vector graphics, catering to the needs of designers, artists, and content creators alike.

Recraft V3 Svg
Recraft V3 SVG generates high-quality, customizable vector graphics with precision and ease. Perfect for logos, infographics, illustrations, and more.

Stable Diffusion 3.5 Turbo Text to Image
Stable Diffusion 3.5 Turbo offers exceptional customizability, efficient performance on consumer hardware, and diverse image outputs that accurately represent different skin tones and features, all while maintaining high-quality results and strong prompt adherence.

Stable Diffusion 3.5 Large Text to Image
Stable Diffusion 3.5 Large offers exceptional customizability, efficient performance on consumer hardware, and diverse image outputs that accurately represent different skin tones and features, all while maintaining high-quality results and strong prompt adherence.
Faceswap V3
Face Swap V3 is a cutting-edge tool that empowers you to seamlessly swap faces in images. With customizable features and advanced technology, you can achieve professional-quality results.
Video Audio Merge
Effortlessly merge audio and video with our intuitive Video Audio Merge model. Create stunning multimedia content with precise timing, fade effects, and customizable audio options. Perfect for content creators, filmmakers, and marketers.
Runway Gen Alpha Turbo Image to Video
Runway Gen-3 AlphaTurbo is a cutting-edge AI tool that transforms static images into dynamic videos with exceptional fidelity and motion
Kling AI Image to Video
Kling AI Image-to-Video is a powerful AI tool that transforms static images into captivating, animated videos. Create high-quality content effortlessly with Kling AI's advanced capabilities.
Kling AI Text to Video
Kling AI Text-to-Video is a cutting-edge AI tool that transforms text into stunning, lifelike videos. Create professional-quality content effortlessly with Kling AI's advanced capabilities.
Meta MusicGen Medium
MusicGen: Transform text into music with AI. Create unique, high-quality audio from simple descriptions. Experience the future of music generation with this innovative AI model.

Video Captioner
With Video Captioner create accurate, customizable subtitles for your videos effortlessly.
Face Detailer
Restore characters' faces to their original glory with Face Detailer. Enhance facial details, eliminate distortion, and upscale images for stunning results.

flux-pro-1.1
Flux Pro 1.1 is a cutting-edge image generation tool offering exceptional speed, quality, and customization. Ideal for digital artists, designers, and content creators.

MyShell Text To Speech
MyShell's Voice Cloning and Text to Speech - Transform your audio content with realistic, personalized voices. Experience high-quality, efficient, and cost-effective audio synthesis.
Video Stitch
Revolutionize your video editing with the Video Stitch Model. Seamlessly stitch clips, add captivating audio, and create professional-looking videos in minutes.

Simple Vector Flux Lora
Flux is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions

Ideogram Text To Image
Ideogram Text to Image: Turn your ideas into stunning visuals instantly with this powerful AI tool. Create captivating designs, realistic images, and more. Perfect for artists, designers, and anyone seeking creative inspiration.

Openvoice
OpenVoice is a versatile voice cloning model that supports multiple languages and offers precise tone replication, flexible style control, and zero-shot cross-lingual capabilities
Cog videoX Image To Video
CogVideoX image-to-video is a cutting-edge AI model that converts static images into dynamic, high-quality videos. Perfect for content creation, animation, and education, it offers high-resolution output, efficient inference, and versatile precision. Transform your images into engaging videos with CogVideoX

Expression Editor
Expression Editor uses reference images to accurately generate new images with desired expressions. Perfect for digital art, memes, and marketing.

Esrgan Video Upscaler
ESRGAN Video Upscaler: Experience sharper, clearer 4k videos with ESRGAN. This AI-powered video upscaler boosts resolution and reduces artifacts, making your video content look its best. Best Topaz alternative.

Consistent Character With Pose
Create images of a given character in different poses

Flux Pulid
Flux PuLID: Customize AI-generated images with your unique identity. Seamlessly integrate faces into text-to-image models for realistic and customizable results. High fidelity, tuning-free customization, and versatile editing options.

Flux Ipadapter
Flux IP Adapter is a cutting-edge AI model that lets you to create stunning, customized images. With its advanced style adaptation capabilities, Flux IP Adapter lets you seamlessly blend different artistic styles into your creations.

Flux Inpaint
Flux Inpainting is a powerful image editing tool designed to effortlessly edit and enhance your images. It's perfect for tasks like removing unwanted objects, restoring damaged photos, and creating artistic effects.
Flux Controlnets
Flux ControlNets is a collection of models that gives you precise control over image generation. By integrating ControlNet with Flux.1, these models enable you to create highly detailed and customized images with unprecedented accuracy.

Text Overlay
Elevate your visuals withText Overlay Model. Easily add customized text to any image, perfect for social media, marketing, and blogs. Enjoy precise positioning, advanced styling, and seamless integration.
Cog Video X 5B
CogVideo is a groundbreaking AI model that turns text into high-quality videos. Create realistic scenes, animations, and more with ease. Ideal for content creators, educators, and businesses.

Fast Flux.1 Schnell
Fast Flux.1 Schnell by Segmind is an optimized text-to-image model designed for developers needing faster image generation. It offers high efficiency without compromising quality. Perfect for startups and engineers seeking quick, resource-efficient AI models.

Flux Realism Lora with Upscale
Flux Realism Lora with upscale, developed by XLabs AI is a cutting-edge model designed to generate realistic images from textual descriptions.

Sam V2 Image
SAM v2, the next-gen segmentation model from Meta AI, revolutionizes computer vision. Building on SAM's success, it excels at accurately segmenting objects in images, offering robust and efficient solutions for various visual contexts.
Sam V2 Video
SAM v2 Video by Meta AI, allows promptable segmentation of objects in videos.

Flux.1 Image To Image
Flux Image-To-Image model by Black Forest Labs is an advanced deep learning tool designed for transforming images based on specific textual prompts.

Flux.1 Dev
Flux Dev is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions

Flux.1 Schnell
Flux Schnell is a state-of-the-art text-to-image generation model engineered for speed and efficiency.

Flux .1 Pro
Flux Pro is a state-of-the-art image generation with top of the line prompt following, visual quality, image detail and output diversity.

Text Embedding 3 Small
Text-embedding-3-small is a compact and efficient model developed for generating high-quality text embeddings. These embeddings are numerical representations of text data, enabling a variety of natural language processing (NLP) tasks such as semantic search, clustering, and text classification

Text Embedding 3 Large
Text-embedding-3-large is a robust language model by OpenAI designed for generating high-dimensional text embeddings for a wide range of natural language processing (NLP) tasks including semantic search, text clustering, and classification.

Realdream Pony V9
Real Dream Pony V9 is an advanced image generation model based on the Stable Diffusion XL (SDXL) architecture, excelling in photorealism.

AI Product Photo Editor
AI Product Photo Editor leverages advanced image-based ML techniques to generate high-quality product visuals using text prompts, product images, and background images.

RealDream Lightning
RealDream is a sophisticated image generation model utilizing SDXL Lightning architecture. It creates incredibly realistic images from textual prompts. With the ability to excellently generate human portraits from the user's descriptive text.

Llama 3.1 70b
Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8 and 70B sizes. The Llama 3 instruction tuned models are optimized for dialogue use cases and outperform many of the available open source chat models on common industry benchmarks.

Llama 3.1 8b
Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8 and 70B sizes. The Llama 3 instruction tuned models are optimized for dialogue use cases and outperform many of the available open source chat models on common industry benchmarks.

Stable Diffusion 3 Medium Image to Image
Stable Diffusion 3 Medium image-to-image is a cutting-edge AI tool that uses advanced image-to-image technology to transform one image into another.

SD3 Medium Tile Controlnet
SD3 Medium Tile ControlNet is a large generative image model designed for generating detailed images based on textual prompts and tile-based input images.
SD3 Medium Canny Controlnet
Stable Diffusion 3 (SD3) Medium Canny ControlNet uses Canny edge detection to provide fine-grained control over the generated outputs.
SD3 Medium Pose Controlnet
Stable Diffusion 3 (SD3) Pose ControlNet is a large generative image model tailored for generating images based on text prompts while using pose information as guidance.
Motion Control SVD
Motion Control SVD is an innovative deep learning framework that breathes life into static images. By intelligently managing both camera and object motion, it empowers creators to achieve precise animation effects.
Live Portrait video to video
Experience the magic of Live Portrait’s Video-to-Video Model! Transform your static images into dynamic videos seamlessly.

Image Superimpose V2
Superimpose V2 elevates image editing! Seamlessly layer images with background removal, precise positioning, and flexible resizing options. Explore 14 blending modes to create stunning effects
Video Faceswap
Video Faceswap is a powerful tool for creators, filmmakers, and meme enthusiasts. With this innovative technology, you can effortlessly replace faces in videos

Aura Flow
Largest completely open sourced flow-based generation model that is capable of text-to-image generation

Live Portrait
Live Portrait animates static images using a reference driving video through implicit key point based framework, bringing a portrait to life with realistic expressions and movements. It identifies key points on the face (think eyes, nose, mouth) and manipulates them to create expressions and movements.

ElevenLabs Dubbing
Instantly dubs audio and video into 29 languages while preserving each speaker's original voice.

Kolors
Kolors is a cutting-edge text-to-image model that bridges language and visual art. Transform your textual ideas into photorealistic images with semantic precision.

Image Superimpose
Superimpose model lets you to create captivating visuals by seamlessly overlaying one image on top of another. It streamlines your image layering process, allowing you to bring your creative vision to life effortlessly.

SDXL Img2Img
SDXL Img2Img is used for text-guided image-to-image translation. This model uses the weights from Stable Diffusion to generate new images from an input image using StableDiffusionImg2ImgPipeline from diffusers
SDXL Controlnet
SDXL ControlNet gives unprecedented control over text-to-image generation. SDXL ControlNet models Introduces the concept of conditioning inputs, which provide additional information to guide the image generation process

Story Diffusion
Story Diffusion turns your written narratives into stunning image sequences.

Elevenlabs Sound Generation
Eleven Labs' Sound Generation API provides a robust development tool for programmatically generating audio content using artificial intelligence. This API empowers developers and creators to integrate sound generation functionalities into their applications and workflows.

Elevenlabs Speech To Speech
Eleven Labs Speech-to-Speech offers AI-powered voice conversion for content creators, media professionals, and anyone seeking to modify or translate audio speech.

Elevenlabs Text To Speech
ElevenLabs TTS transforms text into captivating, human-like speech for diverse applications.

Omni Zero
Omni-Zero: A diffusion pipeline for zero-shot stylized portrait creation.

LLAVA 1.6 7B
LLaVa translates images into text descriptions & captions.

Tooncrafter
Create videos from illustrated input images

V Express
V-Express lets you create portrait videos from single images.

SadTalker
Audio-based Lip Synchronization for Talking Head Video

Hallo
Hallo lets you create portrait videos from single images.

Relighting
Prompts to auto-magically relight your images.

Automatic Mask Generator
Automatic Mask Generator is a powerful tool that automates the creation of precise masks for inpainting

Magic Eraser
LaMA Object Removal- AI Magic Eraser

Inpaint Mask Maker
Real-Time Open-Vocabulary Object Detection

Background Eraser
Background Eraser helps in flawless background removal with exceptional accuracy.

Clarity Upscaler
High resolution creative image Upscaler and Enhancer. A free Magnific alternative.

Consistent Character
Create images of a given character in different poses

IDM VTON
Best-in-class clothing virtual try on in the wild

Stable Diffusion 3 Medium Text to Image
Stable Diffusion is a type of latent diffusion model that can generate images from text. It was created by a team of researchers and engineers from CompVis, Stability AI, and LAION. Stable Diffusion v2 is a specific version of the model architecture. It utilizes a downsampling-factor 8 autoencoder with an 865M UNet and OpenCLIP ViT-H/14 text encoder for the diffusion model. When using the SD 2-v model, it produces 768x768 px images. It uses the penultimate text embeddings from a CLIP ViT-H/14 text encoder to condition the generation process.

Fooocus
Fooocus enables high-quality image generation effortlessly, combining the best of Stable Diffusion and Midjourney.

IPAdapter Style Transfer
Style & Composition Transfer with Stable Diffusion IP Adapter

Profile Photo Style Transfer
Turn any image of a face into artwork using Stable Diffusion Controlnet and IPAdapter

illusion-diffusion-hq
Monster Labs QrCode ControlNet on top of SD Realistic Vision v5.1

PuLID
Novel tuning-free ID customization method for text-to-image generation.

GPT 4 turbo
GPT-4 outperforms both previous large language models and as of 2023, most state-of-the-art systems (which often have benchmark-specific training or hand-engineering). On the MMLU benchmark, an English-language suite of multiple-choice questions covering 57 subjects, GPT-4 not only outperforms existing models by a considerable margin in English, but also demonstrates strong performance in other languages. Currently points to gpt-4-turbo-2024-04-09.

GPT 4o
GPT-4o (“o” for “omni”) is our most advanced model. It is multimodal (accepting text or image inputs and outputting text), and it has the same high intelligence as GPT-4 Turbo but is much more efficient—it generates text 2x faster and is 50% cheaper. Additionally, GPT-4o has the best vision and performance across non-English languages of any of our models. GPT-4o is available in the OpenAI API to paying customers.

GPT 4
GPT-4 outperforms both previous large language models and as of 2023, most state-of-the-art systems (which often have benchmark-specific training or hand-engineering). On the MMLU benchmark, an English-language suite of multiple-choice questions covering 57 subjects, GPT-4 not only outperforms existing models by a considerable margin in English, but also demonstrates strong performance in other languages.

Mixtral 8x22b
Mistral MoE 8x22B Instruct v0.1 model with Sparse Mixture of Experts. Fine tuned for instruction following.

face-to-many
Turn a face into 3D, emoji, pixel art, video game, claymation or toy

face-to-sticker
Turn a face into a sticker

Llama 3 8b
Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8 and 70B sizes. The Llama 3 instruction tuned models are optimized for dialogue use cases and outperform many of the available open source chat models on common industry benchmarks.

material-transfer
Transfer a material from an image to a subject

Faceswap V2
Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desired face. No dataset, no training

Insta Depth
InstantID aims to generate customized images with various poses or styles from only a single reference ID image while ensuring high fidelity

Background Removal V2
This model removes the background image from any image

NewReality Lightning SDXL
NewReality Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

DreamShaper Lightning SDXL
DreamShaper Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

Colossus Lightning SDXL
Colossus Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

Samaritan Lightning SDXL
Samaritan Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

Realism Lightning SDXL
Realism Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

ProtoVision Lightning SDXL
ProtoVision Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

NightVis Lightning SDXL
NightVis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

WildCard Lightning SDXL
WildCard Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

Dynavis Lightning SDXL
Dynavis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

Juggernaut Lightning SDXL
Juggernaut Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

Realvis Lightning SDXL
Realvis Lightning SDXL is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.

Try-On Diffusion
Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on

Fooocus Inpainting
Fooocus Inpainting is a powerful image generation model that allows you to selectively edit and enhance images.

Fooocus Outpainting
Fooocus Outpainting transforms ordinary images into extraordinary works of art by seamlessly expanding their boundaries.

InstantID
InstantID aims to generate customized images with various poses or styles from only a single reference ID image while ensuring high fidelity

Samaritan 3D XL
Samaritan 3D XL leverages the robust capabilities of the SDXL framework, ensuring high-quality, detailed 3D character renderings.

Stable Video Diffusion
Takes image as input and returns a video.

Segmind-Vega
The Segmind-Vega Model is a distilled version of the Stable Diffusion XL (SDXL), offering a remarkable 70% reduction in size and an impressive 100% speedup while retaining high-quality text-to-image generation capabilities.

Segmind-VegaRT
Segmind-VegaRT a distilled consistency adapter for Segmind-Vega that allows to reduce the number of inference steps to only between 2 - 8 steps.

IP-adapter Depth XL
IP Adapter Depth XL is built on the SDXL framework. This model integrates the IP Adapter and Depth preprocessor to offer unparalleled control and guidance in creating context-rich images.

SDXL Inpaint
This model is capable of generating photo-realistic images given any text input, with the extra capability of inpainting the pictures by using a mask

SSD Img2Img
This model uses SSD-1B to generate images by passing a text prompt and an initial image to condition the generation

SDXL-Openpose
This model leverages SDXL to generate the images with ControlNet conditioned on Human Pose Estimation.

SSD-Depth
This model leverages SSD-1B to generate the images with ControlNet conditioned on Depth Estimation

SSD-1B
SSD-1B efficiently generates high-quality, diverse images from text prompts in real-time.

Copax Timeless SDXL
The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.

Zavychroma SDXL
The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.

Realvis SDXL
The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.

Dreamshaper SDXL
The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software.

Word2img
Create beautifully designed words using Segmind’s word to image for your marketing purposes

Stable Diffusion Inpainting
Stable Diffusion Inpainting is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input, with the extra capability of inpainting the pictures by using a mask

Stable Diffusion img2img
This model uses diffusion-denoising mechanism as first proposed by SDEdit, Stable Diffusion is used for text-guided image-to-image translation. This model uses the weights from Stable Diffusion to generate new images from an input image using StableDiffusionImg2ImgPipeline from diffusers

Stable Diffusion XL 1.0
The SDXL model is the official upgrade to the v1.5 model. The model is released as open-source software

Reliberate
This model corresponds to the Stable Diffusion Reliberate checkpoint for detailed images at the cost of a super detailed prompt

Realistic Vision
This model corresponds to the Stable Diffusion Realistic Vision checkpoint for detailed images at the cost of a super detailed prompt

SD Outpainting
Stable Diffusion Outpainting can extend any image in any direction

Juggernaut Final
The most versatile photorealistic model that blends various models to achieve the amazing realistic images.

Epic Realism
This model corresponds to the Stable Diffusion Epic Realism checkpoint for detailed images at the cost of a super detailed prompt

Edge of Realism
This model corresponds to the Stable Diffusion Edge of Realism checkpoint for detailed images at the cost of a super detailed prompt

Cyber Realistic
The most versatile photorealistic model that blends various models to achieve the amazing realistic images.

ControlNet Soft Edge
This model corresponds to the ControlNet conditioned on Soft Edge.

ControlNet Scribble
This model corresponds to the ControlNet conditioned on Scribble images.

ControlNet Depth
This model corresponds to the ControlNet conditioned on Depth estimation.

ControlNet Canny
This model corresponds to the ControlNet conditioned on Canny edges.

Codeformer
CodeFormer is a robust face restoration algorithm for old photos or AI-generated faces.

Segment Anything Model
The Segment Anything Model (SAM) produces high quality object masks from input prompts such as points or boxes, and it can be used to generate masks for all objects in an image.

Faceswap
Take a picture/gif and replace the face in it with a face of your choice. You only need one image of the desired face. No dataset, no training

Background Removal
This model removes the background image from any image

ESRGAN
ERGAN is an Image Super-Resolution (upscaler) model that enhances images with stunning, high-quality upscaling while preserving the exact composition of the original source. It improves detail without altering the image content.

ControlNet Openpose
This model corresponds to the ControlNet conditioned on Human Pose Estimation.