AI Models Guide
Understand the AI models available in Genvid and what you can control when generating images, videos, and audio.
Overview
Genvid integrates with multiple AI providers to offer a diverse selection of models for content generation. Each model has different strengths, costs, and capabilities. This guide helps you understand which model to choose for your creative needs and what parameters you can adjust.
Key Concepts
- Model: An AI system trained to generate specific types of content (images, videos, or audio)
- Render Type: The category of generation task (T2I for text-to-image, I2V for image-to-video, etc.)
- Aspect Ratio: The width-to-height proportion of your output (16:9, 9:16, 1:1, etc.)
- Resolution: The pixel dimensions of your output (720p, 1080p, 4K, etc.)
- Duration: For video models, how long the generated video will be
How Models Are Organized
Models in Genvid are organized by media type:
| Media Type | Description |
|---|---|
| Audio | Voice synthesis and text-to-speech |
| Image | Still images from text or image editing |
| Video | Moving content from various inputs |
Within each type, models are further categorized by their specific task (render type).
Quick Reference: All Models at a Glance
This table provides a quick overview of all 58 active models. Refer to detailed sections below for full information.
Audio Models (2 models)
| Model | Cost | Best For |
|---|---|---|
| ElevenLabs v2 | $0.003/request | General voice synthesis (default) |
| FAL ElevenLabs Dialogue v3 | $0.0005/second | Long-form dialogue |
Image Models - Text-to-Image (13 models)
| Model | Cost | Resolution Control | Best For |
|---|---|---|---|
| FLUX1.1 [dev] | $0.025/megapixel | Full | General purpose, good value |
| Flux Pro v1.1 Ultra | $0.06/image | Aspect ratio only | High quality images |
| Ideogram v2 | $0.05/image | Aspect ratio only | Text in images |
| Ideogram v2 Turbo | $0.02/image | Aspect ratio only | Fast text rendering |
| Ideogram v3 | $0.08/image | Aspect ratio only | Best text quality |
| Ideogram v3 Turbo | $0.04/image | Aspect ratio only | Fast quality text |
| Imagen 4 Preview | $0.04/image | Full | Photorealistic content |
| Imagen 4 Ultra | $0.08/image | Full | Premium photorealism |
| Imagen 4 Ultra [Fast] | $0.06/image | Full | Fast photorealism |
| Nano-Banana Pro T2I | $0.0375/megapixel | Full | Budget-friendly |
| Qwen2.5 VL Image | $0.003/image | Full | Most affordable |
| Seedream 4.0 | $0.0175/image | Full | Budget quality |
| Leonardo Origin | $0.035/image | Full | Artistic styles |
Image Models - Image-to-Image Edit (6 models)
| Model | Cost | Resolution Control | Best For |
|---|---|---|---|
| FLUX Kontext Pro | $0.04/image | None (fixed output) | Smart image editing |
| Nano-Banana Pro I2I | $0.0375/megapixel | Full | Budget editing |
| Nano-Banana I2I | $0.0375/megapixel | Full | Legacy editing |
| Qwen2.5 VL Image Edit Plus | $0.004/image | Full | Affordable editing |
| Seedream v4 Edit | $0.0175/image | Full | Budget editing |
| Ideogram Edit | $0.08/image | None | Text-focused editing |
Image Models - Special (2 models)
| Model | Cost | Use Case |
|---|---|---|
| FLUX Kontext Multi (Merge) | $0.06/image | Combine multiple images |
| Remove Background | $0.02/image | Background removal |
Video Models (35+ models)
| Category | Model Count | Cost Range |
|---|---|---|
| Image-to-Video (I2V) | 18 models | $0.002 - $0.50/second |
| Text-to-Video (T2V) | 10 models | $0.0024 - $0.50/second |
| Keyframe-to-Video (KF2V) | 5 models | $0.0032 - $0.27/second |
| Video-to-Video (V2V) | 1 model | $0.27/second |
| Reference-to-Video (R2V) | 1 model | $0.27/second |
| Upscale | 1 model | $0.02/second |
What You Can Control vs. What Is Fixed
Understanding what parameters you can adjust is crucial for getting the results you want.
Parameters You CAN Control
These settings are available in the generation interface:
| Parameter | Where Used | Description |
|---|---|---|
| Prompt/Description | All models | The text describing what you want to generate |
| Aspect Ratio | Most image and video models | Shape of output (16:9, 9:16, 1:1, etc.) |
| Resolution | Some image models | Output size (720p, 1080p, 2K, 4K) |
| Duration | Video models with ranges | Length of video in seconds |
| Model Selection | All generation types | Which AI model to use |
Parameters Fixed by the Model
These are determined by the model and cannot be changed:
| Parameter | Description |
|---|---|
| Output File Format | Models output specific formats (MP4 for video, PNG/JPG for images) |
| Frame Rate | Video models have fixed frame rates |
| Internal Processing | Some models process at fixed resolutions regardless of input |
| Color Depth | Bit depth and color space are model-determined |
| Audio Sample Rate | Audio models have fixed quality settings |
Audio Models
Audio models generate voice content from text, enabling realistic character dialogue.
ElevenLabs v2 (Default)
The default text-to-speech model, providing high-quality voice synthesis.
| Attribute | Value |
|---|---|
| Cost | $0.003 per request |
| Type | Text-to-Speech (TTS) |
| Status | Default audio model |
What you control:
- Voice selection (from your cast members)
- The text/dialogue to speak
- Pacing through punctuation
What is fixed:
- Audio sample rate and quality
- Output format
Best for: General voice synthesis needs, shorter dialogue segments.
FAL ElevenLabs Dialogue v3
A cost-effective option for longer dialogue, charged per second of generated audio.
| Attribute | Value |
|---|---|
| Cost | $0.0005 per second |
| Type | Text-to-Speech (TTS) |
What you control:
- Voice selection
- Dialogue text
- Natural pacing through script writing
What is fixed:
- Audio quality parameters
- Output format
Best for: Longer dialogue sequences where per-second pricing is more economical.
Image Models: Text-to-Image (T2I)
These models create images from text descriptions. Use them to generate keyframes, asset images, and visual concepts.
Understanding Size Configuration
Different models handle image dimensions in different ways:
| Strategy | How It Works | Models Using This |
|---|---|---|
| Aspect Ratio + Resolution | You choose aspect ratio AND resolution separately | FLUX1.1 [dev] |
| Aspect Ratio Only | You choose aspect ratio, model determines resolution | Flux Pro Ultra, Ideogram |
| Width/Height | You specify exact pixel dimensions | Imagen 4, Seedream, Leonardo |
FLUX1.1 [dev]
A versatile development model offering excellent quality at reasonable cost.
| Attribute | Value |
|---|---|
| Cost | $0.025 per megapixel |
| Size Control | Full (aspect ratio + resolution) |
Supported Aspect Ratios:
- 21:9 (Ultra-wide)
- 16:9 (Widescreen)
- 4:3 (Standard)
- 1:1 (Square)
- 3:4 (Portrait standard)
- 9:16 (Vertical/mobile)
- 9:21 (Ultra-tall)
What you control:
- Prompt description
- Aspect ratio
- Resolution/image size
Best for: General-purpose image generation with flexible size options.
Flux Pro v1.1 Ultra
Premium quality model for high-end image generation.
| Attribute | Value |
|---|---|
| Cost | $0.06 per image |
| Size Control | Aspect ratio only |
Supported Aspect Ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, 9:21
Special Feature: Supports "raw" mode for unprocessed aesthetic.
What you control:
- Prompt description
- Aspect ratio
- Raw mode (true/false)
What is fixed:
- Output resolution (model-determined based on aspect ratio)
Best for: Premium quality images where you need top-tier results.
Ideogram v2 / v2 Turbo
Excellent for images containing text, logos, or typography.
| Attribute | v2 | v2 Turbo |
|---|---|---|
| Cost | $0.05/image | $0.02/image |
| Speed | Standard | Fast |
Supported Aspect Ratios: 16:9, 4:3, 1:1, 3:4, 9:16, 10:16, 16:10
What you control:
- Prompt description (including text to render)
- Aspect ratio
What is fixed:
- Output resolution
Best for: Any image that needs legible text, signs, logos, or typography.
Ideogram v3 / v3 Turbo
The latest Ideogram models with expanded aspect ratio support and improved text rendering.
| Attribute | v3 | v3 Turbo |
|---|---|---|
| Cost | $0.08/image | $0.04/image |
| Speed | Standard | Fast |
Supported Aspect Ratios: Includes all v2 ratios plus 3:1 and 1:3 for panoramic/banner formats.
Best for: Highest quality text rendering in images, banner and panoramic formats.
Imagen 4 Preview / Ultra / Ultra [Fast]
Google's photorealistic image generation models.
| Attribute | Preview | Ultra | Ultra [Fast] |
|---|---|---|---|
| Cost | $0.04/image | $0.08/image | $0.06/image |
| Quality | Good | Best | Very Good |
| Speed | Standard | Slow | Fast |
Size Control: Full control with separate width and height parameters. The system provides presets for each aspect ratio and resolution combination.
What you control:
- Prompt description
- Width and height (via presets)
- Aspect ratio
Best for: Photorealistic images, realistic human subjects, natural environments.
Nano-Banana Pro T2I
Budget-friendly option with full size control.
| Attribute | Value |
|---|---|
| Cost | $0.0375 per megapixel |
| Size Control | Full (width/height) |
What you control:
- Prompt, width, height
Best for: Budget-conscious generation where you need size control.
Qwen2.5 VL Image
The most affordable text-to-image option.
| Attribute | Value |
|---|---|
| Cost | $0.003 per image |
| Size Control | Full (width/height) |
Best for: Rapid prototyping, concept exploration, high-volume generation.
Seedream 4.0
Balanced quality and cost for general use.
| Attribute | Value |
|---|---|
| Cost | $0.0175 per image |
| Size Control | Full (width/height) |
Best for: Everyday image generation with good quality-to-cost ratio.
Leonardo Origin
Known for artistic and stylized outputs.
| Attribute | Value |
|---|---|
| Cost | $0.035 per image |
| Size Control | Full (width/height) |
Best for: Artistic styles, illustrations, stylized content.
Image Models: Image-to-Image Edit (I2I)
These models modify existing images based on your instructions. Use them for edits, touch-ups, and transformations.
FLUX Kontext Pro
A smart image editing model that understands context and makes intelligent modifications.
| Attribute | Value |
|---|---|
| Cost | $0.04 per image |
| Size Control | None - fixed output size |
Supported Aspect Ratios: Yes, but only affects composition, not output resolution.
What you control:
- Edit prompt (what changes to make)
- Aspect ratio preference
What is fixed:
- Output resolution (~1344x786)
- Cannot output at input resolution
Best for: Quick conceptual edits where final resolution is not critical. Use for previewing changes before applying with a different tool.
Not recommended for: Final production images, high-resolution requirements.
Nano-Banana Pro I2I / Nano-Banana I2I
Budget editing options with full size control.
| Attribute | Pro | Legacy |
|---|---|---|
| Cost | $0.0375/megapixel | $0.0375/megapixel |
| Size Control | Full | Full |
What you control:
- Edit prompt
- Output width and height
Best for: Budget editing where you need to preserve or control resolution.
Qwen2.5 VL Image Edit Plus
Affordable image editing with good results.
| Attribute | Value |
|---|---|
| Cost | $0.004 per image |
| Size Control | Full (width/height) |
Best for: High-volume editing tasks, affordable touch-ups.
Seedream v4 Edit
Budget-friendly editing with full control.
| Attribute | Value |
|---|---|
| Cost | $0.0175 per image |
| Size Control | Full (width/height) |
Best for: General editing at reasonable cost.
Ideogram Edit
Specialized for editing images with text content.
| Attribute | Value |
|---|---|
| Cost | $0.08 per image |
| Size Control | None |
Best for: Editing images that contain text, modifying typography, fixing text in images.
Image Models: Special Purpose
FLUX Kontext Multi (Merge)
Combines multiple images into a single coherent output.
| Attribute | Value |
|---|---|
| Cost | $0.06 per image |
| Type | Image Merge |
What you control:
- Multiple input images
- Merge prompt describing how to combine them
Best for: Character turnarounds, composite scenes, combining elements from multiple sources.
Remove Background
Automatically removes image backgrounds.
| Attribute | Value |
|---|---|
| Cost | $0.02 per image |
| Type | Background Removal |
What you control:
- Input image only
Output:
- Transparent PNG with background removed
Best for: Extracting subjects from images, creating assets for compositing.
Video Models
Video models are categorized by their input type and generation approach.
Understanding Video Render Types
| Render Type | Full Name | Input | Output |
|---|---|---|---|
| I2V | Image-to-Video | Single image + prompt | Video starting from that image |
| T2V | Text-to-Video | Text prompt only | Video from scratch |
| KF2V | Keyframe-to-Video | Start + end images + prompt | Video interpolating between them |
| V2V | Video-to-Video | Existing video + prompt | Transformed video |
| R2V | Reference-to-Video | Reference images + prompt | Video using those references |
| UPSCALE | Video Upscale | Existing video | Higher resolution video |
Video Duration: What You Can Control
Video models handle duration differently:
| Duration Type | Description | Examples |
|---|---|---|
| Fixed | Model only produces one duration | Veo 3 (8 seconds only) |
| Range | Choose from valid options | Kling (5 or 10 seconds) |
| Flexible | Specify within a range | Veo 2 (5-8 seconds) |
Image-to-Video Models (I2V)
These models animate a single image into a video. The most common workflow is to create a keyframe and then generate a video from it.
Kling v2.1 Master I2V (Default)
The default and recommended I2V model, offering excellent quality and motion.
| Attribute | Value |
|---|---|
| Cost | $0.2688 per second |
| Duration Options | 5 or 10 seconds |
| Default Duration | 5 seconds |
What you control:
- First frame (keyframe image)
- Video prompt
- Duration (5s or 10s)
- Aspect ratio
Best for: High-quality character animation, professional production.
Kling Model Family (I2V)
| Model | Cost/Second | Notes |
|---|---|---|
| Kling v2.1 Master | $0.2688 | Default, highest quality |
| Kling v2.0 | $0.0672 | Good quality, more affordable |
| Kling v1.6 Pro | $0.1344 | Professional tier |
| Kling v1.5 Pro | $0.112 | Legacy professional |
All Kling I2V models support: 5 or 10 second durations, aspect ratio control.
Veo Models (I2V)
Google's video generation models.
| Model | Cost/Second | Duration |
|---|---|---|
| Veo 3 | $0.50 | Fixed 8 seconds |
| Veo 2 | $0.25 | 5-8 seconds |
Best for: Photorealistic motion, natural movement.
LTX Video Models (I2V)
Ultra-affordable video generation.
| Model | Cost/Second |
|---|---|
| LTX 0.9.7 | $0.0024 |
| LTX 0.9.5 | $0.0018 |
Best for: Prototyping, concept testing, high-volume generation.
Other I2V Models
| Model | Cost | Notes |
|---|---|---|
| Minimax Video 01 | $0.12/second | Good motion quality |
| Luma Ray 2 | $0.0032/frame | Per-frame pricing |
| Luma Ray Flash 2 | $0.00048/frame | Fast, affordable |
| Sora Turbo | $0.21/second | OpenAI model |
| Runway Gen-4 | $0.25/second | Latest Runway |
| Runway Gen-3 Alpha | $0.10/second | Previous Runway |
| Seedance 1.0 | $0.14/second | Good value |
| Wan 2.1 | $0.008/second | Budget option |
| Wan 2.1 1.3B | $0.002/second | Most affordable |
| Pika 2.2 | $0.20/second | Character motion |
Text-to-Video Models (T2V)
Generate videos from text descriptions without any input images.
When to Use T2V
- Abstract concepts without visual reference
- Establishing shots and backgrounds
- Exploring ideas before creating keyframes
- When no keyframe exists yet
Kling T2V Models
| Model | Cost/Second |
|---|---|
| Kling v2.1 Master | $0.2688 |
| Kling v2.0 | $0.0672 |
| Kling v1.6 Pro | $0.1344 |
Other T2V Models
| Model | Cost | Notes |
|---|---|---|
| Veo 3 | $0.50/second | Fixed 8s, premium quality |
| Veo 2 | $0.25/second | 5-8s range |
| LTX 0.9.7 | $0.0024/second | Ultra-affordable |
| Minimax Video 01 | $0.12/second | Balanced option |
| Luma Ray 2 | $0.0032/frame | Per-frame pricing |
| Sora Turbo | $0.21/second | OpenAI quality |
| Runway Gen-4 | $0.25/second | Creative results |
Keyframe-to-Video Models (KF2V)
Generate videos that smoothly transition between a starting image and ending image.
When to Use KF2V
- Cinematic transitions
- Controlled camera movements
- Character repositioning
- When you know exactly how a shot should start and end
KF2V Models
| Model | Cost/Second |
|---|---|
| Kling v2.1 Master (Default) | $0.2688 |
| Kling v2.0 | $0.0672 |
| Veo 2 | $0.25 |
| Minimax Video 01 | $0.12 |
| Luma Ray 2 | $0.0032/frame |
What you control:
- First frame (start image)
- Last frame (end image)
- Transition prompt
- Duration (where applicable)
Video-to-Video Models (V2V)
Transform existing videos with style changes or enhancements.
Kling v2.1 Master V2V
| Attribute | Value |
|---|---|
| Cost | $0.2688 per second |
What you control:
- Source video
- Transformation prompt
- Optional dialog sync
Best for: Style transfer, restyling existing videos, syncing to dialog.
Reference-to-Video Models (R2V)
Generate videos using reference images for character/scene consistency.
Kling v2.1 Master R2V
| Attribute | Value |
|---|---|
| Cost | $0.2688 per second |
What you control:
- 1-N reference images (cast, locations, props)
- Video prompt
- Duration
Best for: Maintaining character consistency across shots, using specific asset appearances.
Upscale Models
Enhance video resolution and quality.
Topaz Video Upscale
| Attribute | Value |
|---|---|
| Cost | $0.02 per second |
What you control:
- Source video
- Scale factor
- Target resolution
- Target FPS
What is fixed:
- Upscaling algorithm
- Output format
Best for: Final polish before delivery, enhancing resolution of selected videos.
Choosing the Right Model
For Images
| Scenario | Recommended Models |
|---|---|
| General purpose | FLUX1.1 [dev], Seedream 4.0 |
| Highest quality | Flux Pro Ultra, Imagen 4 Ultra |
| Text in images | Ideogram v3, Ideogram v2 |
| Fastest generation | Ideogram Turbo variants, Qwen2.5 |
| Budget-conscious | Qwen2.5 VL ($0.003), Seedream ($0.0175) |
| Photorealistic | Imagen 4 family |
| Artistic styles | Leonardo Origin |
For Video
| Scenario | Recommended Models |
|---|---|
| Best quality | Kling v2.1 Master, Veo 3 |
| Good value | Kling v2.0, Minimax Video 01 |
| Budget generation | LTX Video, Wan 2.1 |
| Fast prototyping | Luma Ray Flash 2, LTX |
| Character animation | Kling family, Pika |
| Photorealistic | Veo family, Sora |
For Audio
| Scenario | Recommended Model |
|---|---|
| Short dialogue | ElevenLabs v2 |
| Long dialogue | FAL ElevenLabs Dialogue v3 |
Cost Optimization Tips
Image Generation
- Use Qwen2.5 for prototyping ($0.003) then switch to premium for finals
- Ideogram Turbo vs Standard - Turbo is half the price with good quality
- Megapixel pricing - FLUX1.1 dev charges by megapixel; smaller images cost less
Video Generation
- Use LTX for concepts ($0.002/second) before committing to expensive models
- 5 seconds vs 10 seconds - Choose shorter durations when possible
- Upscale last - Only upscale your selected, final videos
- Per-frame vs per-second - Calculate which is cheaper for your video length
General Tips
- Generate multiple variations with budget models first
- Refine your prompt before using premium models
- Track credit usage in Budget Management
- Delete rejected content to declutter your workspace
Common Questions
Why can't I control the resolution on some models?
Different AI models have different architectures. Some are trained to output specific resolutions and cannot change them. Models like FLUX Kontext Pro produce fixed-size outputs regardless of input. Always check the model's "Size Control" specification before generating.
Why is my edited image smaller than my original?
If you used FLUX Kontext Pro for image editing, it outputs at approximately 1344x786 regardless of input size. For edits that preserve resolution, use Nano-Banana I2I, Qwen2.5 VL Edit, or Seedream v4 Edit.
Why can Veo 3 only make 8-second videos?
This is a fundamental characteristic of how Veo 3 was trained. The model optimizes for quality at a specific duration and does not support other lengths. If you need different durations, use Veo 2 or Kling models.
What's the difference between I2V and R2V?
| Aspect | I2V | R2V |
|---|---|---|
| Input | Single keyframe | 1-N reference images |
| Purpose | Animate specific frame | Maintain consistency |
| Output | Video starting from keyframe | Video influenced by references |
| Best for | Standard shot generation | Multi-shot consistency |
Why are some models more expensive?
Price generally correlates with:
- Output quality - Higher fidelity, better details
- Speed - Faster generation often costs more
- Features - More controllable parameters
- Training cost - Models trained on larger datasets
How do I know which models support my aspect ratio?
The interface filters models by compatibility. When you select a generation mode and aspect ratio, only compatible models appear in the dropdown. You can also refer to the detailed model information in this guide.
Can I use multiple models together?
Yes! A common workflow is:
- Generate keyframe with an image model
- Create video with an I2V model
- Upscale with the upscale model
- Add dialog with an audio model
Each step can use a different model optimized for that task.
Next Steps
- Video Editor - Generate videos using these models
- Image Editor - Create and edit images
- Keyframe Editor - Compose keyframes for video generation
- Subscriptions & Credits - Understand credit costs
- Budget Management - Track spending across models
