Genvid
Documentation

AI Models Guide

Understand the AI models available in Genvid and what you can control when generating images, videos, and audio.

Overview

Genvid integrates with multiple AI providers to offer a diverse selection of models for content generation. Each model has different strengths, costs, and capabilities. This guide helps you understand which model to choose for your creative needs and what parameters you can adjust.

Key Concepts

  • Model: An AI system trained to generate specific types of content (images, videos, or audio)
  • Render Type: The category of generation task (T2I for text-to-image, I2V for image-to-video, etc.)
  • Aspect Ratio: The width-to-height proportion of your output (16:9, 9:16, 1:1, etc.)
  • Resolution: The pixel dimensions of your output (720p, 1080p, 4K, etc.)
  • Duration: For video models, how long the generated video will be

How Models Are Organized

Models in Genvid are organized by media type:

Media TypeDescription
AudioVoice synthesis and text-to-speech
ImageStill images from text or image editing
VideoMoving content from various inputs

Within each type, models are further categorized by their specific task (render type).


Quick Reference: All Models at a Glance

This table provides a quick overview of all 58 active models. Refer to detailed sections below for full information.

Audio Models (2 models)

ModelCostBest For
ElevenLabs v2$0.003/requestGeneral voice synthesis (default)
FAL ElevenLabs Dialogue v3$0.0005/secondLong-form dialogue

Image Models - Text-to-Image (13 models)

ModelCostResolution ControlBest For
FLUX1.1 [dev]$0.025/megapixelFullGeneral purpose, good value
Flux Pro v1.1 Ultra$0.06/imageAspect ratio onlyHigh quality images
Ideogram v2$0.05/imageAspect ratio onlyText in images
Ideogram v2 Turbo$0.02/imageAspect ratio onlyFast text rendering
Ideogram v3$0.08/imageAspect ratio onlyBest text quality
Ideogram v3 Turbo$0.04/imageAspect ratio onlyFast quality text
Imagen 4 Preview$0.04/imageFullPhotorealistic content
Imagen 4 Ultra$0.08/imageFullPremium photorealism
Imagen 4 Ultra [Fast]$0.06/imageFullFast photorealism
Nano-Banana Pro T2I$0.0375/megapixelFullBudget-friendly
Qwen2.5 VL Image$0.003/imageFullMost affordable
Seedream 4.0$0.0175/imageFullBudget quality
Leonardo Origin$0.035/imageFullArtistic styles

Image Models - Image-to-Image Edit (6 models)

ModelCostResolution ControlBest For
FLUX Kontext Pro$0.04/imageNone (fixed output)Smart image editing
Nano-Banana Pro I2I$0.0375/megapixelFullBudget editing
Nano-Banana I2I$0.0375/megapixelFullLegacy editing
Qwen2.5 VL Image Edit Plus$0.004/imageFullAffordable editing
Seedream v4 Edit$0.0175/imageFullBudget editing
Ideogram Edit$0.08/imageNoneText-focused editing

Image Models - Special (2 models)

ModelCostUse Case
FLUX Kontext Multi (Merge)$0.06/imageCombine multiple images
Remove Background$0.02/imageBackground removal

Video Models (35+ models)

CategoryModel CountCost Range
Image-to-Video (I2V)18 models$0.002 - $0.50/second
Text-to-Video (T2V)10 models$0.0024 - $0.50/second
Keyframe-to-Video (KF2V)5 models$0.0032 - $0.27/second
Video-to-Video (V2V)1 model$0.27/second
Reference-to-Video (R2V)1 model$0.27/second
Upscale1 model$0.02/second

What You Can Control vs. What Is Fixed

Understanding what parameters you can adjust is crucial for getting the results you want.

Parameters You CAN Control

These settings are available in the generation interface:

ParameterWhere UsedDescription
Prompt/DescriptionAll modelsThe text describing what you want to generate
Aspect RatioMost image and video modelsShape of output (16:9, 9:16, 1:1, etc.)
ResolutionSome image modelsOutput size (720p, 1080p, 2K, 4K)
DurationVideo models with rangesLength of video in seconds
Model SelectionAll generation typesWhich AI model to use

Parameters Fixed by the Model

These are determined by the model and cannot be changed:

ParameterDescription
Output File FormatModels output specific formats (MP4 for video, PNG/JPG for images)
Frame RateVideo models have fixed frame rates
Internal ProcessingSome models process at fixed resolutions regardless of input
Color DepthBit depth and color space are model-determined
Audio Sample RateAudio models have fixed quality settings

Audio Models

Audio models generate voice content from text, enabling realistic character dialogue.

ElevenLabs v2 (Default)

The default text-to-speech model, providing high-quality voice synthesis.

AttributeValue
Cost$0.003 per request
TypeText-to-Speech (TTS)
StatusDefault audio model

What you control:

  • Voice selection (from your cast members)
  • The text/dialogue to speak
  • Pacing through punctuation

What is fixed:

  • Audio sample rate and quality
  • Output format

Best for: General voice synthesis needs, shorter dialogue segments.


FAL ElevenLabs Dialogue v3

A cost-effective option for longer dialogue, charged per second of generated audio.

AttributeValue
Cost$0.0005 per second
TypeText-to-Speech (TTS)

What you control:

  • Voice selection
  • Dialogue text
  • Natural pacing through script writing

What is fixed:

  • Audio quality parameters
  • Output format

Best for: Longer dialogue sequences where per-second pricing is more economical.


Image Models: Text-to-Image (T2I)

These models create images from text descriptions. Use them to generate keyframes, asset images, and visual concepts.

Understanding Size Configuration

Different models handle image dimensions in different ways:

StrategyHow It WorksModels Using This
Aspect Ratio + ResolutionYou choose aspect ratio AND resolution separatelyFLUX1.1 [dev]
Aspect Ratio OnlyYou choose aspect ratio, model determines resolutionFlux Pro Ultra, Ideogram
Width/HeightYou specify exact pixel dimensionsImagen 4, Seedream, Leonardo

FLUX1.1 [dev]

A versatile development model offering excellent quality at reasonable cost.

AttributeValue
Cost$0.025 per megapixel
Size ControlFull (aspect ratio + resolution)

Supported Aspect Ratios:

  • 21:9 (Ultra-wide)
  • 16:9 (Widescreen)
  • 4:3 (Standard)
  • 1:1 (Square)
  • 3:4 (Portrait standard)
  • 9:16 (Vertical/mobile)
  • 9:21 (Ultra-tall)

What you control:

  • Prompt description
  • Aspect ratio
  • Resolution/image size

Best for: General-purpose image generation with flexible size options.


Flux Pro v1.1 Ultra

Premium quality model for high-end image generation.

AttributeValue
Cost$0.06 per image
Size ControlAspect ratio only

Supported Aspect Ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, 9:21

Special Feature: Supports "raw" mode for unprocessed aesthetic.

What you control:

  • Prompt description
  • Aspect ratio
  • Raw mode (true/false)

What is fixed:

  • Output resolution (model-determined based on aspect ratio)

Best for: Premium quality images where you need top-tier results.


Ideogram v2 / v2 Turbo

Excellent for images containing text, logos, or typography.

Attributev2v2 Turbo
Cost$0.05/image$0.02/image
SpeedStandardFast

Supported Aspect Ratios: 16:9, 4:3, 1:1, 3:4, 9:16, 10:16, 16:10

What you control:

  • Prompt description (including text to render)
  • Aspect ratio

What is fixed:

  • Output resolution

Best for: Any image that needs legible text, signs, logos, or typography.


Ideogram v3 / v3 Turbo

The latest Ideogram models with expanded aspect ratio support and improved text rendering.

Attributev3v3 Turbo
Cost$0.08/image$0.04/image
SpeedStandardFast

Supported Aspect Ratios: Includes all v2 ratios plus 3:1 and 1:3 for panoramic/banner formats.

Best for: Highest quality text rendering in images, banner and panoramic formats.


Imagen 4 Preview / Ultra / Ultra [Fast]

Google's photorealistic image generation models.

AttributePreviewUltraUltra [Fast]
Cost$0.04/image$0.08/image$0.06/image
QualityGoodBestVery Good
SpeedStandardSlowFast

Size Control: Full control with separate width and height parameters. The system provides presets for each aspect ratio and resolution combination.

What you control:

  • Prompt description
  • Width and height (via presets)
  • Aspect ratio

Best for: Photorealistic images, realistic human subjects, natural environments.


Nano-Banana Pro T2I

Budget-friendly option with full size control.

AttributeValue
Cost$0.0375 per megapixel
Size ControlFull (width/height)

What you control:

  • Prompt, width, height

Best for: Budget-conscious generation where you need size control.


Qwen2.5 VL Image

The most affordable text-to-image option.

AttributeValue
Cost$0.003 per image
Size ControlFull (width/height)

Best for: Rapid prototyping, concept exploration, high-volume generation.


Seedream 4.0

Balanced quality and cost for general use.

AttributeValue
Cost$0.0175 per image
Size ControlFull (width/height)

Best for: Everyday image generation with good quality-to-cost ratio.


Leonardo Origin

Known for artistic and stylized outputs.

AttributeValue
Cost$0.035 per image
Size ControlFull (width/height)

Best for: Artistic styles, illustrations, stylized content.


Image Models: Image-to-Image Edit (I2I)

These models modify existing images based on your instructions. Use them for edits, touch-ups, and transformations.


FLUX Kontext Pro

A smart image editing model that understands context and makes intelligent modifications.

AttributeValue
Cost$0.04 per image
Size ControlNone - fixed output size

Supported Aspect Ratios: Yes, but only affects composition, not output resolution.

What you control:

  • Edit prompt (what changes to make)
  • Aspect ratio preference

What is fixed:

  • Output resolution (~1344x786)
  • Cannot output at input resolution

Best for: Quick conceptual edits where final resolution is not critical. Use for previewing changes before applying with a different tool.

Not recommended for: Final production images, high-resolution requirements.


Nano-Banana Pro I2I / Nano-Banana I2I

Budget editing options with full size control.

AttributeProLegacy
Cost$0.0375/megapixel$0.0375/megapixel
Size ControlFullFull

What you control:

  • Edit prompt
  • Output width and height

Best for: Budget editing where you need to preserve or control resolution.


Qwen2.5 VL Image Edit Plus

Affordable image editing with good results.

AttributeValue
Cost$0.004 per image
Size ControlFull (width/height)

Best for: High-volume editing tasks, affordable touch-ups.


Seedream v4 Edit

Budget-friendly editing with full control.

AttributeValue
Cost$0.0175 per image
Size ControlFull (width/height)

Best for: General editing at reasonable cost.


Ideogram Edit

Specialized for editing images with text content.

AttributeValue
Cost$0.08 per image
Size ControlNone

Best for: Editing images that contain text, modifying typography, fixing text in images.


Image Models: Special Purpose

FLUX Kontext Multi (Merge)

Combines multiple images into a single coherent output.

AttributeValue
Cost$0.06 per image
TypeImage Merge

What you control:

  • Multiple input images
  • Merge prompt describing how to combine them

Best for: Character turnarounds, composite scenes, combining elements from multiple sources.


Remove Background

Automatically removes image backgrounds.

AttributeValue
Cost$0.02 per image
TypeBackground Removal

What you control:

  • Input image only

Output:

  • Transparent PNG with background removed

Best for: Extracting subjects from images, creating assets for compositing.


Video Models

Video models are categorized by their input type and generation approach.

Understanding Video Render Types

Render TypeFull NameInputOutput
I2VImage-to-VideoSingle image + promptVideo starting from that image
T2VText-to-VideoText prompt onlyVideo from scratch
KF2VKeyframe-to-VideoStart + end images + promptVideo interpolating between them
V2VVideo-to-VideoExisting video + promptTransformed video
R2VReference-to-VideoReference images + promptVideo using those references
UPSCALEVideo UpscaleExisting videoHigher resolution video

Video Duration: What You Can Control

Video models handle duration differently:

Duration TypeDescriptionExamples
FixedModel only produces one durationVeo 3 (8 seconds only)
RangeChoose from valid optionsKling (5 or 10 seconds)
FlexibleSpecify within a rangeVeo 2 (5-8 seconds)

Image-to-Video Models (I2V)

These models animate a single image into a video. The most common workflow is to create a keyframe and then generate a video from it.

Kling v2.1 Master I2V (Default)

The default and recommended I2V model, offering excellent quality and motion.

AttributeValue
Cost$0.2688 per second
Duration Options5 or 10 seconds
Default Duration5 seconds

What you control:

  • First frame (keyframe image)
  • Video prompt
  • Duration (5s or 10s)
  • Aspect ratio

Best for: High-quality character animation, professional production.


Kling Model Family (I2V)

ModelCost/SecondNotes
Kling v2.1 Master$0.2688Default, highest quality
Kling v2.0$0.0672Good quality, more affordable
Kling v1.6 Pro$0.1344Professional tier
Kling v1.5 Pro$0.112Legacy professional

All Kling I2V models support: 5 or 10 second durations, aspect ratio control.


Veo Models (I2V)

Google's video generation models.

ModelCost/SecondDuration
Veo 3$0.50Fixed 8 seconds
Veo 2$0.255-8 seconds

Best for: Photorealistic motion, natural movement.


LTX Video Models (I2V)

Ultra-affordable video generation.

ModelCost/Second
LTX 0.9.7$0.0024
LTX 0.9.5$0.0018

Best for: Prototyping, concept testing, high-volume generation.


Other I2V Models

ModelCostNotes
Minimax Video 01$0.12/secondGood motion quality
Luma Ray 2$0.0032/framePer-frame pricing
Luma Ray Flash 2$0.00048/frameFast, affordable
Sora Turbo$0.21/secondOpenAI model
Runway Gen-4$0.25/secondLatest Runway
Runway Gen-3 Alpha$0.10/secondPrevious Runway
Seedance 1.0$0.14/secondGood value
Wan 2.1$0.008/secondBudget option
Wan 2.1 1.3B$0.002/secondMost affordable
Pika 2.2$0.20/secondCharacter motion

Text-to-Video Models (T2V)

Generate videos from text descriptions without any input images.

When to Use T2V

  • Abstract concepts without visual reference
  • Establishing shots and backgrounds
  • Exploring ideas before creating keyframes
  • When no keyframe exists yet

Kling T2V Models

ModelCost/Second
Kling v2.1 Master$0.2688
Kling v2.0$0.0672
Kling v1.6 Pro$0.1344

Other T2V Models

ModelCostNotes
Veo 3$0.50/secondFixed 8s, premium quality
Veo 2$0.25/second5-8s range
LTX 0.9.7$0.0024/secondUltra-affordable
Minimax Video 01$0.12/secondBalanced option
Luma Ray 2$0.0032/framePer-frame pricing
Sora Turbo$0.21/secondOpenAI quality
Runway Gen-4$0.25/secondCreative results

Keyframe-to-Video Models (KF2V)

Generate videos that smoothly transition between a starting image and ending image.

When to Use KF2V

  • Cinematic transitions
  • Controlled camera movements
  • Character repositioning
  • When you know exactly how a shot should start and end

KF2V Models

ModelCost/Second
Kling v2.1 Master (Default)$0.2688
Kling v2.0$0.0672
Veo 2$0.25
Minimax Video 01$0.12
Luma Ray 2$0.0032/frame

What you control:

  • First frame (start image)
  • Last frame (end image)
  • Transition prompt
  • Duration (where applicable)

Video-to-Video Models (V2V)

Transform existing videos with style changes or enhancements.

Kling v2.1 Master V2V

AttributeValue
Cost$0.2688 per second

What you control:

  • Source video
  • Transformation prompt
  • Optional dialog sync

Best for: Style transfer, restyling existing videos, syncing to dialog.


Reference-to-Video Models (R2V)

Generate videos using reference images for character/scene consistency.

Kling v2.1 Master R2V

AttributeValue
Cost$0.2688 per second

What you control:

  • 1-N reference images (cast, locations, props)
  • Video prompt
  • Duration

Best for: Maintaining character consistency across shots, using specific asset appearances.


Upscale Models

Enhance video resolution and quality.

Topaz Video Upscale

AttributeValue
Cost$0.02 per second

What you control:

  • Source video
  • Scale factor
  • Target resolution
  • Target FPS

What is fixed:

  • Upscaling algorithm
  • Output format

Best for: Final polish before delivery, enhancing resolution of selected videos.


Choosing the Right Model

For Images

ScenarioRecommended Models
General purposeFLUX1.1 [dev], Seedream 4.0
Highest qualityFlux Pro Ultra, Imagen 4 Ultra
Text in imagesIdeogram v3, Ideogram v2
Fastest generationIdeogram Turbo variants, Qwen2.5
Budget-consciousQwen2.5 VL ($0.003), Seedream ($0.0175)
PhotorealisticImagen 4 family
Artistic stylesLeonardo Origin

For Video

ScenarioRecommended Models
Best qualityKling v2.1 Master, Veo 3
Good valueKling v2.0, Minimax Video 01
Budget generationLTX Video, Wan 2.1
Fast prototypingLuma Ray Flash 2, LTX
Character animationKling family, Pika
PhotorealisticVeo family, Sora

For Audio

ScenarioRecommended Model
Short dialogueElevenLabs v2
Long dialogueFAL ElevenLabs Dialogue v3

Cost Optimization Tips

Image Generation

  1. Use Qwen2.5 for prototyping ($0.003) then switch to premium for finals
  2. Ideogram Turbo vs Standard - Turbo is half the price with good quality
  3. Megapixel pricing - FLUX1.1 dev charges by megapixel; smaller images cost less

Video Generation

  1. Use LTX for concepts ($0.002/second) before committing to expensive models
  2. 5 seconds vs 10 seconds - Choose shorter durations when possible
  3. Upscale last - Only upscale your selected, final videos
  4. Per-frame vs per-second - Calculate which is cheaper for your video length

General Tips

  1. Generate multiple variations with budget models first
  2. Refine your prompt before using premium models
  3. Track credit usage in Budget Management
  4. Delete rejected content to declutter your workspace

Common Questions

Why can't I control the resolution on some models?

Different AI models have different architectures. Some are trained to output specific resolutions and cannot change them. Models like FLUX Kontext Pro produce fixed-size outputs regardless of input. Always check the model's "Size Control" specification before generating.

Why is my edited image smaller than my original?

If you used FLUX Kontext Pro for image editing, it outputs at approximately 1344x786 regardless of input size. For edits that preserve resolution, use Nano-Banana I2I, Qwen2.5 VL Edit, or Seedream v4 Edit.

Why can Veo 3 only make 8-second videos?

This is a fundamental characteristic of how Veo 3 was trained. The model optimizes for quality at a specific duration and does not support other lengths. If you need different durations, use Veo 2 or Kling models.

What's the difference between I2V and R2V?

AspectI2VR2V
InputSingle keyframe1-N reference images
PurposeAnimate specific frameMaintain consistency
OutputVideo starting from keyframeVideo influenced by references
Best forStandard shot generationMulti-shot consistency

Why are some models more expensive?

Price generally correlates with:

  • Output quality - Higher fidelity, better details
  • Speed - Faster generation often costs more
  • Features - More controllable parameters
  • Training cost - Models trained on larger datasets

How do I know which models support my aspect ratio?

The interface filters models by compatibility. When you select a generation mode and aspect ratio, only compatible models appear in the dropdown. You can also refer to the detailed model information in this guide.

Can I use multiple models together?

Yes! A common workflow is:

  1. Generate keyframe with an image model
  2. Create video with an I2V model
  3. Upscale with the upscale model
  4. Add dialog with an audio model

Each step can use a different model optimized for that task.


Next Steps