How an AI video generator works: text, photo and references to video
An AI video generator turns a sentence, a single photo or a handful of reference images into a short clip. Here is what you give it, what it does with that, and what comes back, with credits, wait times and a first-clip walkthrough.
· 5 min read · 1.2K views · by the Cortex Digital Hub team
Type a sentence or drop a photo, pick a length and format, and the button tells you the exact credits before you commit. Your first clip is a few minutes away. From ₹24 a clip · no watermark · runs in your browser
Make a video → → Opens AI Video Generator · runs in your browser
An AI video generator takes something you can describe or show, a sentence, a product photo, a few reference images, and returns a short clip with movement no camera filmed. The AI Video Generator on this site puts 21 models behind one form, so the useful question is "what do I give it, and what comes back". This post answers that, then makes a first clip for a Hyderabad bakery so you can see the whole loop.
The three things you can give it
Every clip starts from one of three inputs, and the generator reads the mode from what you drop in.
The first is a prompt and nothing else. You type what you want to see, and the model invents the picture from scratch: subject, setting, light, camera. This is text to video, and it is the right choice when you have no photo, or when the photo you have is too weak to build on.
The second is a single image. Drop one photo and it becomes the first frame of the clip. The model does not reinterpret your picture; it starts from it and adds motion, so your prompt should describe only how things move, such as "the camera pushes in slowly". This is image to video, and for a small business it is usually the most useful mode, because the product, the flat or the face stays recognisably yours.
The third is several references. Drop two or more images, a short video or an audio file, and the model treats them as things to keep to rather than as a starting frame. Two to four clear pictures of the same person or product hold them consistent through the clip. A video reference drives the motion, which is what Motion control does.
| What you give | Mode | What the model does | Good for |
|---|---|---|---|
| A sentence, no files | Text to video | Invents subject, setting, light and camera | Festival teasers, concept shots, b-roll |
| One photo | Image to video | Uses your photo as frame one, adds only motion | Products, flats, portraits |
| 2–4 photos of one product or person | References | Keeps that product or face consistent in a new scene | Mascots, a model in several settings |
| A still plus a short video | Motion control | Moves your still the way the video moves | Gestures, product spins |
| A finished clip plus a prompt | Edit or Extend | Changes an element, or continues the clip by 4–30 s | Fixes, longer stories |
What the model does with it
A video model has learned from an enormous amount of footage how the world tends to move: how steam drifts, how a dolly feels, how fabric settles. When you press Generate, it starts from noise (or from your first frame) and refines every frame at once, guided by your prompt, until the sequence looks like something that could have been filmed. It is predicting plausible motion, not retrieving a stock clip, which is why the same prompt gives a different result each run unless you fix the seed.
Each model has its strengths. Lumen Motion 3 is the all-rounder and makes multi-shot clips with native sound. Vela 2.5 is the most cinematic and goes up to 30 seconds. Nova 2K gives the sharpest single frame. Orbit 3 and Nova Lite are the cheap draft models for testing a prompt before you spend on a big one. Sound, when the toggle is on, is generated with the picture: ambience, effects and music that follows the motion, with no voice-over unless the prompt asks for speech.
Length, format and resolution
Length runs from 3 to 30 seconds depending on the model, and most business clips do their job in 5 to 10. Format is the aspect ratio: 9:16 for Reels, Shorts and WhatsApp Status, 16:9 for YouTube and websites, 1:1 for a Meesho or Flipkart tile. Resolution goes from 480p drafts to 1080p for most work, and 2K or 4K where a model offers it.
Beyond those there is a Quality switch (Standard or High), a Bitrate setting, and camera presets that write a dolly or an orbit into your prompt for you; the General preset leaves it to you. Under Advanced sit an Avoid box for things you do not want, Prompt strength (default 0.5), an optional Seed, and "let the model expand my prompt" for short descriptions. Most first clips need none of these.
Credits: you see the cost before you press anything
Credits are charged per second of finished video, and 1 credit is ₹0.50. The Generate button shows the exact total for your model, length and resolution. Credits are deducted when the clip starts and returned automatically if the model fails or declines the prompt.
As of October 2026, the ₹249 pack gives 520 credits, which is roughly four 5-second clips on Lumen Motion 3 Standard at 23 credits a second. Nova Lite at 8 credits a second takes a 6-second draft down to 48 credits, about ₹24. Vela 2.5 at 1080p sits at the other end, 421 credits a second, so the usual advice is to draft cheap and spend only on the final.
The wait, the history and the watermark
A 5 to 10 second clip at 1080p usually takes 2 to 4 minutes. 4K and 30-second clips can take 8 minutes or more. You can leave the page; clips in progress finish on their own and appear under "Your videos".
Clips stay in that history for 7 days, then they are deleted, so download the ones you want. Your image goes to the video model only to make the clip, and nothing you upload is used to train anything, the same principle behind the browser-based tools on this site. Downloads carry no watermark and no logo.
Your first clip: a bakery in Kondapur
Say you run a bakery and want a Reel for the Diwali box. Take a phone photo of the box open on the counter, in window light. If the counter is cluttered, lift the box out with the background remover first.
Open the generator, drop the photo, and it becomes the first frame. Pick Lumen Motion 3 Standard, 5 seconds, 9:16, 1080p, sound off since you will add a track in Instagram. Type one line of motion: "the camera pushes in slowly on the open box, a little steam from the fresh batch". The button reads 185 credits, about ₹92. Press it, and the clip is waiting in Your videos in a few minutes.
The same loop works for a flat listing. Drop the exterior photo, choose 16:9 and 10 seconds, and write "the camera rises slowly from the gate to the top floor, clear morning sky". The 3D plan to video tool makes the interior walkthrough to pair with it.
Finishing the clip for free
Once the download lands, the rest is on-device and free. Resize to 1080x1920 if you generated in 16:9 and want a Status version, trim the first half second if the motion starts slow, and compress for WhatsApp before sending it to a channel partner. All of the video tools run in your browser, and none of them send your clip anywhere.
Frequently asked questions
Do I need a photo to use an AI video generator?
No. With a prompt alone the model invents the whole picture. A photo is only needed when the clip must show your actual product, flat or face, in which case it becomes the first frame.
How long does a clip take to generate?
Usually 2 to 4 minutes for a 5 to 10 second clip at 1080p. 4K and 30-second clips can take 8 minutes or more. You can leave the page; finished clips appear under Your videos.
Is there a watermark on the clips?
No. Clips download without a watermark or logo, and they stay in your history for 7 days before being deleted.
What happens to my photo after the clip is made?
It is sent to the video model only to make that clip. Nothing you upload is used to train anything.