First and last frame AI video: control where a clip starts and ends
Most AI video prompts control the start and leave the ending to chance. Give the model a first frame and a last frame and you decide both: land on a logo, reveal a furnished flat from an empty one, turn a 2D plan into 3D, or make clips that join cleanly. Here is which models do it and how to make the frames.
· 5 min read · 763 views · by the Cortex Digital Hub team
Pick Lumen Frames, Lumen Omni or Prism 6, drop your start image and your end image, and describe the change between them. The clip opens on one and lands on the other. From ₹24 a clip · no watermark · runs in your browser
Make a video → → Opens AI Video Generator · runs in your browser
Most people prompting an AI video control the beginning and hope for the ending. You drop a photo, describe the motion, and the model decides where the camera is when the clip stops. That is fine for a mood piece. It is a problem when the clip has to end on your logo, on the furnished version of the flat, or on the exact view the next clip starts from. A first frame and a last frame fix that, and the AI Video Generator has three models that take both.
Why the last frame matters
There are three jobs where the ending is the whole point.
Landing on a logo or a pack shot. An ad for a Hyderabad sweet shop can wander through the kitchen for eight seconds, but it needs to finish on the box with the name readable. With a last frame you give the model the pack shot and it steers towards it.
Matching the next clip. No model on the site goes past 30 seconds and most stop at 15, so anything longer is made in parts and joined. Parts join cleanly only when the end of one is the start of the next. A last frame on clip one that is also the first frame on clip two gives you a cut nobody notices.
Before and after reveals. Empty to furnished, day to night, plan to render, boxed to unboxed. These are the most-shared clips in real estate and D2C because the transformation does the talking. Without an end frame the model invents its own "after"; with one, the after is your actual render or your actual product.
Which models take both frames
As of October 2026, three models in the picker accept a first and a last frame.
Lumen Frames is the default choice. It makes clips up to 15 seconds with sound, takes references as well, and costs 23 credits a second, so a 5-second clip is 115 credits (₹57.50) and a 15-second one is 345 (₹172.50). Lumen Omni makes 5- or 10-second clips at 30 credits a second (150 for 5 seconds) and also takes references and edits. Prism 6 makes 1- to 15-second clips at 1080p with sound at 52 credits a second (260 for 5 seconds); use it when the final has to be full HD.
Every other model takes a single image at most. On those, your image becomes the first frame and the ending is up to the model.
How to make a matching pair of frames
The model interpolates between the two frames, so the more they agree on camera position and framing, the smoother the clip. Two ways to make a pair that agree.
Edit the same photo. Take one image and make the "after" version of it in the photo editor: brighten or darken it for day to night, paste the furniture in, drop the product out of the box. Because both frames come from the same source, the walls, the floor and the horizon line up and the model only has to animate the difference.
Use a still from the previous clip. For stitching, the last frame of clip one is the only correct first frame of clip two. Open the finished clip in the trim tool, cut it to the final moment, and save that frame. The trim without re-encoding guide shows how to do it without the frame going soft.
Whichever route you take, resize both frames to the same pixel dimensions with the image resizer before you upload. A 1080x1920 first frame and a 1080x1350 last frame forces the model to invent the missing strip.
Four reveals that work
| Scenario | First frame | Last frame | Prompt |
|---|---|---|---|
| Empty flat to furnished | Site photo of the bare living room | 3D render of the same room, furnished, from the same corner | "Camera holds still. Furniture, curtains and lamps fade into place one by one, warm evening light, no people." |
| Raw plan to 3D plan | The sanctioned 2D floor plan, flat on white | The 3D render of the same plan, same orientation | "Slow tilt down. The flat drawing rises into walls and rooms, top-down view, clean white background." |
| Product in box to out of box | Closed box on a table, front view | Same table, box open, product standing beside it | "Static camera. The lid lifts, the product rises out and settles on the table, soft studio light." |
| Day to night | Building exterior at 4 pm | Same exterior, edited dark with windows lit | "Fixed camera. Sky turns from blue to deep orange to night, windows light up floor by floor." |
The plan-to-3D pair is the easiest to make because the site produces both frames for you: upload the drawing to the 2D to 3D floor plan tool, keep the render, and you have a matched pair with identical orientation.
Writing the prompt between two frames
With two frames in place, the prompt's job shrinks. The model already knows what the start and end look like, so describe the camera and the change, nothing else.
Keep the camera still or give it one move. "Static camera" or "slow push in" is enough. Two frames plus a wandering camera is asking the model to solve three things at once, and it usually drops one.
Describe the transition, not the objects. "Furniture fades in one by one" tells the model how to get from frame A to frame B. Listing the sofa, the rug and the lamp does not, because they are already in frame B.
Leave people out of before/after clips. A person who exists in one frame and not the other has to appear or vanish, and that is where the model's seams show.
If a 5-second draft lands on the right frame but rushes the middle, make the clip longer rather than rewriting; a 10-second Lumen Frames clip is 230 credits and gives the transition room.
Stitching parts into one longer video
A 45-second clubhouse walkthrough for a Bangalore launch is three 15-second clips on the same model, each with its first frame taken from the previous clip's last frame. Keep the model, the resolution, the aspect ratio and the lighting phrase in the prompt the same across all three, or the join will show as a colour shift even when the geometry lines up.
Once the parts are done, join them in the free video converter, check the cuts, and then size the result for where it is going: 9:16 for a Reel or a WhatsApp status, 16:9 for a brochure microsite. Clips stay in your history for seven days, so download each part as it finishes. The sharing floor plans on WhatsApp post covers the compression step for channel partners who will forward it on.
Frequently asked questions
Which models support first and last frame control?
As of October 2026, Lumen Frames (up to 15 seconds, with sound), Lumen Omni (5 or 10 seconds) and Prism 6 (1 to 15 seconds at 1080p, with sound). The other models accept a first frame only or no frame at all.
Do the first and last frames have to be the same size?
Keep them the same aspect ratio and ideally the same pixel size. If one is cropped differently the model has to invent what is missing at the edges. Resize both to the same dimensions before uploading.
Can I use the last frame of one clip as the first frame of the next?
Yes, and that is the standard way to build something longer than one model's limit. Trim the previous clip to its final moment, save that frame, and use it as the first frame of the next clip with the same model and style.
Does an end frame cost extra?
No. Credits are charged per second of finished video at the model's rate. A 5-second clip on Lumen Frames is 115 credits (about ₹58) whether you give it one frame or two.