Three modes sit in the composer, and they look like three ways to do the same thing. They are not. They run on different engines, they cost different amounts, and the one thing that will cost you the most is reaching for the expensive one when the free one is better.
Here is the whole decision, in one line each:
- Image : you want a picture that does not exist yet.
- Video : you want motion that does not exist yet.
- Composition : you want your own facts, on screen, exactly right.
Composition is the one people skip, and the one that wins most often
Start here, because it is counter-intuitive.
Image and Video are generative: a model invents pixels. Composition is deterministic: it fills a template you chose with text you wrote, and renders it. Same input, same output, every time. No surprise, no re-roll, no "close but the text is misspelled".
That matters more than it sounds. A generative model cannot reliably render your price, your percentage or your product name: it approximates letterforms. Composition does not approximate anything, because it is not generating text, it is typesetting it.
So: a number, a claim, a comparison, a quote, an announcement, a step list. Anything where being wrong is worse than being pretty. That is Composition.
Ten shapes, six formats, and the shapes declare which formats they support:
| Shape | What it is for | |---|---| | Hook + proof | A strong claim, then the facts that hold it | | Steps | A sequence the viewer has to follow | | Feature showcase | One capability, shown rather than described | | Chart reveal | A number that moves | | Stat card | One number, made large | | Quote | Someone's words, attributed | | Comparison | Two columns, honestly | | Announcement | Something new, dated | | OG card | The image a link shows when shared | | Icon tile | The small square, for a grid |
Formats: reel-9x16, landscape-16x9, square-1x1, portrait-4x5,
og-1200x630, tile-512.
How to fill one
Write one line per field, in order. The composer reads your lines and maps them onto the shape's fields. Required fields are filled first; what is left over goes to the optional ones.
The Studio renders 10 shapes in 6 formats
Deterministic: same input, same file, every time
No re-rolls, no misspelled prices
Three lines on a Hook + proof shape gives you the claim, then two supporting facts. Give it one line and it tells you which field is starving rather than inventing content for it.
Image: when nothing exists yet
Reach for Image when the picture has to be made: a scene, a mood, a product in a context you have not photographed.
What a good prompt has, in order of how much it changes the result:
- The subject, concretely. "A ceramic mug" beats "a product".
- The setting. "On a linen tablecloth, morning light through a window" is the difference between a stock look and a brand look.
- The framing. Close-up, three-quarter, flat lay.
- The mood, last. "Warm, calm" is a nudge, not a description.
What does not help: adjective stacking. "Beautiful, stunning, professional, high-quality, 8k, masterpiece" moves nothing. The model already tries to be good. Spend those words on the subject instead.
References travel with the prompt. Drop an image into the composer and it conditions the result: a product you already photographed, a colour palette, a pose. That is almost always faster than describing the same thing in words.
Video: when the motion is the point
Video generates a clip from your prompt, and optionally from a start frame you supply. It is billed per second, so the duration pill is not a preference, it is the price.
Two things decide whether a clip is usable:
- One action, not three. "The mug steams gently" works. "The mug steams, then a hand enters, then the camera pulls back" gives you three half-actions in five seconds.
- The start frame. Supplying one turns "generate something like this" into "animate exactly this". If you already have the image, use it: you get your product rather than a model's idea of it.
The mistake that costs the most
Using Video for something Composition renders perfectly.
A stat card with your number, animated, is a Composition. Asking a video model for "a card showing 47% with a subtle animation" costs per second, takes minutes, and comes back with 4l% or 47/. about as often as not.
The rule: if the words on screen have to be right, it is a Composition.
What the composer will not do yet
Two modes carry a PLANNED badge: Nova and Replace. They are
visible so the shape of the toolbar does not change under you, and they
are disabled because nothing behind them generates anything. A button
that looks like it works and does not is the one thing this design
system refuses.
When they ship, they ship with an engine, not with a badge removed.