The economics of video will shape what people can build with it. Our research begins with a question: how much less work could it take to produce a useful visual result?
A finished film can be created once and watched by millions. An interactive experience has a different equation. It may generate something new for every person, throughout a session. As the experience grows, the cost grows with it.
Making video an everyday AI capability therefore requires a different cost curve. The goal is to give creators and developers room to generate, revise, and interact as a normal part of using an application.
There are two costs to bring down
The first is the cost of a single generation. At a high level, this depends on how much computation is required and how efficiently it runs. A useful cost analysis asks how much work is necessary for a result at the desired quality, and where repeated effort can be reduced.
The second is the cost of getting a result someone can actually use. If a creator has to try several times to get the right character, motion, or camera angle, every attempt belongs in that bill. Better control and consistency could reduce these retries, creating a second source of savings.
Two sources of improvement
Why a different cost curve is possible
A generation bill combines the amount of visual information produced, the computation used per pass, the number of repeated passes, and the hardware throughput doing useful work. These factors can compound: less work per pass and fewer passes can together produce a larger saving than either change alone. Our bet is that visual intelligence still has substantial room to become more efficient.
We describe the direction simply: scale logic up. Scale vision down. The aim is to make capability and efficiency improve together.
Several improvements can compound. Reducing the computation required for an attempt, running that computation more efficiently, and needing fewer attempts all affect the final cost. The value of any one improvement depends on what happens to the rest of the system.
Illustrative arithmetic
Same desired result100 cost units per attempt
4 attempts per usable result
10 cost units per attempt
2 attempts per usable result
Compare the whole experience
A smaller bill for one part of a system does not establish a smaller bill overall. Any comparison needs to include the full inference and serving cost, and use the same duration, resolution, quality requirement, and workload. Response time and failed attempts matter, too.
API prices are another layer. They can include provider margins and packaging choices, so a change in an API bill is not automatically a change in underlying generation cost. We separate those questions when evaluating the economics.
Our current ambition is a 100× reduction in generation cost, with target economics of $1 per hour of generated video. These are research targets. Proving them requires comparable visual quality and a complete accounting of the system’s cost.
Lower cost changes the medium
The reward is more than a cheaper clip. Lower cost could let a creator explore more alternatives, a game generate throughout a session, or an interface respond visually to each person. Affordability makes room for uses where generation happens continuously.
That is the economic idea behind our work: make useful visual generation inexpensive enough to become part of everyday creation and interaction.
From economics to possibility
Read The Visual Renaissance for the broader ambition, or explore our recorded research prototypes.
