We want video generation to become an everyday AI capability. A medium people can create with, think with, and interact with.
A remarkable clip gives a glimpse of what is possible. An everyday capability goes further. It lets people explore an idea, change their minds, and keep going. It becomes part of a workflow, a story, or an experience.
We call that ambition the Visual Renaissance: a new period of visual expression, made possible by models that are useful enough and accessible enough to use routinely.
A medium for motion and change
Text can describe an idea. An image can show a moment. Video can show how something changes. It brings motion, space, and time into the way we communicate, and opens possibilities beyond a finished piece of content.
Describe an idea
Show a moment
Explore how it changes
A creator could explore different directions for a story. A player could shape what happens next. A visual interface could respond to the task in front of someone. Each asks for more than an attractive frame. The result needs to follow an intention and remain coherent as the experience develops.
A different direction
Our guiding idea is deliberately simple: scale logic up. Scale vision down. We are exploring a different balance for visual intelligence, with a goal of expanding what models can do while making them more practical to use.
A different balance
Scale logic up.
Scale vision down.
The measure of progress is the experience it makes possible: a useful revision, a coherent scene, an interaction that feels responsive. Strong visual quality, control, and affordable generation all belong to the same ambition.
From clips to experiences
Our research explores video and UGC, generative gaming, generative software, and robotics. These settings help us ask concrete questions of the models we are building.
For creation, can someone shape a story without starting over at every revision? For games, can an experience develop in response to a player? For software, can a visual interface become more useful to the person working with it? For robotics, can visual generation contribute to understanding possible actions and outcomes?
These are research questions. Our recorded prototypes communicate the directions we are exploring, and the possibilities that motivate us.
Everyday means room to explore
When each attempt is expensive, people ration their experiments. When a revision is slow, they stop exploring. Bringing down the cost of a useful result gives people more room to create and gives developers more room to build.
That is why economics is central to the model capability. Our ambition is a 100× reduction in generation cost and target economics of $1 per hour of generated video, alongside strong visual quality. These are research goals that we are working to validate.
We want the next generation of visual models to make an experience possible throughout a session, as naturally as text and images are used today.
Build the capability. Discover the possibilities.
Umwelt Lab brings research and engineering together around this goal. Our team’s experience spans visual generation, world models, 3D, and large-scale model training, with research that has reached products and interactive systems.
Our focus is the foundation model capability and what people can do with it. The Visual Renaissance begins when visual generation becomes a medium people can work with every day.
The economics behind the ambition
Read The Economics of Everyday Video or explore our recorded prototypes.
