All journal articles

Research direction

The Visual Renaissance

A new visual medium, and a different direction for making it everyday.

Minimal conceptual illustration, The Visual Renaissance
Conceptual illustration. Not model output.

We want video generation to become an everyday AI capability. A medium people can create with, think with, and interact with.

A remarkable clip gives a glimpse of what is possible. An everyday capability goes further. It lets people explore an idea, change their minds, and keep going. It becomes part of a workflow, a story, or an experience.

We call that ambition the Visual Renaissance: a new period of visual expression, made possible by models that are useful enough and accessible enough to use routinely.

A medium for motion and change

Text can describe an idea. An image can show a moment. Video can show how something changes. It brings motion, space, and time into the way we communicate, and opens possibilities beyond a finished piece of content.

Text

Describe an idea

Image

Show a moment

Video

Explore how it changes

A visual medium for both creation and interaction.

A creator could explore different directions for a story. A player could shape what happens next. A visual interface could respond to the task in front of someone. Each asks for more than an attractive frame. The result needs to follow an intention and remain coherent as the experience develops.

A different direction

Our guiding idea is deliberately simple: scale logic up. Scale vision down. We are exploring a different balance for visual intelligence, with a goal of expanding what models can do while making them more practical to use.

A different balance

Up.

Scale logic up.

Down.

Scale vision down.

More capable. More accessible. A research direction for visual intelligence.

The measure of progress is the experience it makes possible: a useful revision, a coherent scene, an interaction that feels responsive. Strong visual quality, control, and affordable generation all belong to the same ambition.

From clips to experiences

Our research explores video and UGC, generative gaming, generative software, and robotics. These settings help us ask concrete questions of the models we are building.

For creation, can someone shape a story without starting over at every revision? For games, can an experience develop in response to a player? For software, can a visual interface become more useful to the person working with it? For robotics, can visual generation contribute to understanding possible actions and outcomes?

A recorded exploration of generative software from our research materials.

These are research questions. Our recorded prototypes communicate the directions we are exploring, and the possibilities that motivate us.

Everyday means room to explore

When each attempt is expensive, people ration their experiments. When a revision is slow, they stop exploring. Bringing down the cost of a useful result gives people more room to create and gives developers more room to build.

That is why economics is central to the model capability. Our ambition is a 100× reduction in generation cost and target economics of $1 per hour of generated video, alongside strong visual quality. These are research goals that we are working to validate.

We want the next generation of visual models to make an experience possible throughout a session, as naturally as text and images are used today.

Build the capability. Discover the possibilities.

Umwelt Lab brings research and engineering together around this goal. Our team’s experience spans visual generation, world models, 3D, and large-scale model training, with research that has reached products and interactive systems.

Our focus is the foundation model capability and what people can do with it. The Visual Renaissance begins when visual generation becomes a medium people can work with every day.

The economics behind the ambition

Read The Economics of Everyday Video or explore our recorded prototypes.