Command search

Building Worlds Beyond the Screen With Google DeepMind

Experimentation typically begins with a question and the instinct to chase an answer. For Moment Factory and Google DeepMind, that instinct was the foundation of a collaboration built on shared curiosity and the belief that, together, they could unlock a world of possibility.

The question they posed was, “What happens when generative AI video leaves the screen?”

To answer it, Moment Factory provided the constraints and realities of large-scale  location-based experiences, while Google DeepMind contributed its expertise in AI, across both multimodal and world models. Together, they formed a team of researchers, artists, and technologists who explored generative AI video in real-world spatial applications, testing its limits beyond traditional screen-based contexts.

In recent years, generative AI video has evolved within a familiar set of boundaries: short clips, standard screens, and carefully controlled viewing conditions. Inside those limits, progress has been remarkable, and the promise of generative AI video has reached artists, professionals and AI first-timers at unprecedented speeds. Over time, better models have produced sharper clips, fewer glitches, and more convincing realism — on screen.

A small rectangle can be incredibly forgiving; a massive architectural canvas less so.

Architecture has its own scale, texture, and intent. Layered with visual content, it can be completely transformed to transport audiences to other worlds. But that only works if the illusion holds over time, as people move through space and examine it up close. A minor glitch in a traditional screen can translate into a 10-meter-wide rupture in an immersive environment: a fatal flaw in the experience.

Seeing whether generative AI video could survive spatial integration required testing across multiple environments of varying complexity. The experiments began in Montreal, first at Moment Factory’s AI Studio, a dedicated prototyping space, then at the Society for Arts & Technology’s immersive dome, the Satosphère. Eventually, the team was off to New York City, where the ornament, scale, and architectural intricacy of Cipriani 25 Broadway provided the experiment's toughest challenge yet.

Each space raised the stakes. Scale exposed what screens concealed, prolonged periods of time revealed repetition, and the human mind’s natural ability for pattern recognition proved a relentless challenger.

There were no shortcuts or magic presto buttons. Every obstacle made the work smarter. Every bit of trial, debate, questioning, and trust was needed for the team to arrive at a breakthrough. They landed on a modular tiling system that divides mega canvases into smaller, more manageable units capable of maintaining crisp detail, high fidelity, and strong pixel density.

Above all, it’s a system made by and for artists. It relies on their judgement, instinct, expertise, and that indispensable human eye, while giving them the power to shape, stretch, and choreograph AI-generated content across physical spaces and into immersive experiences that other humans can enjoy. 

Throughout their partnership, the teams at Moment Factory and Google DeepMind were reminded that those who remain open to wherever an experiment’s guiding question leads may create the conditions for endless, exciting answers, not just one. In this case, they believe they're onto something that has the potential to shift how we collaborate on location-based entertainment in the future and what we are able to create next.

Read about the experiment and collaboration in this case study

Here is a selection of projects that illustrate what we can create together.

--