Where is my robot butler? Why we’re in a watershed moment for world models

Michael Davies, PhD, Lead Data Scientist at Faculty, explores why world models are the missing ingredient for enterprise AI - and why building a digital representation of your business could become the critical source of competitive advantage.

2026-09-30
Frontier

Just a few weeks ago, Tesla announced they have rolled their Optimus robot onto their pilot production line, and Elon has promised that it could be in our homes in 2 years. 

It immediately took me back to my interview to join Faculty, where I was asked “why are you interested in AI?”.

‘Where is my robot butler?’ was my response.

My logic? It was back in 2022, and two recent breakthroughs had promised that my days of getting up off the couch, heaving myself across my flat to the fridge, opening the door, finding a beer, closing the door, and nestling back to my couch should long be in the past.

First, computer vision had been solved* for quite some time already then, around 7 to 10 years, with the breakthroughs of deep learning. 

Second, Boston Dynamics had recently dropped their viral video of their robots impressively break dancing to the Motown classic “Do you love me?”.

So it seemed perfectly reasonable, on the face of it, that a robot can not only see my fridge, it can moonwalk its way over there, open the fridge, grab me a beer, then moonwalk its way back to me. Perhaps congratulating me on “yet another excellent choice, sir”.

*slightly facetiously, I am defining “solved” here as surpassing human-level accuracy and replacing traditional hand-crafted algorithms across core visual benchmark tasks.


The missing ingredients

What is the problem then? It would be reasonable to jump to cost. But the gap is more fundamental. It is the reason Optimus is “two years away” (not the first time someone’s said that in the field of robotic butlers…) and, it turns out, it is a gap that is also sinking most enterprise AI. It’s referred to as the world model gap. Let me explain.

The dancing robot has extraordinary individual capabilities. It can moonwalk into a backflip and that’s a party trick I’d like to have in the locker.

However, a robot that can moonwalk has no idea that it should, or when it should.

The individual capabilities are impressive, but they are just that, individual. And individual capabilities do not add up to coherent, intelligent behaviour. 

To actually fetch my beer, the robot is missing two key ingredients.

(1) Agency: a brain to decide and sequence its capabilities towards my goal. 

(2) A world model: an understanding of its environment, e.g. an updating map of my flat and its inhabitants, so that “go to the fridge” actually means something.

World models are the new frontier

That was my pitch back in 2022: just solve two giant frontiers of technology, and my days of schlepping to the fridge would be over.

In the few years since, AI has rocketed in the collective consciousness from "if you know, you know” to “what rock have you been living under?!”, driven by leaps in that first ingredient, agency. “Agent” is now part of our daily vernacular. “World model” less so.

Looking now in 2026 at my robot butler, the agency required — “Should I grab one beer or two? Fetch it now or wait for the commercial break?” — is now trivial compared to frontier-model capabilities.

Yet, here I am, still walking back-and-forth to the fridge.

For all that agency, the world model is entirely missing. So my robot has no idea that my fridge is in my kitchen, that my hypothetical cat is blocking the fridge door, or that I already drank the last beer.

Let’s step out of my flat, into the wider world. The core principle comes out the front door with us. My butler is the toy version. The trillion-dollar version is playing out across the world’s enterprises putting AI agents to work.

Consider a large pharmaceutical company designing a clinical trial. An agent today can analyse previous trial data to identify high-performing clinical sites, recommend potential patient recruitment strategies and scan the competitor landscape. Impressive, but notice that each of these is a reading from the past. It tells you what trials like this have looked like.

To actually help a development team decide, the agent needs to understand the world those decisions play out in. How will the three other trials competing for the same pool of patients affect my recruitment plan? If a key site drops out at month 6, how does that ripple through timelines and costs everywhere else? And how do each of my decisions ripple across our clinical trial portfolio as a whole?

For agents to confidently answer these critical operational and strategic questions, what’s missing is a true understanding of the system they operate within.

Indeed, the biggest question facing large enterprises now is not “what can my agents do?” but “how do I enable them to deliver the results I want?”. World models are the answer.


What are world models?

So, what do I mean by world model?

A world model is, at a fundamental level, a representation of how a system works. In software, it is a digital simulation. One that allows AI to build a “mental model” of the environment it is operating in, and thus understand cause and effect.

The mechanism is fundamentally different from an LLM; which operates on a game of pattern matching to predict what comes next. An LLM predicts what sounds right based on past text; a world model simulates what will actually happen if a specific action is taken.

Take my robot in 2026. Having read the entire internet, it knows exactly what opening a fridge looks like, but not what happens if it opens this particular fridge, right now: whether the door clears my sleeping cat, or sweeps a glass off the counter. Prediction knows fridges in general; a world model knows mine.

Cause and effect. Actions and reactions. Playing parallel realities against each other, to then map the optimum path forward. 

This is how the businesses of tomorrow will power themselves. And we don’t need to look to our favourite Sci-Fi novels to imagine it. Some already do.

Self-driving cars are an intuitive example, as physical world models are being built. Residents of London are all too aware of the very manual world-model building that Waymo is currently undertaking.

Or look to finance, where digital world models have long been around. High-frequency trading firms don’t simply predict the market. Prediction is not decision: a forecast that tells you the market is heading one way is worthless if your own subsequent trade then sends it the other. Instead, they model how the order book will react to potential trades. The battleground here is autonomous agents, operating inside a world model, modelling cause and effect, to make optimal decisions at sub-microsecond speeds.


A watershed moment

We are in a watershed moment for world models then, not because they are a new invention, but because the individual capabilities of agents have reached new levels, leading to an incredible growth in the number of domains in which world models are now readily applicable. 

What was once confined to self-driving cars, my future butler, or highly specialised areas such as financial trading, is now spreading into the whole of “intelligent work”: supply chain and logistics, insurance pricing, workforce planning, marketing spend allocation, telecoms, traffic routing, hospital admissions and discharges, drug discovery, clinical trial design.

Everywhere players are now racing to develop their competitive edge. And two questions come up again and again.

First, “How do I get the best results for my business with agents?”. 

As we’ve discussed, world models are the essential missing ingredient. If AI agents are your new workers, world models are your new offices. 

Second, “How do I build and maintain a competitive strategy, in an environment where the major gains in productivity are coming from openly available products?”. 

Indeed, your new workers are openly available, and you can swap them in and out off the shelf. There is no moat there. But the world model of your business can not be swapped in and out, or bought off-the-shelf. It is unique, and bespoke to your company: your operations, your values, your market, your customers, your product…

The catch is exactly the bespoke nature of it. It is not easy. One tempting shortcut is to shrink the world. For example, I could lay train tracks in my flat and cut my robot’s world model to “forwards or backwards”. It’s how most automation has worked to date, rigid assembly lines, with rule-based workflows, are powering most modern manufacturing. 

But I don’t want train tracks in my flat. And you don’t want your business on brittle rails from which an unexpected event (likely from a competitor) can derail it. Your world model is what lets you react, or not, to the next market shock. 

Indeed, how you build that world model is now a fundamental strategic bet in this new paradigm. Compare Waymo’s manual world-model building with HD images (“white-box”) to the approach of Tesla/Wayve to learn directly from raw data (“black box”): these choices are the heart of the strategic bets they are making against each other. The former gets you out of the gates quicker, the latter promises a far higher ceiling on capabilities if your pockets are deep enough to get there. I refer you to our blog “Agents give us the automation for the white-box world of decision making” for more on that here. Tesla’s Optimus is that black-box bet again: billions wagered that a humanoid can learn the world model of a home from fleet-scale data.

Take our pharmaceutical company again. True competitive advantage doesn’t come from having the latest AI model or agent – those are available to everyone. It comes from building a digital representation of how its own end-to-end clinical development portfolio works: its trials, resources, constraints, dependencies and objectives. With that in place, the questions we posed earlier become answerable, and the decisions far more consequential: what happens to cost, timelines and quality across my portfolio if we accelerate this trial, change its recruitment strategy, or reallocate resources from another?

The question has now changed

So, will Optimus be fetching my beer in just two years time, as Musk promises? Or will Moravec’s paradox bite again, and I’ll be left watching a moonwalking robot that somehow still can’t quite navigate my modest 50 m2 flat in London. Tesla’s bet lies on solving the world model problem of the home. Either way, that is very different to the world model of your business – and therein lies your moat.

For enterprises, then, the question has now changed from: “which latest agent are you using?” to: “how have you built the world model for your agents to operate in?”

Get that right, and your days of going back and forth to the proverbial fridge might finally be over…

Michael Davies
Lead Data Scientist
Michael is a Lead Data Scientist at Faculty, with over four years' experience applying machine learning, deep learning and computational simulation to model complex real-world systems. Michael holds a PhD in Computational Physics from UCL and the University of Cambridge, where his research used deep learning and simulation to uncover new insights into the atomic-scale structure of water and ice. His work has been published in leading journals including Science and PNAS, and featured in global media such as The New York Times and Der Spiegel. In 2023, he won the Institute of Physics Award for Computational Physics PhD in the UK and Ireland.