Fei-Fei Li is walking through the park beside her company’s offices when she explains why she prefers meetings in motion. ‘We sit too much,’ she says, and then, almost without pausing, the conversation turns to artificial intelligence, because that is the only thing she talks about. For anyone trying to understand where AI goes after ChatGPT, Li is the right person to be walking with. The scientist who gave AI its eyes is now trying to give it a world to live in. The question is whether the world is ready to follow her there.
Li grew up in Chengdu in the 1980s and 1990s, chasing bugs with her father, sketching in the mountains, and reading voraciously. She cut her hair short, loved aerospace and physics, and, by her own account, had to be at least a little rebellious just to stake out an identity. When her family emigrated to New Jersey in the 1990s, she arrived as a teenager who did not know the language, finished high school, and then put herself through Princeton studying physics while running the family dry-cleaning business on weekends. Her doctorate came later, at Caltech, where she became a computer scientist and began asking the question that would define her career: how does AI see, and how is that different from how a person sees?
The catalog that sparked a revolution
In 2006, when the broader research world was barely paying attention to data, Li started building ImageNet, a catalog of 14 million pictures across more than 21,000 categories, the largest assembled at that time. The logic was simple, even if the execution was not: algorithms needed data to get smarter, and nobody was giving them enough of it. She turned ImageNet into a competition in 2010. In 2012, a team from the University of Toronto entered with an algorithm called AlexNet, powered by NVIDIA graphics cards, and the result rewrote the field. Massive data, neural networks, and GPU computing became, as Li describes it, ‘the golden recipe for modern AI.’ Standing in front of a photograph from roughly a hundred years ago, the kind of industrial skyline image that ImageNet would have catalogued, she is direct about what the moment meant: ‘It’s not the eight of us. It’s the entire humanity. That’s really the main story here.’
One thing on the table
While the rest of the AI industry sprints to build larger language models, Li launched her own startup in 2024. World Labs is betting on what she calls world models, a technology that aims to do something language alone cannot. ‘Can words put down fires?’ she asks. ‘Can words cook an omelet?’ The distinction matters to her. Language models predict the next word in a sentence. World models aim to predict what happens next in the physical world, capturing geometry, physics, and spatial relationships so that machines can plan and act, not just describe.
World Labs has already shipped its first product, a platform called Marble, which generates explorable, editable 3D environments from a single image or text prompt. Movie productions are using it to shoot actors in any environment. Game developers are cutting the time and resources required to build worlds. NVIDIA is using Marble environments to help train robots. The company has roughly 50 people and has raised a billion dollars in investment, with total investment in world models across the industry reaching $3 billion and growing. Li describes the competitive pressure without flinching. ‘I’m paranoid every day,’ she says. ‘But am I paralyzed? Absolutely not. We have one thing on the table, and we’re focusing on that thing.’
One of her earliest collaborators, who began his PhD with Li at Stanford in 2012 and now works at World Labs, frames the appeal this way: at a large company your hands often cannot change the trajectory, but at 50 people they can.
Li is honest that the technology is early, perhaps as early as 2019 was for chatbots. The field has not yet agreed on how to build world models, and getting them out of the demo phase will take more time and capital. When asked whether a breakthrough moment is coming, her answer is characteristically precise: ‘If we do it right, yes.’
A name she did not claim, and will not reject
For all her accomplishments, the title ‘Godmother of AI’ did not come from Li herself. When it was first put to her, she was taken aback. ‘I don’t naturally think about myself as godmother of anything,’ she says. ‘That’s not my personality. I focus on work.’ But she chose not to deflect it, either. ‘If I rejected that on that spot, we once again would be in a situation where women just don’t get recognized the same way as men.’ What she actually wants, she makes clear, is more women being called godmother of whatever they created.
The responsibility she feels extends well beyond her startup. She has advised US presidents and the United Nations on AI policy, and she carries a consistent message into those rooms: root policy in science, not science fiction, and resist the temptation toward either utopia or doomsday. She is critical of leaders who speak as though they alone know what is best for humanity. ‘I think it’s dangerous for any individual to think they know better than everybody else,’ she says. Her framing for AI is insistently human: ‘There’s nothing artificial about artificial intelligence. It’s inspired by people. It’s created by people. And most importantly, it has an impact on people.’
Still dreaming about what is out there
On a bench at the edge of the park, the conversation edges briefly into childhood. Li mentions that as a kid she was fascinated by UFOs, that she has read Drake’s equation computing the probability of extraterrestrial life, and that her all-time favorite film is ‘Contact.’ The detail sits there quietly, a reminder that the scientist who built the world’s largest image catalog and is now training AI to navigate physical space started as a curious kid looking up.
Li gave AI its eyes with ImageNet. Now, with World Labs, she is building it something larger: a world with structure, physics, and consequence. Whether that bet pays off is still an open question, but the woman asking it has been right about the next frontier before.


