What good looks like: showing kids lived culture, not postcards
Updated: Jul 24
I typed a plain prompt into a new image model: children greeting each other at school in Tokyo. No cultural coaching. No steering. I pressed go and waited.
The result was not a cherry blossom postcard. It was not a row of children bowing in costume. It was a moment. The right uniforms for a contemporary Japanese school. The right architecture. Kids doing what kids do, arriving at a place they know. A lived morning, not a souvenir.
That stopped me. Because it told me something had changed about what AI can do, and what we now have to hold it to.
Why the postcard version is the problem
When my children open Unomundi and explore Japan, they should see Japan as Japanese children live it, not Japan as a foreign photographer imagined it thirty years ago. That distinction matters more than it might sound.
Children are pattern-matchers. Show them enough postcards, and they build a mental model of a place from those postcards. Show them enough lived moments and they build something much closer to truth. The image on screen is not decoration. It is part of what a child learns to believe about the world.
That is the standard we hold ourselves to. Not beautiful. Not diverse-looking. Lived and contemporary. A school morning that could have been taken this week.

What the AI actually did differently
Earlier models flattened culture. Ask for a school scene and you might get something geographically generic, or worse, something that leaned on the most recognisable visual cliche of that place. Postcard-brain, baked in.
What I found in recent testing was different. The model appeared to reason about the scene rather than just render it. It reached for contemporary specifics: architecture, clothing, context. Not perfectly, and not every time. But consistently enough to matter.
Character consistency, which had always been a problem with image generation, also improved sharply. Our guide, Una, has to look like herself across dozens of scenes and storylines. In more than fifty consecutive generations, she held: her height, her features, her presence. Scene consistency across a whole storyline went from painstaking to manageable. That is a real change in what production looks like.
The pace changed, too. We went from three iterations per image down to two, on average. Less than half the human rework. Mid-production, when you are building a great deal of content, that is not a small thing. It is the difference between sustainable and not.
Does AI get culture right on its own?
No. That is the honest answer. AI-generated images are not free of cultural bias or stereotyping, not completely, not all the time. The leap was real. The gap is not closed.
Human review stays in our pipeline. Always. Every image that goes into Unomundi is checked by a person who knows what they are looking for. Not because the AI is broken, but because this is not the kind of thing you automate and walk away from. The stakes are a child's mental model of the world. We do not outsource the final judgment on that.
What shifted is the starting point. We used to begin from a souvenir-shop baseline and push hard toward something true. Now we begin closer to true and push to make it excellent. That changes what the human reviewer is doing. Less correcting. More refining.
What you actually see on screen, and why it matters
As a parent, you are probably not thinking about image pipelines when you hand your child a phone. You are thinking about whether what they are looking at is good for them.
So here is the practical version. When your child explores Nigeria in Unomundi, they should see Lagos as a city, not as a village from a nature documentary. When they explore Japan, they should see a school morning, not a temple tour. When they explore Brazil, they should see a neighbourhood, not a carnival float.
Those choices are not automatic. They are deliberate. They require a standard, a pipeline, and a human saying this one, yes, and this one, not yet.
We build stories that help children understand each other before they judge each other. That only works if what children are seeing on screen is real. Not a filtered version of real. Not a safe version of real. The lived, contemporary thing.
The standard, after the test
The Tokyo school prompt was a test. The model passed. But passing a test is not the same as meeting a standard.
Unomundi is a cultural-discovery app where children aged 6 to 12 explore countries and cultures through stories, games, and conversations with Una, a warm AI guide. Every image Una appears in, every scene a child moves through, goes through that standard. The AI brings us closer. The human check is what makes it ours.
The postcard was never good enough. Now we have fewer excuses than ever to settle for it.
FAQs
How does Unomundi make sure images of other cultures are accurate and not stereotypical?
Every image goes through human review before it reaches the app. AI generation has improved enough to start from a contemporary, lived baseline, but the final check is always made by a person.
Can AI image tools show real, everyday culture rather than tourist clichés?
Recent models have improved significantly, reasoning about contemporary specifics like school architecture and clothing rather than defaulting to the most recognisable postcard version of a place. But they are not bias-free, and human oversight is still essential.
What is Unomundi?
Unomundi is a cultural-discovery app where children aged 6 to 12 explore countries and cultures through stories, games, and conversations with Una, a warm AI guide.