Cover: How AI Keeps Your Child's Face Consistent | MyOwnChildbook
24 August 2026

How AI Keeps Your Child's Face Consistent | MyOwnChildbook

Generate the same text prompt twice in an AI image generator and you get two different faces. That is not a glitch, it is how the model works: every generation starts from random noise and follows its own path to the final image. For a single illustration, that is harmless. For a story spanning eleven pages, where the same child has to reappear again and again, it is exactly the core of the problem.

Why consistency is not a built-in property

Text-to-image models have no memory between generations. Every image is computed independently, even when the prompt is nearly identical to the previous one. Researchers now treat this as a distinct, formally named problem: “consistent character generation” cannot be solved simply by training a bigger model. In “The Chosen One: Consistent Characters in Text-to-Image Diffusion Models” (Avrahami et al., 2023), the authors show that a standard model, without any additional technique, produces a slightly different face, hairstyle or build for every new scene, even when the text prompt stays word-for-word identical.

Two families of technical solutions

Research into this problem has broadly produced two approaches. The first, described in that same paper, generates a large number of variants from the same description and then clusters the results that resemble each other most closely into one consistent identity, without needing a reference photo at all.

The second approach works the other way round: give the model an actual reference image alongside the text prompt, so it can literally “anchor” itself to visual features. This is called reference-conditioning, and it is set out in papers such as “IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models” (Ye et al., 2023). The model learns not only from text but also from a separately supplied image, via a dedicated attention layer that focuses specifically on that image.

Colourful lines of programming code on a screen, symbolising the technique behind AI illustrations

Our approach: one anchor image, sent along eleven times

At MyOwnChildbook, we take the second route. As soon as the first illustration in a book is approved, that image becomes the anchor: it is sent along as a reference with every subsequent page prompt, together with a detailed written character description (hair colour, hair type, skin tone, clothing item). Edwin, the data engineer who built the pipeline: “Without that anchor, you might get a boy with brown hair on page 7 when he still had red hair on page 2. With the anchor, the model stays close to what has already been approved.”

It is not entirely flawless. On roughly one in eight pages, an automatic check detects noticeable character drift, such as a slightly different hair colour or a shifted face shape, and the page is automatically regenerated before you ever see the book. Our detailed walkthrough of the full gpt-image-2 pipeline shows how this step fits into the whole process, from noise to pixel.

When even this approach is not perfect

It matters to be honest about this: reference-conditioning delivers recognisability, not a pixel-for-pixel identical copy. The red hair, the freckles, the general face shape stay stable enough across eleven pages to trigger that “that’s me” reaction, but the illustration on page 3 is never an exact digital clone of the one on page 9, and that was never the goal.

Anyone looking for a single hyper-realistic portrait where every facial feature matches a photograph exactly, for example to frame on a wall, is better served by a printed photograph or a photo book than by an illustrated story. And a source photo that is blurry, dark or low resolution simply gives the model less to anchor to, so the sharper the uploaded photo, the more stable the consistency across pages tends to be. Our explanation of embeddings and style recognition goes deeper into what a model can and cannot extract from a single image.

Family sharing a moment of recognition together on the sofa

The alternative: a human who keeps comparing

There is a non-AI route that solves this problem in a structurally different way: hiring a freelance children’s book illustrator, for example through a platform like Etsy or Fiverr, who lays each page by hand next to the same reference photo. A human deliberately compares against earlier pages and corrects themselves; that costs no computing power but does cost time, often two to six weeks per book, and a price considerably higher than an automated pipeline. For anyone with no deadline pressure who loves a fully unique, hand-drawn result, that is a fair alternative.

A craftsperson working by hand on a creative piece, symbolising a human illustrator

Edwin’s perspective

“I saw this problem crop up again and again in my own tests before we built the anchor mechanism,” Edwin says. “That is exactly the level of detail at which consistency needs to work, not photographically perfect, but recognisable on every single page.”

Character consistency is not a visual coincidence. It is an explicit technical choice, an anchor image, a reference input, an automatic check, that decides whether a child still recognises themselves on page 9 as well as they do on the cover. Curious what that looks like in practice? Take a look at the illustration style examples before you start your own.

👉 Create your child’s own book

Your child as the hero of the story

Create the complete book in a few minutes and read every page first. You only pay once you have seen the whole thing.

  • Read the entire book before you pay
  • Your child on every page, drawn from your photo
  • Premium hardcover, usually delivered within 6 working days
Start for free