Skip to main content
ai search aeo07 Sept 2026·7 min read

How a machine chooses the right photos for an article

Dragoș-Adrian BuhoiuFounder · Digital Ecosystem Architect
How a machine chooses the right photos for an article
FEATURED.IMG
How a machine chooses the right photos for an article

The real photo always wins; generation is the fallback, anchored in the client's photos and declared as such. What we measured — and the rule we never touch.

When our content engine publishes an article on a client's site, the cover photo comes in this order: a REAL photograph from their site, if one fits the topic; if not, a generated image, anchored in their real work photos and declared as generated in the metadata; and never an image that reproduces the premises, the people or the logo. The order is not an aesthetic preference — it is the result of measurements that contradicted us three times. Below, the whole chain, the numbers and the rule we never touch.

The problem it solves

An article without an image is an article nobody shares. An article with a stock image is an article everybody has already seen. And an article with a randomly generated image — a generic office, a tradesman with three hands — betrays, in two seconds, that nobody was there.

The client almost always has real photos: on the site, in the gallery, on service pages. Few, sometimes badly taken, but theirs. The question is not "how do we generate a beautiful image", but "how does a machine use the photos the client already has, without lying".

The chain, step by step

The inventory:Which photos actually exist

At the first reading of the site, the machine collects all public images — from pages, from the gallery, from articles — with their address, size and the context they appear in. Only what the client published themselves; nothing from private accounts or folders. The inventory is surprisingly uneven: one services company had 7-8 photographs in total; one guesthouse, 126.

Triage:What is visible in each photo

Every image gets a measured description — what is VISIBLE, not what we would like it to be — and a verdict on three questions: does it show work, tools, materials, products or craft? does it show an identifiable place, a façade, a room? does it show a person with a visible face, or a logo?

This is where the first surprise came. Many good photos were "spoiled" by one detail: a corner of a logo, a sign on the wall, a face in the background. We added cropping: the machine cuts the problem area and re-judges the remaining pixels. The result, measured on two real sites: at the services company, from 0 usable photos out of 7 to 5 out of 8; at the guesthouse, from 9% to 20% of the inventory. Cropping recovered more than any prompt.

The real photo first — the decision that contradicted our plan

The initial plan was: real photos are references for generation. The numbers said otherwise. The guesthouse had 72 photos rejected as references — they all showed the property, so they were forbidden for generation — and was left with generic covers. Exactly the niche where the photo matters most was getting the weakest images.

The way out: the photos rejected as references are rejected because the model would reproduce them. As photos, they are perfect — they are the client's own site. So, wherever a real photograph fits the topic and is large enough (we measured the width threshold on the two sites: 768px phone photos are weak as covers, 1,400px ones excellent), it BECOMES the cover. One hundred percent true, zero model cost, and — importantly — no "generated" label: it is not generated, and labelling it so would be the inverse lie.

The only exclusion that remains even for real photos: faces. An identifiable person on an article's cover requires a consent we cannot verify.

Generation, only when there is no photo — and anchored

When the inventory has no fitting photograph, the machine generates one. But not from bare text: it selects, by closeness of meaning to the article's title, the client's most fitting WORK photos as references — their overalls, their tools, their materials — and generates around them.

The number that decided how many references: three. With one, the clothes came out wrong; with three, the company's equipment came out right; with four, the model mixed indoor equipment into an outdoor scene. The meaning-closeness threshold is calibrated on fitting topics versus topics from another niche — the exact figures are ours and get recalibrated with every new model; the principle is public: below the threshold, no reference is better than a wrong one.

The rule we never touch

A reference may show work, tools, materials, products, craft. It may not show the identifiable premises, rooms, façades, people with visible faces or logos. The reason is measured, not principled: given the photo of a guesthouse's property, the model reproduced the same property — with invented icicles and with the carved balcony posts replaced by gaps. The visual judge answered two questions: "is it the same place?" — YES; "would it fool a client?" — YES. An image that shows the client's house other than it is, is not a cover — it is false advertising.

The gate and the declaration

No image reaches the site without a pixel-level check: a second model looks at the result and answers concrete questions (does it fit the topic? does it contain invented text? impossible elements?). An "ok" from a prompt is not a measurement — we learned that the hard way, with three "good" covers that were broken.

And generated images carry in their metadata the IPTC "Digital Source Type" label — the official vocabulary Google reads — as trainedAlgorithmicMedia or compositeWithTrainedAlgorithmicMedia, as the AI Act's transparency (Art. 50) requires. Real photos do not carry the label, because it would not be true.

What it costs and what it does not

A generated cover starts at ~1.8 MB as PNG and lands on the site at ~140 KB as JPEG, compressed without visible loss. A real cover costs zero at the model and nothing worth mentioning in compression. The real cost of the chain is in judgement — triage, cropping, verification — not in generation. And that is exactly the cost we do not cut: without it, the machine would publish beautiful, false photos.

How it connects to the rest: the image enters the article with a real description in the alt, and the article declares its author, sources and structure in the schema — the schema markup guide. On generated content and what Google expects of it: information gain, not volume and what the detectors actually detect.

What this article does NOT say

  • We do not publish the numeric thresholds for meaning-closeness and cropping — they get recalibrated with every model and would mean nothing in another implementation. The principles, yes.
  • We do not say generation is "as good" as the real photo. It is not; that is why it is the fallback.
  • The numbers come from two real sites (a services company and a guesthouse), measured in August 2026. Other niches may have other proportions; the chain is the same.

FAQ.PROTOCOL

Frequently Asked Questions

Because "beautiful and false" is worse than "real and imperfect". The reader recognises the generic image in two seconds, and the client sees themselves represented by something that is not theirs. The real photo always wins when it exists; generation is only when it does not.
As references, yes — they show exactly your equipment, materials and craft. As covers, those under ~1,000px wide come out poorly; we measured the threshold on real sites. Send them anyway: cropping recovers more than you think.
No — the rule forbids identifiable premises, faces and logos as references, for a measured reason: the model reproduced a client's property with invented details. The real photo of the premises can be a cover (it is true); its generated reproduction cannot.
Yes, in the IPTC metadata ("Digital Source Type"), in the official vocabulary Google reads, in line with the transparency the AI Act requires. Real photos are NOT marked — it would not be true.
A second model, at pixel level, with concrete questions — not an "ok" from a prompt. We had three "good" covers that were broken; since then, verification is on the artefact, not on trust. What fails gets regenerated or replaced with a real photo.
INITIATE.SEQUENCE
// 01_OF_01
// Next Step

Let's build something remarkable.

30-min discovery call — no cost, no pitch. We audit your digital architecture and deliver a clear operational plan.

  1. 01Short message with your business context
  2. 02Reply within 24h with a discovery-call proposal
  3. 03Operational plan + scope recommendation
Contact usor browse resources
24h replyZero spamDirect with the founder

Digital engineering notes

One measurement on a real site, every Tuesday. Numbers, method, and what does not flatter us.

Unsubscribe anytime. Privacy

Related Articles

AUDIT · 24–48H

See exactly where you stand — no promises.

We run your site’s real data — speed, indexation, visibility in Google and AI search — and send you a concrete report.

✓ YES

  • Concrete, prioritised issues
  • Manually verified, not a tool
  • No card, no strings

✕ NO

  • “#1 guaranteed in 7 days”
  • Generic AI-spat report
  • Aggressive sales call
Request the audit →
WhatsApp direct