Labelling AI-Generated Images: Provenance Metadata and What Actually Reaches the User
IPTC digital source type, C2PA Content Credentials and invisible watermarks — what each one records, what survives a CMS upload, and what search engines currently do with any of it.

Two developments arrived at roughly the same time, and they are frequently discussed as one thing. The first is technical: standardised ways to record how an image was made, readable by machines. The second is regulatory: rules in several jurisdictions requiring that synthetic media be identifiable as such. They interact, but they are not the same problem and they do not have the same answer.
This piece covers what the technical mechanisms actually record, what happens to them in a normal publishing workflow, and what search engines do with any of it as of August 2026.
Three mechanisms, doing three different jobs
| Mechanism | What it is | Tamper-evident | Where it lives |
|---|---|---|---|
| IPTC digital source type | A metadata field with a controlled vocabulary for how media was produced | No | Inside the file's IPTC block |
| C2PA Content Credentials | A cryptographically signed manifest of creation and edit history | Yes | Attached to the file, or a sidecar |
| Invisible watermarking | A signal embedded in the pixels themselves | Partially | The image data, not the metadata |
The distinction that matters: the first two are metadata, which means they can be removed by anything that rewrites the file. The third is in the pixels, so it survives a metadata strip — but it can only be read by whoever holds the corresponding detector, which in practice means the organisation that applied it.
None of the three is a ranking signal. All three are disclosure mechanisms.
IPTC digital source type
The IPTC Photo Metadata Standard, maintained by the International Press Telecommunications Council, has carried structured photo metadata since long before generative imagery existed — creator, credit line, copyright notice, description. Its Digital Source Type property records how the media was produced, drawing on a published controlled vocabulary that distinguishes straight photographic capture from composites and from media created by a trained algorithm.
This is the least glamorous mechanism and the most immediately practical one, for a simple reason: it sits in the same metadata block your existing tools already read and write. If your workflow already writes a copyright notice into the file — and it should — writing a digital source type is the same operation into a neighbouring field. The guide to what EXIF and IPTC metadata actually does for SEO covers how that block behaves in practice and which fields survive a typical upload.
C2PA and Content Credentials
The Coalition for Content Provenance and Authenticity publishes a specification for signed provenance manifests; Content Credentials is the user-facing branding for it. A conforming file carries a manifest saying what produced it, what edited it, and when — each step signed, so alteration is detectable rather than merely undeclared.
The advantage over a plain metadata field is exactly that signature. Anyone can type a digital source type value into a file, truthfully or otherwise. A C2PA manifest that has been tampered with fails validation.
The disadvantage is ecosystem coverage. A signed manifest is only meaningful where something checks it, and provenance survives only where every tool in the chain preserves it. Adoption among generation platforms and major editors has grown substantially, but a manifest that passes through one non-conforming resize step is gone.
Invisible watermarking
Several model providers embed a signal directly into generated pixels — a modification imperceptible to a viewer but detectable by a matching classifier. Because it lives in the image data, it survives cropping, format conversion and metadata stripping to a useful degree.
The limitation is access. Detection generally requires the provider's own tooling, so watermarking answers "is this ours?" better than it answers "what is this?" for a third party. It is a useful backstop, not a disclosure mechanism you can rely on to inform your users.
What search engines do with it
Less than the coverage implies, and that is the honest summary.
Google has built provenance data into informational surfaces — the "About this image" panel that can tell a user where an image has appeared and, where the data exists, how it was made. That is an information feature, not a ranking one. There is no published evidence that a declared digital source type changes an image's position in Google Images, and no reason from Google's stated principles to expect it to.
What does matter is the deception question, and it predates generative imagery entirely. An image implying something untrue about a business — premises that do not exist, a product that looks nothing like what ships, a photograph of a team that was never assembled — is a misrepresentation problem whether it was generated, composited or staged. Search engines are one of several parties who will eventually take a view on that, and usually not the one whose view costs you most.
The recurring reasonable question — whether to use generated imagery at all for a given slot — is a judgement about fitness rather than about labelling. Our comparison of stock photography and AI-generated images works through which slots each option genuinely suits, and the answer is neither "always" nor "never".
The regulatory layer, briefly and cautiously
The EU AI Act includes transparency obligations for generated and manipulated content, with the relevant provisions beginning to apply in August 2026. The obligations are directed at synthetic media capable of deceiving — in particular realistic depictions of real people, events or places — and at making such content identifiable in a machine-readable way. Several other jurisdictions have adopted or proposed measures with similar aims and different scope.
Two observations that are safe to make and useful:
Scope is about deception potential, not production method. The concern driving these rules is synthetic media presented as a record of something real. An illustrative image on a service page is a different object from a fabricated photograph of a named person.
Machine-readable is the recurring requirement. Where disclosure obligations exist, they tend to ask that the disclosure be present in a form a system can read, not only in visible small print. That is precisely what the metadata mechanisms above provide, which is a good reason to write them even where you are confident you are out of scope.
This is not legal advice, the boundaries are genuinely unsettled, and if your imagery depicts identifiable people or real events the question deserves a lawyer rather than a blog post.
A workable policy for an ordinary business site
Most sites do not need a provenance programme. They need a defensible default and a rule for the edge cases.
Default: write the digital source type into every generated image at creation, alongside the attribution fields — creator, credit, copyright notice — you should be writing anyway. Cost: nothing, if the tool does it. Benefit: an accurate machine-readable record you never have to reconstruct later.
Preserve where you control the pipeline. Check your CMS. Upload a generated image, publish it, download the published version, and inspect its metadata. If the fields are gone, your resize step is stripping them, and that is usually a configuration option rather than a rewrite.
Do not label what you do not know. An unfilled field is neutral. A field asserting straight photographic capture on an image of unknown origin is a false statement in a place designed to be trusted.
Draw the line at depiction, not decoration. Generated imagery illustrating a concept is unremarkable. Generated imagery that a reasonable visitor would read as documentary evidence — your premises, your staff, your finished work — is where accuracy stops being a metadata question. If you want images of the actual thing, photograph the actual thing.
Keep rights and origin separate. Provenance answers how an image was made. Licensing answers who may use it. They are different fields and different questions; the copyright and commercial use position on AI images covers the second, and the structured data that declares licence terms is a third thing again.
Answering the question when a client asks
Agencies and freelancers get a version of this question regularly, usually phrased as "are these real photos?" It deserves a direct answer, and the answer is easier if the policy above is already in place.
A serviceable form of it: the photographs of your premises, your team and your completed work are real and were taken on site. The contextual and illustrative images were generated, they carry metadata recording that, and none of them claims to depict anything specific about your business.
That answer works because it is true, because it distinguishes the two categories the client actually cares about, and because it does not require the client to have an opinion about generative imagery in the abstract. The version that causes problems is the vague one — "we source them from a library" — which is technically true of stock as well and postpones a conversation that will be less comfortable later.
Two follow-ups worth pre-empting. Clients ask whether generated images will hurt their rankings: no evidence supports that, and the answer above is a good moment to say so plainly. And they ask who owns the images, which is a licensing question rather than a provenance one, and is answered by the contract rather than by the file.
What to check this week
A short audit, in the order that finds problems fastest:
- Download three published images from your live site and inspect their metadata. This tells you immediately whether your pipeline preserves anything.
- If the fields are being stripped, find the resize step and check for a preserve-metadata option before rewriting anything.
- Confirm your generated images carry a correct digital source type and accurate attribution fields at creation.
- Review any image on the site that a visitor could reasonably read as documentary — premises, team, completed work — and decide whether it should be a photograph.
- Leave the rest alone. Decorative and illustrative imagery does not need a provenance programme.
SEOpix writes IPTC attribution and copyright fields into every file it generates, so the metadata half of step three happens at creation rather than in a post-processing pass. See what is written at each plan level, or read the FAQ for the specifics of which fields land in the file.
Frequently asked questions
Do I have to disclose that a website image was AI-generated?+
It depends on jurisdiction and on what the image is doing. Transparency obligations in the EU AI Act, which began applying to generated content in August 2026, are aimed principally at synthetic media that could deceive — realistic depictions of real people, events or places — rather than at every decorative illustration on a business website. This is not legal advice and the boundaries are genuinely unsettled, so if your images depict identifiable people or events, take advice rather than a blog post's word for it.
Does Google penalise AI-generated images?+
There is no evidence of a penalty for generated imagery as such. Google's stated position on content generally is that it rewards helpfulness and originality rather than production method. What does cause problems is generated imagery used deceptively — a fake product shot, a fabricated photograph of premises that do not exist — which is a misrepresentation issue rather than an image-format one.
What is the IPTC digital source type field?+
It is a standard metadata field that records how a piece of media was produced, using a controlled vocabulary that includes values for straight photography, composites and algorithmically generated media. It is the most widely readable way to declare an image's origin because it uses the same IPTC block that news and stock workflows have used for decades, so tools that already read photo metadata can read it.
What are Content Credentials?+
Content Credentials are the user-facing name for provenance data built on the C2PA specification — a cryptographically signed manifest attached to a file recording what created it and what edited it since. Unlike a plain metadata field, the signature makes tampering detectable, and the manifest can carry a chain of edits rather than a single claim.
Will provenance metadata survive my CMS?+
Often not. Many content management systems and image pipelines strip or rewrite metadata during resizing, and most social platforms remove it on upload. Treat provenance metadata as durable where you control the pipeline and fragile everywhere else — test by downloading a published image from your live site and inspecting it, which takes two minutes and is the only reliable answer for your specific stack.
Does labelling an image as AI-generated hurt its performance in search?+
No measurable effect has been demonstrated, and search engines currently use provenance data for informational panels rather than ranking. The practical risk of labelling is reputational rather than algorithmic, and it runs the other way for most businesses — being straightforwardly accurate about your imagery is safer than being discovered to have implied otherwise.
Should I add provenance metadata to images I did not generate?+
Only if you know the answer. A digital source type value asserting straight photography on an image whose origin you cannot verify is a false claim in a machine-readable field, which is worse than leaving it absent. When origin is unknown, leave the field out and fill in the attribution fields you can substantiate.
Let SEOpix handle the metadata
Filenames, alt text, EXIF fields and GPS coordinates written automatically as each image is generated. Start with 10 free images a month — no credit card required.


