All guidesVisual Search

Optimising for Google Lens and Visual Search: What Actually Applies

Visual search matches pixels, not keywords — but the things that get you found are mostly the image SEO fundamentals plus two that are specific to Lens. Where the hype ends.

July 26, 20267 min read
A smartphone propped against a stack of books on a sunlit wooden floor, facing the legs of a wooden chair

Visual search attracts a particular kind of writing: large usage numbers, a claim that it changes everything, and a checklist that turns out to be ordinary image SEO with the word "Lens" added. It is worth separating what is genuinely different about visual matching from what is the same work you should already be doing.

The short version, as of July 2026: there is no Lens-specific tag, no visual-search schema, and no way to submit an image for matching. What exists are two composition and structure requirements that ordinary image SEO does not emphasise — and those are worth knowing if you sell physical objects.

In text search, someone types words and Google matches them against text. In visual search, someone points a camera at an object and Google matches the visual content against indexed images, then tries to identify the object and connect it to useful information.

That difference produces three consequences.

There is no query to target. You cannot pick a keyword, because the input is a photograph taken in someone's living room under their lighting at their angle. What you influence is whether your image is a good match candidate.

The subject must be identifiable in isolation. A busy lifestyle shot with five objects gives the matcher an ambiguous target. A single object, clearly separated from its background, in even light, is a strong candidate. This is the one composition rule that is genuinely specific to visual search — and it runs against the current fashion for atmospheric, cluttered product photography.

Matching and understanding are separate steps. Winning the pixel match gets you nothing if Google cannot then work out what the object is and where to send the person. That second step runs on exactly the signals ordinary image SEO produces: page text, structured data, alt text, filenames.

What actually applies, ranked

Lever Effect on visual search Same as ordinary image SEO?
Image is indexed and crawlable Prerequisite — nothing works without it Yes
Single clear subject, uncluttered background High — the one Lens-specific composition rule No
Valid Product structured data High for retail — makes results shoppable Yes, but underused
Multiple angles of the same object Moderate — more match candidates Partly
Adequate served resolution Moderate — very small images match poorly Yes
Descriptive filename and alt text Supports the understanding step Yes
Surrounding page text naming the object Supports the understanding step Yes
Image sitemap coverage Aids discovery at scale Yes
"Optimising for Lens" tags or markup None — does not exist

Two rows are worth dwelling on, because they are where visual search asks something different of you.

The composition requirement

If someone photographs a chair, Google is comparing that photograph against indexed images of chairs. Your image is a strong candidate if the chair is the unmistakable subject: centred or clearly dominant, separated from the background, lit evenly enough that its shape and material read clearly, shown from a common viewing angle.

Your image is a weak candidate if the chair is one of six objects in a styled room shot at an unusual angle in dramatic side light. That photograph may be far better marketing — but it is a worse match target.

The resolution is not to abandon lifestyle photography. It is to make sure that for every product, at least one image is a clean isolated shot, and to accept that this is the one that does the visual-search work while the styled ones do the persuading. Our product image SEO post covers the full shot priority list; the clean primary shot sits at the top of it for feed compliance too, so this costs nothing extra.

Multiple angles help for the same reason. Someone photographing a chair from behind will not match your front-only catalogue.

The structured-data requirement

Once an object is identified, what turns that into a useful result is information about the thing. For retail, that is Product structured data: name, availability, price, and image values that are absolute and crawlable.

Google's Product structured data documentation lists image as required, and it is the field most commonly broken — relative URLs, values pointing at a blocked CDN path, or a single image where multiple aspect ratios would serve better. Fixing that is worth more to visual search than any amount of Lens-specific speculation, and it also affects ordinary rich results, so the return is doubled.

If you do not sell products, this lever largely does not apply to you, which is part of the honest assessment below.

Where generated imagery fits — and does not

Visual search is the clearest case for a line most image guidance draws vaguely.

Lens is pointed at a real physical object. If your page shows a generated approximation of a product rather than the product, two things happen: the match is weaker, because the generated version differs in the details a matcher attends to; and if it does match, the visitor arrives at a page whose imagery does not depict what they photographed. That is a trust failure and, for retail, a returns problem.

So: photograph products. Use generated imagery for the roles where the subject is context rather than an identifiable object — category headers, editorial illustration, blog imagery, background plates. That is the same allocation our stock versus generation comparison arrives at from a cost angle, and the copyright post reaches from a legal one. Three different lines of reasoning, same conclusion.

The other visual surfaces, briefly

Google Lens dominates the conversation but it is not the only camera pointed at products, and the requirements are close enough that one body of work serves all of them.

Pinterest is the most commercially significant of the others for retail and home goods, and it rewards a different aspect ratio — vertical imagery performs far better in its feed than the square or landscape crops that suit Google. Its visual search works on pins, so being present in the ecosystem is the prerequisite rather than anything you do on your own site.

Bing Visual Search works on the same principles as Lens and draws on the same signals: indexed images, clear subjects, product markup. Nothing separate to do, though it is worth remembering that Bing's index is fed increasingly by content that AI assistants also draw on.

Retailer and marketplace search inside Amazon, eBay and similar platforms is its own discipline governed by each platform's rules, not by your site's markup. If a meaningful share of your sales happen there, image requirements on those listings deserve their own checklist.

The practical takeaway is that vertical crops are the only genuinely distinct asset requirement in this group. If you are producing a clean isolated product shot anyway, exporting a vertical crop of it costs nothing and covers the Pinterest case.

Where AI assistants fit into this

Worth addressing because it is the adjacent question people actually have. Assistants that read pages — and increasingly interpret images on them — rely on the same understanding layer as visual search: the text around the image, the structured data, the alt attribute, the filename.

There is no separate optimisation for it, and any claim otherwise deserves scepticism. An image with a descriptive filename, honest alt text and body copy discussing the subject is legible to a crawler, a visual matcher and a language model alike, because all three are trying to answer the same question about what the picture shows. That convergence is the strongest argument for treating the descriptive layer as foundational rather than as an SEO tactic.

The measurement problem, stated plainly

There is no Google Lens report. Search Console's Performance report filtered to Search type = Image is the nearest proxy, and it does not isolate camera-initiated queries. Lens-driven visits land in your analytics as ordinary organic traffic.

This matters because it makes visual search a poor candidate for its own project. You cannot demonstrate the return, and any agency proposing a "visual search optimisation package" with reporting attached is selling you image SEO with a premium label.

The correct posture is to treat visual search as a beneficiary of work you do anyway. Index your images, give every product one clean isolated shot, fix your Product schema, write real filenames and alt text. Visual search improves as a side effect, and every one of those items has an independently measurable payoff.

The fundamentals still decide it

Strip out the composition rule and the schema requirement and what remains is the standard list — because the understanding step runs entirely on it. An image that wins a pixel match but sits on a page with IMG_2291.jpg, an empty alt attribute and no text naming the object gives Google nothing to work with.

Which is the same conclusion as everywhere else in image SEO: the delivery layer is configured once, and the descriptive layer is a per-file decision nobody automates for you. Our checklist orders the whole list, and the alt text and filename guides cover the two items that carry the most weight.

For the images you generate rather than photograph, SEOpix writes the filename, alt text and EXIF block at generation time so the descriptive layer exists before the file reaches your site. For your product shots, photograph the product — that is the one place in this whole subject where there is no shortcut.

Do the two things that actually apply: confirm every product has one clean isolated image, and validate your Product schema image values. Then stop reading about visual search and go fix your alt text. Ten images a month are free if the generation side is where your bottleneck is, and the pricing tiers cover real volume.

Frequently asked questions

Can you optimise for Google Lens directly?+

Not with a dedicated set of tags or a Lens-specific markup. Lens matches a camera image against indexed visual content, so being findable requires your images to be indexed, clearly show a single identifiable subject, and sit on pages with structured data and text that let Google connect the visual match to a product or entity. The levers are ordinary image SEO plus composition discipline.

Does Lens use alt text and filenames?+

Indirectly. The visual match is made on image content, but once matched, Google needs to understand what the thing is and what page to send someone to — and that understanding comes from the surrounding text, structured data, filenames and alt text. Good metadata does not win the match; it makes the match useful.

Which industries see the most visual search traffic?+

Anything where people identify objects in the physical world: retail and fashion, furniture and home goods, plants and gardening, auto parts, tools and hardware, and food. Service businesses see far less, since there is rarely a physical object to point a camera at, though parts and equipment photography can still surface.

Do I need special structured data for visual search?+

No new type, but existing Product structured data does a lot of work here. When Lens identifies an object, availability and pricing information from Product markup is what lets the result become a shoppable one. If you sell physical goods, valid Product schema with correct image values is the highest-leverage item.

Are AI-generated images matched by Lens?+

They are indexed and matched like any other image, with an important caveat: Lens is being pointed at a real physical object, so a generated approximation of a product will produce weaker matches than a photograph of the actual item — and a mismatch between what the camera sees and what your page sells is a bad outcome for everyone. Generated imagery belongs in context and editorial roles, not product identification.

How do I measure visual search performance?+

There is no separate Lens report. Search Console's Performance report with Search type set to Image is the closest available proxy, and Lens-driven visits arrive as ordinary organic traffic. This measurement gap is real and a reason to treat visual search as a beneficiary of image SEO work rather than a channel with its own budget.

Does image resolution affect visual matching?+

Up to a point. Very small or heavily compressed images give the matcher less to work with, so an image that renders at 200 pixels wide is a poor candidate. Beyond a reasonable display size the returns flatten quickly, and oversized files cost page speed — so serve appropriate dimensions rather than maximum ones.

Let SEOpix handle the metadata

Filenames, alt text, EXIF fields and GPS coordinates written automatically as each image is generated. Start with 10 free images a month — no credit card required.

Keep reading