Google was questioned on Bluesky about hreflang tags potentially being ignored, showing a Belgian-French page in France. John responded, “I suspect this is a “same language” case where our systems just try to simplify things for sites.” I have addressed this in previous podcasts and on Hreflang Builder the challenge of same-language duplicates that seems to align with how Google processes pages.
The original poster did not provide the web page, so we cannot compare, but based on my experience, especially with that market combination, the pages would be near-identical. When you launch identical product pages for different French-speaking regions, such as France and Belgium, with the same language, text, images, and price, Google handles these pages depending on several technical and policy factors. Here are the key facts, supported by patents, documentation, and expert commentary:
Launching localized websites for distinct markets like France and Belgium is a critical international SEO strategy. However, when those pages are identical in content, structure, language, and pricing, Google’s systems must decide whether to index and rank them independently or group them as duplicates and select a canonical—even across different domains.
Duplicate Content Detection and Canonicalization
Google has patented systems for detecting duplicate and near-duplicate content. When its crawler encounters pages with “substantially identical content” but different URLs, it uses a process (including content fingerprinting) to group these into an “equivalence class” and then selects one as the canonical version. The canonical URL may not always be what you want them to select. Multiple factors, such as site structure, authority, and possibly user signals, may influence the decision. Suppose Google determines that two pages are exact or near-exact duplicates. In that case, it will typically index and rank only one version, filtering the others out of search results to avoid redundancy. This will often trigger a Duplicate Canonical error in the Search Console.
Supporting Patents and Research
- US Patent 8,543,398: “Identifying duplicate and near-duplicate files” describes using content fingerprinting and shingling techniques (breaking text into word sequences) to determine similarity between documents.
- US Patent 8,341,091: Discusses how URLs with substantially similar content are assigned into equivalence classes, and only the strongest (most relevant/trusted) is surfaced in search results.
- Google Research Paper: “Near Duplicate Detection in Web Search” (Google, 2005) explains scalable methods Google uses to identify and de-duplicate content at the web scale, regardless of domain.
Takeaway: Google’s systems don’t treat domain boundaries as sacred; identical content across example.fr and example.be is at risk of being collapsed into a single representation if it appears redundant.
Factors Influencing Google’s Choice of a Canonical URL
Google has stated it uses a complex system relying on approximately 40 different signals to determine which URL should be treated as the canonical (preferred) version when it encounters duplicate or near-duplicate content. Unlike what others have written, in this cross-market and sale language scenario, some of the strongest signals, like canonical tags, redirects, sitemaps, and internal linking structures, are irrelevant to identifying the canonical version as these signals apply to the same website, not cross-markets. This essentially defaults to hreflang, geotargeting, and relevance signals as the factors used to determine which page to use.
In a 2023 Google Search Central Office Hours, John Mueller said:
“If the content is completely the same, and we can’t tell any difference, then for simplicity and user experience we may just show one version—even if hreflang is present.”
John again stated that hreflang alone is not a guarantee of indexing when Google believes content is duplicated and not meaningfully localized in this BlueSky thread.

Takeaways:
- URLs with identical content in the same language and different markets are often grouped, and one is selected for search—typically the more authoritative or better-linked version. This requires market relevance and geographical targeting.
hreflangis a suggestion, not a directive. It influences serving, but doesn’t override duplicate detection if two pages are effectively the same. This is critical and explained below in Google’s process chain.- Canonical tags are irrelevant across different domains and do not establish cross-site canonicalization unless explicitly confirmed via cross-domain
rel=canonical(rare and discouraged in this context).
Geographical Targeting Challenge
Unlike language detection, determining a website’s target geographic market is fraught with ambiguity and technical challenges. Search engines must assemble a complex puzzle of often contradictory signals to make educated guesses about market targeting. This process occurs at multiple stages of the search engine’s processing pipeline, creating challenges as different signals may be evaluated at various points in the algorithm.
I wrote a fairly extensive article on this challenge of non-existent and conflicting geographic signals on most multinational websites. This is especially problematic for cloned e-commerce sites targeting different markets; the problem is especially pronounced. Search engines have almost no reliable signals to determine the intended market when a website uses identical templates and repurposes product descriptions. The subtle differences that do exist (currency symbols, contact information) may be outweighed by the overwhelming similarity in content and structure. At a minimum, search engines need multiple consistent signals to determine geographic targeting confidently.
Algorithm Sequence Challenge
A few years ago, I showed this flow representing what I thought was Google’s processing pipeline in my PubCon keynote. We did significant analysis to understand why it was taking longer for hreflang to kick in to ensure the correct pages were showing in markets.

What deduced a few things. First, if hreflang was not accurate and in place when the page was first detected, it would take at least two months to be correctly applied. Second, reasoning Google’s priorities around reducing indexing and identifying high-value content, we believe that hreflang processing has moved further down the processing flow.
I am extending this thinking better to understand geographical target detection during Google’s processing pipeline when this determination occurs. This timing itself presents challenges:
- Crawling Phase: Some basic geographic signals (like ccTLDs) might be recognized during initial crawling
- Indexing Phase: Language detection typically occurs during indexing, but regional variants may be harder to distinguish at this stage.
- Duplicate Content Filtering: Critical for geographic targeting, this occurs before final ranking but after basic content analysis.
- Query Processing: Some geographic determinations happen at query time based on user location and intent
- Results Ranking: Final geographic relevance adjustments may occur during the ranking phase
Suppose geographic signals aren’t strong enough during the indexing and duplicate filtering phases, and if Google has moved hreflang processing later in the processing chain. In that case, content may be incorrectly grouped or filtered before reaching the stage where geographic relevance for specific queries is evaluated. This means weak geographic signals can cause problems early in the search engine’s processing pipeline that cannot be corrected later. This is further complicated if hreflang is incorrect when it is eventually processed.
Business Implications
Google’s handling of same-language, cross-market pages isn’t a bug — it’s a consequence of how their systems are designed to reduce redundancy and maximize relevance. When content across regional sites is effectively identical, and when geographic signals are weak or inconsistent, Google’s systems will often treat the pages as duplicates and suppress one in favor of what it deems the most authoritative or relevant version. As John Mueller alluded to, this simplification is not malicious — it’s pragmatic.
However, for businesses operating across same-language markets like France and Belgium, this simplification can lead to serious misalignments in user experience, analytics, and revenue attribution. If your Belgian customers land on your French site—or worse, if they’re filtered out altogether—you’re losing not just visibility but potential conversions, trust, and local market presence.
Understanding where in the pipeline these decisions are made — and how early weak signals can doom a page’s visibility — is critical to designing a resilient international SEO strategy.
Recommended Actions
To reduce the risk of geo-targeting misalignment and ensure Google can recognize and retain both pages in the index and serve them correctly, you need to differentiate them intentionally:
- Ensure Valid Hreflang at Creation
Your hreflang must be set and correct when the page goes live so that it has the best chance of being identified as having a market purpose and not be viewed as a duplicate. - Integrate Unique Market Content
Create distinct content differences between country versions (e.g., cultural nuances, dialect variations, or additional content to help Google see them as different. - Introduce Regional Differentiators
Modify more than just the currency or contact info and integrate multiple geolocation signals. Consider localizing aspects of your product pages: availability, reviews, return policies, shipping options, or even imagery. These help Google distinguish pages during content deduplication. - Validate Hreflang Implementation Early and Often
Ensure your hreflang tags are present and correct before the page is first crawled. Retroactive fixes can take months to register, and by then the damage may already be done. - Audit Duplicate Canonical Conflicts
Monitor Search Console regularly for “Duplicate without user-selected canonical” or “Alternate page with proper canonical tag” warnings, especially across markets with same-language content. - Monitor Regional Performance Trends
If you see unexplained dips or flatlining traffic in a specific country where you operate, investigate whether the right version of the page is ranking and appearing. Tools like GSC, logs, and third-party rank trackers can reveal mismatches.
Closing Thought
Google aims to reduce redundant content in its index while improving user experience. When two same-language pages target different markets but are indistinguishable in content, Google may choose only one. Businesses must prove market relevance through real content variation, not just metadata or hreflang.
This isn’t just an International SEO or technical issue — it’s a strategic one. Your ability to appear in the correct country’s search results isn’t just about adding hreflang tags. It’s about recognizing that search engines need help differentiating the goal of the website and web pages, and that help must be baked into every layer of your global web infrastructure. Treat geo-targeting as a first-class citizen in your international strategy, or risk collapsing your efforts into the wrong market or being ignored altogether.