Google UGC Fresh Data Program: Eligibility and How to Prepare

1. What Happened? – Google UGC Fresh Data Program: Eligibility and Prep

On 8 October 2026, Google added a new page to its Search Central documentation: About the UGC Fresh Data Program. The Search Central changelog records it the same day, with the stated purpose of explaining the programme and helping prospective applicants understand the eligibility criteria.

In Google’s words, the programme is “designed to help platforms display fresh, high quality, and authentic user-generated content (UGC) across various features on Search.” Approved platforms get “a dedicated data ingestion pipeline to send content and interaction signals to Google”, so that fresh content “is processed and quickly updated in Search.”

Three things about how it works stand out.

  • It is a separate route in. Google describes it as “a specialized pipeline designed exclusively for trending UGC content”, “independent of the Google Indexing API and traditional organic crawling for general web content.”
  • It carries engagement as well as content. Platforms must be able to send content “as fresh as possible, ideally sent within minutes”, and then send “regular updates of engagement counters (within 72 hours of creation).”
  • It guarantees nothing about visibility. The page says the programme “doesn’t guarantee that content will appear in Search.” The application form says it again: serving “is determined entirely by ranking algorithms.”

Who can apply

Participation is limited to platforms that meet Google’s quality and volume guidelines:

CriterionWhat Google requires
Content focusThe platform must primarily host UGC, social postings or forum discussions, on dedicated pages with stable URLs – not profile or feed pages. Google defines UGC as material “created entirely by independent users rather than the platform itself.”
Volume and audienceA high volume of UGC and a significant user base. No thresholds are published.
Content accessibilityPublic pages accessible to Googlebot and users. Content behind a login or paywall is ineligible. All UGC must be attributable to a creator with a public profile.
Technical readinessSecure OAuth 2.0 API authentication, strictly validated JSON-LD payloads, and valid on-page schema.org markup – for example SocialMediaPosting or DiscussionForumPosting with interactionStatistic.
Content health and moderationNo illegal content, active moderation, and a user reporting mechanism.
Content freshnessSubmission ideally within minutes, and engagement counter updates within 72 hours of creation.

How to apply

Applications go through a Google Form. It asks for the company name, the primary domain content will come from, 10 recent sample URLs, which “should follow the same format/structure that you intend to send to us”, monthly active users, launch date, the email address of an owner of the domain in Search Console, a contact email, and any questions.

Google says it will respond within 6–8 weeks. Acceptance isn’t guaranteed. Accepted platforms receive the technical integration documentation and onboarding instructions – which means the actual API specification is not public yet.

2. Why Does It Matter?

Google now has a fast lane for every kind of content that outruns the crawler

Normally, content reaches Google Search one way: Googlebot discovers a URL, crawls it, and indexes it, on Google’s schedule. A handful of content types have long had faster routes, because they change faster than crawling can follow:

Content typeRoute inWho controls the timing
Products – prices, stockMerchant Center feedsThe merchant pushes updates
Job postings and livestream videoThe Indexing API, which Google restricts to pages with JobPosting or a BroadcastEvent inside a VideoObjectThe site notifies Google
Conversations – forum threads, social postsThe UGC Fresh Data Program (new)The approved platform pushes content and engagement
Everything elseCrawlingGoogle

Products got a feed. Jobs and livestreams got an API. Conversations now get a pipeline. Everything else still waits for Googlebot.

The gap it closes is measured in weeks, not hours

How slow is the normal route? In internal estimates Gary Illyes presented at Search Central Live in early October, which we covered in Google’s crawl, index and serve timings, Google put typical discovery of a new URL at around 20 hours and refreshing a known URL at around 30 days. These are conference estimates, not commitments, but the scale is what matters.

For most content, that’s tolerable. For a trending conversation, it’s fatal. A thread about a product recall, a service outage or a breaking story does almost all of its useful work in the first day or two. Under normal crawling, Google finds it the next day and may not see its reply and like counts updated for a month.

That is why the 72-hour engagement window is the most revealing detail on the page. On-page interactionStatistic markup is only as current as Google’s last crawl, and on a normal refresh cycle that’s weeks out of date. The pipeline lets the platform report the engagement curve as it happens, for the 72 hours when it says the most. Our reading – and it’s an inference, not something Google has said – is that the counters are how Google tells a post that is trending from one that merely exists. The programme’s own description, “exclusively for trending UGC content”, points the same way.

From private deals to a published door – but still a narrow one

Google has fed fresh social content into Search before, through one-off agreements. It had a real-time search deal with Twitter from 2009 until it lapsed in 2011, and a new one in 2015. In February 2024, Google and Reddit announced a partnership giving Google access to Reddit’s Data API.

What’s new is a documented programme with published criteria and an application form. That is a different kind of access: any platform can read the bar and work towards it.

The door is still narrow. “High volume” and “significant user base” have no published numbers, acceptance is discretionary, and Google hasn’t said which platforms are already in. Coverage of the launch has assumed the largest platforms are; that’s plausible, but unconfirmed.

It’s a speed lane, not a ranking boost

Google says it three times. The documentation page says the programme “doesn’t guarantee that content will appear in Search”. The form says “we never guarantee serving”, and that Google “has no obligation to incorporate your content in our features”. The pipeline gets content processed faster. Whether it’s shown is still down to ranking.

For leadership teams, that distinction needs saying early. A successful application is not a traffic forecast. The commercial case is that, for the posts that would rank anyway, the window in which they rank opens hours or days sooner and stays current while it matters.

The criteria are a spec for trustworthy UGC – even if you never apply

Most platforms reading this won’t meet the volume bar. The criteria are still worth reading closely, because they are the most explicit statement Google has published of what it considers trustworthy user-generated content:

  • One post, one permanent URL. Not feeds, not profiles, not URLs that change when a title is edited.
  • Every post attributable to a real, public creator. Anonymous content is out.
  • Public access. No login walls, no “sign up to read more”.
  • Visible moderation. Active moderation and a way for users to report content.
  • Machine-readable engagement. Structured data with real interaction counts.
  • Authenticity. Content “created entirely by independent users rather than the platform itself.”

That last definition has teeth. Staff posts, content a platform has seeded to make a new community look busy, and AI-generated answers presented as community contributions are, by Google’s own definition, not UGC. Google’s forum markup documentation adds a related detail: the digitalSourceType property tells Google when a post was generated by an AI model or a bot, and “if this property is not specified, Google will assume the content is human-generated.” An unlabelled machine-written post is therefore a claim of human authorship, whether you meant it as one or not.

We used the same approach with the UCP location search spec: even where you can’t or won’t implement something, a published spec makes a useful, free audit standard.

Moderation now has to be faster than submission

On a normally crawled site, there’s a buffer between a post going live and Google seeing it: hours at least, often days. Moderators work inside that buffer. A spam post removed within an hour has usually never been crawled.

The pipeline removes the buffer. If content goes to Google “within minutes”, moderation must finish within minutes too – or the platform must decide that some posts wait. Google’s spam policies already cover user-generated spam, and a fast lane that delivers spam in minutes is a fast lane to a policy problem. We set out how spam policies are applied, and how site-wide consequences follow from page-level problems, in the September 2026 spam update.

There’s a security angle too. User-generated content is already a route for prompt injection into AI systems that read the web, as we covered in Gemini 4 Argon as an SEO guide. A direct pipeline into Google’s systems raises the stakes for what gets through moderation.

UK platforms have an additional reason to get this right: user-to-user services within scope of the Online Safety Act already carry duties on illegal content and user reporting. Google’s moderation criterion and those duties overlap, and the evidence for one should support the other.

For brands that don’t host UGC, the conversation layer just got faster

Most businesses will never apply, because they don’t run a UGC platform. They’re affected anyway.

In many categories, the conversations about a brand already outrank the brand. In our Air Compressors Category Leaderboard analysis, Facebook was the most visible domain in the category and none of the top 20 were manufacturers – because, as we put it there, manufacturers publish specifications and communities publish consequences.

If the largest platforms can now get trending threads into Search within minutes, a complaint, a product fault or a service outage discussed on those platforms can become searchable before the brand has said anything. The practical response isn’t SEO on the brand’s own site. It’s being present, quickly and credibly, in the conversations themselves.

If you join without a baseline, you’ll never know what it did

One more reason to start preparing now. If you’re accepted, you’ll want to know whether the pipeline made a difference. That needs a measure of how long your new posts currently take to be crawled and indexed – taken before integration. Afterwards, the before-picture is gone.

3. Who Is Affected?

Large UGC platforms – forums, social networks, Q&A sites and community review sites. The direct audience. If you meet the volume bar, this is the most significant change to how your content reaches Google in years. Prepare now, apply when your sample URLs are clean.

Brand-owned communities – telecoms, consumer technology, gaming, finance and software support communities. Some of these are large enough to qualify, and they hold exactly the first-hand problem-solving content the programme describes. Two catches. Staff and moderator posts aren’t UGC by Google’s definition, so they must be excluded from submission. And many brand communities run on hosted community software, so eligibility depends on whether the vendor can support OAuth, strict JSON-LD and custom data flags.

Health, finance and other high-stakes communities. Patient forums, money forums and legal-advice communities hold valuable first-hand content, but the moderation burden is higher, and errors in fast-tracked content carry more risk. The moderation and hold-back recommendations in section 4C matter most here.

Smaller and niche forums. Probably below the volume bar. The on-page requirements are still worth meeting, because Google already uses the same forum markup for its Discussions and forums features through normal crawling.

Publishers with comment sections and retailers with reviews. Probably not eligible: the platform must primarily host UGC. Product reviews have their own markup and their own features.

Platforms behind login walls – private groups, members-only communities, sites that show a sign-up wall after a few pages. Ineligible as they stand. “Soft” walls that block the reader after a set number of page views are a grey area Google hasn’t addressed; treat them as a risk.

Brands with no community of their own. Affected through the competitive effect described above. The conversation layer in Search gets faster and more current, and the brand’s own pages don’t.

Less affected: sites whose content isn’t conversational, and B2B businesses in categories with little public discussion – though even there, it’s worth checking where your customers do talk.

4. What Should Businesses Do?

The work splits into three tracks: decide whether you’re a candidate; make the site meet the spec, whether or not you apply; and, if you do apply, build the pipeline behind proper moderation and measurement.

4A. Everyone: find out which group you’re in

Step 1. Score yourself against the criteria. For each, write down the evidence you’d show Google.

CriterionThe testEvidence to have ready
Content focusIs most of the platform’s content created by users, on per-post pages?Share of content by author type; URL patterns for post pages
Volume and audienceWould a reasonable reviewer call it “high volume” and “significant”?Monthly active users, posts per day, launch date
AccessibilityCan a logged-out visitor read every post you’d submit?Logged-out crawl of a sample; no soft walls
AttributionDoes every eligible post link to a public creator profile?Profile pages with ProfilePage markup
Technical readinessCan you produce valid forum markup and run an OAuth 2.0 integration?Rich Results Test passes; engineering owner named
ModerationActive moderation and a user reporting mechanism?Published community guidelines, report button, moderation response times
FreshnessCan you send within minutes and update counters for 72 hours?Event-driven architecture, or a plan for one

Step 2. Decide which track you’re on.

  • Apply. You meet every criterion, or are close and can close the gap during the 6–8 week wait.
  • Prepare. You meet the content and moderation criteria but not the technical ones yet. Do the 4C work first, then apply.
  • Adopt the spec without applying. You’re below the volume bar or not a UGC platform. Do the on-page work anyway — it serves the Discussions and forums features through normal crawling – and focus on the brand-side response in 4B.

Step 3. Take a baseline now. Whatever track you’re on, measure how long new posts currently take to be crawled and indexed. Sample a few dozen new posts a day through the URL Inspection API and record when each was first crawled – the method is in Google’s crawl, index and serve timings. Keep this running. It’s the only way to show what the pipeline changes, if you’re accepted.

4B. For the content and marketing team

Treat the 10 sample URLs as the application. Google asks for 10 recent URLs that follow the same format you intend to send. They’re the only content a reviewer is certain to see. Choose them deliberately:

  • Recent. Posts from the last few days, not a hand-picked archive.
  • Representative. The same template and markup as everything you’ll send. Ten different page types suggests you don’t have one.
  • Genuinely user-generated. No staff announcements, no moderator posts, no AI-assisted answers.
  • Clean. Every one passes the audit script in 4C, logged out, with no warnings you can’t explain.
  • Uncontroversial. Avoid threads with open reports, heavy moderation or sensitive subjects. A reviewer will read them.

Get the details on the form right.

  • Monthly active users: use the same definition as your annual report or investor materials. A number that doesn’t match what you’ve published elsewhere is a credibility problem.
  • Search Console owner: the form asks for “an owner of your domain on Search Console”. Use a role account held by the team that will own the integration, not a personal account that leaves when someone does. Check that it’s a verified owner, not a user with delegated access.
  • Additional comments: use them. The public page leaves important questions open. Ask how deleted and edited posts are handled; whether moderation retractions are supported; what rate limits apply; whether historical content can be sent; which Search features and “other surfaces” pipeline content can appear on; and how participating platforms can measure the effect. Applicants who ask precise questions read as applicants who’ll integrate properly.

Make moderation visible. Publish community guidelines at a stable URL. Make the report button obvious on every post. If you publish moderation statistics or a transparency report, link to it. Google is asking for active moderation; make it easy to see.

Make creator profiles public by design, not by force. Every submitted post needs a creator with a public profile. Users who choose private profiles are entitled to that choice – exclude their posts from submission rather than changing their settings. Make sure public profiles carry ProfilePage markup, which Google documents for exactly this context.

Set expectations with leadership in writing. Acceptance is discretionary. Visibility isn’t guaranteed. The likely benefit is speed for content that would rank anyway. Put that in the business case before someone else puts a traffic target on it.

If you’re a brand without a qualifying community, work on the conversations themselves:

  • Monitor the platforms where your customers talk – forums, Reddit, Facebook groups, review communities – for your brand, your products and your category’s problems.
  • Agree a response time for public threads that matches the speed of the platforms, not the speed of your content calendar. If a thread can be in Search within minutes, a reply three days later arrives after the story has set.
  • Answer as a named, accountable person, with a public profile, disclosing who you work for. That’s what “attributable to a creator” means in practice, and it’s what readers trust.
  • Feed what you learn back into your own content. The questions customers ask in communities are the content your site is missing. We set out a UK community listening programme in our Air Compressors Category Leaderboard analysis.

4C. For the development team

The integration specification isn’t public until acceptance, so the work now is everything the specification will sit on: a data model that knows what’s eligible, URLs that don’t move, one markup builder used everywhere, a moderation gate, a counter schedule and the logging to prove it works.

1. Add the flags the criteria need to your post model. Each post needs to know:

FieldValuesWhy
author_typeuser, staff, moderator, botOnly independent users’ posts are UGC
visibilitypublic, members, privateGated content is ineligible
author.profile_publictrue/falseEvery post needs a public creator profile
moderation_stateapproved, pending, reported, removedNothing unmoderated goes into a fast lane
ai_sourcenone, llm, botLabel in markup, exclude from submission
deletedtrue/falseMark tombstones; never submit
created_attimezone-aware timestampFreshness and the 72-hour window depend on it
countslikes, dislikes, views, comments, sharesMaps onto Google’s supported interaction types

2. Give every post a permanent URL. The URL must carry a permanent ID, so editing a title never changes it; a slug can follow the ID and redirect when it changes. The canonical tag must equal the URL you’d submit. Feeds, profiles and sorted or paginated views are not posts.

3. Build markup once and use it everywhere. Generate the on-page JSON-LD and anything you send through the pipeline from the same function and the same record. If what you push says one thing and the page Googlebot later crawls says another, you’ve built an inconsistency into your own integration. Merchant Center works on the same principle: feed data has to match the landing page.

4. Gate submission behind moderation, and decide what waits. Moderation can’t take hours if submission takes minutes. Practical options:

  • Submit approved posts immediately, where automated checks run at creation and pass.
  • Hold back higher-risk posts – new accounts, first posts, posts with links, posts in sensitive sections – until a moderator or a stronger automated check clears them. Lateness is a cost; spam in Google’s index is a bigger one.
  • Never submit staff, bot, AI-generated, deleted, members-only or private-profile posts.

5. Schedule engagement updates inside the 72 hours. Send counters on a fixed schedule – for example at 15 minutes, 1, 4, 12, 24, 48 and 72 hours – and stop. Make the schedule configurable, because the developer documentation may specify rate limits or intervals.

6. Prepare the OAuth 2.0 side properly. Expect a service-account-style integration. Keep credentials in your secrets manager, never in the repository; rotate them; and give the integration account no permissions beyond what the pipeline needs.

7. Build deletion and edit hooks now. Users delete posts, moderators remove them, and UK GDPR erasure requests remove whole accounts. The public documentation doesn’t say how the pipeline handles any of this – ask during the application. Either way, fire an event on every deletion, edit and erasure, so you can connect it to whatever the specification provides. On the page, mark kept tombstones with creativeWorkStatus: Deleted, and return 404 or 410 for posts that are gone.

8. Log everything. For each submission, record the post ID, a hash of the payload, the response, and the time from creation to submission and from approval to submission. Those two latencies are your operational KPIs, and the evidence if anything goes wrong.

The builder, validator and gate

This script builds the posting markup from your own post record, validates it strictly against Google’s documented forum markup and the programme’s criteria, decides whether a post should be pushed, and schedules counter updates. It’s deliberately independent of the pipeline’s transport, which isn’t public yet.

python

"""Build, validate and gate UGC posts — one source of truth for on-page markup and any push pipeline."""
import re, json
from datetime import datetime, timedelta, timezone

TYPES = {"forum": "DiscussionForumPosting", "social": "SocialMediaPosting"}
COUNTERS = {"likes": "LikeAction", "dislikes": "DislikeAction", "views": "ViewAction",
            "comments": "CommentAction", "shares": "ShareAction"}
ALLOWED_INTERACTIONS = {f"https://schema.org/{t}" for t in
                        ("LikeAction", "DislikeAction", "ViewAction", "CommentAction", "ReplyAction", "ShareAction")}
AI_SOURCE = {"llm": "https://schema.org/TrainedAlgorithmicMediaDigitalSource",
             "bot": "https://schema.org/AlgorithmicMediaDigitalSource"}

# Adjust to your own routing. Post URLs should carry a permanent ID; feeds and profiles are not eligible.
STABLE_POST_URL = re.compile(r"^https://[^/?#]+/(?:t|p|post|posts|thread|threads|discussion|d|questions)/(?:[^?#]*/)?\d+(?:/[^?#]*)?$")
NOT_A_POST = re.compile(r"/(?:feed|feeds|latest|trending|u|user|users|profile|members?)(?:/|$)|[?&](?:sort|page|ref|utm_[a-z]+)=", re.I)

COUNTER_WINDOW = timedelta(hours=72)
UPDATE_OFFSETS = [timedelta(minutes=m) for m in (15, 60, 240, 720, 1440, 2880, 4320)]   # 4320 min = 72 h


def iso(dt):
    if dt.tzinfo is None:
        raise ValueError("timestamps must be timezone-aware")
    return dt.isoformat(timespec="seconds")


def build_jsonld(post):
    """Build the posting node from the platform's own post record."""
    node = {
        "@context": "https://schema.org",
        "@type": TYPES[post["kind"]],
        "url": post["url"],
        "datePublished": iso(post["created_at"]),
        "author": {"@type": "Person", "name": post["author"]["name"]},
    }
    if post["author"].get("profile_url"): node["author"]["url"] = post["author"]["profile_url"]
    if post.get("title"):       node["headline"] = post["title"]
    if post.get("text"):        node["text"] = post["text"]
    if post.get("images"):      node["image"] = post["images"]
    if post.get("edited_at"):   node["dateModified"] = iso(post["edited_at"])
    if post.get("community_url"): node["isPartOf"] = post["community_url"]
    if post.get("ai_source"):   node["digitalSourceType"] = AI_SOURCE[post["ai_source"]]
    if post.get("deleted"):     node["creativeWorkStatus"] = "Deleted"
    counts = post.get("counts", {})
    if "comments" in counts:    node["commentCount"] = counts["comments"]
    stats = [{"@type": "InteractionCounter", "interactionType": f"https://schema.org/{COUNTERS[k]}",
              "userInteractionCount": v} for k, v in counts.items() if k in COUNTERS]
    if stats:                   node["interactionStatistic"] = stats
    return node


def _walk_nulls(obj, path="$"):
    if obj is None or obj == "":
        yield path
    elif isinstance(obj, dict):
        for k, v in obj.items():
            yield from _walk_nulls(v, f"{path}.{k}")
    elif isinstance(obj, list):
        for i, v in enumerate(obj):
            yield from _walk_nulls(v, f"{path}[{i}]")


def _parse_ts(value):
    try:
        dt = datetime.fromisoformat(str(value).replace("Z", "+00:00"))
    except ValueError:
        return None
    return dt if dt.tzinfo else "naive"


def validate(node, now=None):
    """Strict checks against Google's documented forum markup plus the programme's eligibility rules."""
    now = now or datetime.now(timezone.utc)
    errors = []
    if node.get("@type") not in TYPES.values():
        errors.append(f"wrong_type:{node.get('@type')}")
    url = node.get("url", "")
    if not url:
        errors.append("missing_url")
    elif NOT_A_POST.search(url):
        errors.append("feed_profile_or_parameter_url")
    elif not STABLE_POST_URL.match(url):
        errors.append("url_without_permanent_id")
    ts = _parse_ts(node.get("datePublished"))
    if ts is None:              errors.append("datePublished_not_iso8601")
    elif ts == "naive":         errors.append("datePublished_missing_timezone")
    elif ts > now + timedelta(minutes=5): errors.append("datePublished_in_future")
    author = node.get("author") or {}
    if isinstance(author, list):
        author = author[0] if author else {}
    if not author.get("name"):  errors.append("missing_author_name")
    if not author.get("url"):   errors.append("missing_author_profile_url")
    if not any(node.get(k) for k in ("text", "image", "video", "sharedContent")):
        errors.append("missing_content")
    for i, s in enumerate(node.get("interactionStatistic") or []):
        if s.get("@type") != "InteractionCounter":
            errors.append(f"interactionStatistic[{i}]_not_InteractionCounter")
        itype = s.get("interactionType", "")
        itype = itype.get("@type", "") if isinstance(itype, dict) else itype
        if f"https://schema.org/{itype}" not in ALLOWED_INTERACTIONS and itype not in ALLOWED_INTERACTIONS:
            errors.append(f"interactionStatistic[{i}]_unsupported_type:{itype}")
        c = s.get("userInteractionCount")
        if not isinstance(c, int) or isinstance(c, bool) or c < 0:
            errors.append(f"interactionStatistic[{i}]_bad_count:{c!r}")
    errors += [f"empty_value:{p}" for p in _walk_nulls(node)]
    return errors


def gate(post, now=None):
    """Should this post be pushed at all? Separate from markup: the page can carry markup for posts we never push."""
    now = now or datetime.now(timezone.utc)
    blocks, warnings = [], []
    if post.get("visibility") != "public":                 blocks.append("not_public")
    if not post["author"].get("profile_public"):           blocks.append("author_profile_not_public")
    if post.get("author_type") != "user":                  blocks.append(f"not_independent_user:{post.get('author_type')}")
    if post.get("moderation_state") != "approved":         blocks.append(f"moderation:{post.get('moderation_state')}")
    if post.get("deleted"):                                blocks.append("deleted")
    if post.get("ai_source"):                              blocks.append("machine_generated")
    blocks += validate(build_jsonld(post), now)
    age = now - post["created_at"]
    if age > timedelta(hours=1):                           warnings.append(f"late_submission:{int(age.total_seconds() // 60)}min")
    return (not blocks), blocks, warnings


def counter_update_due(created_at, last_sent_at, now=None):
    """True if an engagement-counter update is due. Updates follow fixed offsets and stop at 72 hours."""
    now = now or datetime.now(timezone.utc)
    if now - created_at > COUNTER_WINDOW + timedelta(minutes=30):   # small grace for the final update
        return False
    due = [created_at + off for off in UPDATE_OFFSETS if created_at + off <= now]
    return bool(due) and (last_sent_at is None or last_sent_at < due[-1])


if __name__ == "__main__":
    import sys
    for line in sys.stdin:                                  # one JSON post record per line
        rec = json.loads(line)
        for key in ("created_at", "edited_at"):
            if rec.get(key):
                rec[key] = datetime.fromisoformat(rec[key])
        ok, blocks, warnings = gate(rec)
        print(json.dumps({"url": rec["url"], "push": ok, "blocks": blocks, "warnings": warnings,
                          "jsonld": build_jsonld(rec)}))

What it enforces. Interaction types are limited to those Google documents for forum markup – likes, dislikes, views, comments or replies, and shares – written as full schema.org URLs, with integer counts. Timestamps must carry a timezone. URLs must carry a permanent ID and must not be feeds, profiles or parameterised views; adjust the two URL patterns to your own routing. AI-generated posts are labelled in the on-page markup and excluded from submission. That’s a conservative policy choice on our part, not a Google rule: Google’s definition of UGC is content “created entirely by independent users”, and we’d rather not test it.

Audit the 10 sample URLs before they go on the form

This script fetches each sample URL as Googlebot and as an ordinary logged-out visitor, and checks what a reviewer would. It looks for noindex, canonical mismatches, login walls, content that’s in the markup but not on the page, invalid JSON-LD, posting markup that fails the validator above, and whether all ten share one markup template.

python

"""Audit the 10 sample URLs before they go on the application form."""
import re, sys, json
from datetime import datetime, timedelta, timezone
from html.parser import HTMLParser
from html import unescape
import requests
from ugc_ready import validate, TYPES, NOT_A_POST

UA = {"googlebot": "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
      "visitor":   "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/124.0.0.0 Safari/537.36"}
LOGIN_WALL = re.compile(r"(log ?in|sign ?in|sign ?up|register|subscribe) to (continue|view|read|see) (this|the|more|full)", re.I)


class _Page(HTMLParser):
    def __init__(self):
        super().__init__(); self.ld, self.meta_robots, self.canonical = [], "", None
        self.password_field, self._in_ld, self._skip, self.text = False, False, 0, []
    def handle_starttag(self, tag, attrs):
        a = {k.lower(): (v or "") for k, v in attrs}
        if tag == "script" and a.get("type", "").lower() == "application/ld+json":
            self._in_ld = True; self.ld.append("")
        elif tag in ("script", "style", "noscript"):
            self._skip += 1
        elif tag == "meta" and a.get("name", "").lower() in ("robots", "googlebot"):
            self.meta_robots += " " + a.get("content", "").lower()
        elif tag == "link" and "canonical" in a.get("rel", "").lower().split():
            self.canonical = a.get("href")
        elif tag == "input" and a.get("type", "").lower() == "password":
            self.password_field = True
    def handle_startendtag(self, tag, attrs):
        if tag not in ("script", "style", "noscript"):
            self.handle_starttag(tag, attrs)
    def handle_endtag(self, tag):
        if tag == "script" and self._in_ld: self._in_ld = False
        elif tag in ("script", "style", "noscript") and self._skip: self._skip -= 1
    def handle_data(self, data):
        if self._in_ld: self.ld[-1] += data
        elif not self._skip: self.text.append(data)


def _nodes(obj):
    if isinstance(obj, list):
        for x in obj: yield from _nodes(x)
    elif isinstance(obj, dict):
        yield obj
        for x in obj.get("@graph", []) if isinstance(obj.get("@graph"), list) else []:
            yield from _nodes(x)


def _norm(u):
    return (u or "").split("#")[0].rstrip("/")


def audit(url, now=None, max_age_days=30):
    now = now or datetime.now(timezone.utc)
    out = {"url": url, "problems": [], "warnings": [], "shape": None}
    if NOT_A_POST.search(url):
        out["problems"].append("feed_profile_or_parameter_url")
    pages = {}
    for who, ua in UA.items():
        r = requests.get(url, headers={"User-Agent": ua}, timeout=20, allow_redirects=False)
        pages[who] = r
        if r.status_code != 200:
            out["problems"].append(f"{who}_status_{r.status_code}")
    r = pages["googlebot"]
    if r.status_code != 200:
        return out
    if "noindex" in r.headers.get("X-Robots-Tag", "").lower():
        out["problems"].append("x_robots_noindex")
    p = _Page(); p.feed(r.text)
    if "noindex" in p.meta_robots:
        out["problems"].append("meta_noindex")
    if _norm(p.canonical) != _norm(url):
        out["problems"].append(f"canonical_mismatch:{p.canonical}")
    visible = re.sub(r"\s+", " ", unescape(" ".join(p.text)))
    if LOGIN_WALL.search(visible):
        out["problems"].append("login_wall_wording")
    if pages["visitor"].status_code == 200 and len(pages["visitor"].text) < 0.5 * len(r.text):
        out["problems"].append("visitor_sees_much_less_than_googlebot")
    posts, bad_json = [], 0
    for block in p.ld:
        try:
            data = json.loads(block)
        except json.JSONDecodeError:
            bad_json += 1; continue
        posts += [n for n in _nodes(data) if n.get("@type") in TYPES.values()]
    if bad_json:
        out["problems"].append(f"invalid_json_ld_blocks:{bad_json}")
    main = [n for n in posts if _norm(n.get("url")) == _norm(url)] or posts[:1]
    if not main:
        out["problems"].append("no_posting_markup")
        return out
    node = main[0]
    out["problems"] += validate(node, now)
    snippet = re.sub(r"\s+", " ", (node.get("text") or "")).strip()[:60]
    text_visible = not snippet or snippet in visible
    if not text_visible:
        out["problems"].append("markup_text_not_in_visible_html")
    if p.password_field:
        # A header login box is normal on public forums; a password field plus missing post text is a wall.
        (out["problems"] if not text_visible else out["warnings"]).append("password_field_on_page")
    try:
        published = datetime.fromisoformat(str(node.get("datePublished")).replace("Z", "+00:00"))
        if published.tzinfo and now - published > timedelta(days=max_age_days):
            out["warnings"].append(f"older_than_{max_age_days}_days")
    except ValueError:
        pass
    out["shape"] = sorted(k for k in node if not k.startswith("@"))
    out["problems"] = list(dict.fromkeys(out["problems"]))
    return out


def audit_set(urls, now=None):
    reports = [audit(u, now) for u in urls]
    shapes = {tuple(r["shape"]) for r in reports if r["shape"]}
    summary = {"urls": len(urls),
               "clean": sum(1 for r in reports if not r["problems"]),
               "consistent_template": len(shapes) <= 1,
               "distinct_markup_shapes": len(shapes)}
    if len(urls) != 10:
        summary["note"] = "the form asks for exactly 10 sample URLs"
    return summary, reports


if __name__ == "__main__":
    urls = [l.strip() for l in open(sys.argv[1]) if l.strip()]
    summary, reports = audit_set(urls)
    for r in reports:
        print(("OK   " if not r["problems"] else "FIX  ") + r["url"], "; ".join(r["problems"] + r["warnings"]))
    print(json.dumps(summary))

Run it with a text file of the ten URLs: python3 audit_samples.py samples.txt. Every line should read OK. A password_field_on_page warning on its own is normal – most public forums have a login box in the header. It only becomes a problem when the post text is missing from the page as well, which is what a login wall looks like.

Know what it can’t tell you. It tests from your network, not Google’s. It checks the page Google would crawl, not the pipeline payload, whose format isn’t public. And it can’t judge whether your volume or moderation meets Google’s bar.

4D. Rolling it out

PhaseTimingWork
1. Assess and baselineWeeks 1-2Score against the criteria (4A). Start the crawl-lag baseline. Name an owner.
2. Data model and URLsWeeks 2-6Add the post flags. Fix URL stability and canonicals. Ship the shared markup builder to every post page.
3. Audit and applyWhen readyRun the audit on a fresh set of 10 sample URLs. Apply when all ten are clean. Use the comments field to ask the open questions.
4. Build behind a flagDuring the 6-8 week waitModeration gate, hold-back queue, counter scheduler, deletion hooks, logging and secrets – everything except the transport.
5. Integrate and pilotOn acceptanceIntegrate to the developer documentation. Start with one section of the platform. Compare crawl lag and visibility against the baseline.
6. Expand or stopAfter 4-8 weeks of pilot dataWiden to the whole platform if the pilot shows a benefit and moderation is holding up. If not, keep the on-page work and revisit.

If you’re on the “adopt the spec” track, phases 1 and 2 are the whole job, plus the brand-side work in 4B.

4E. Governance

Name one owner. The integration crosses product, engineering, trust and safety, legal and SEO. Without a single owner, it stalls at the first disagreement — usually between how fast engineering can send and how fast moderation can clear.

Set moderation service levels before you send anything. Agree target times from post to moderation decision, and which classes of post wait for a human. Review them monthly against the creation-to-submission and approval-to-submission latencies you’re logging.

Run a data protection review. You’ll be sending users’ content, usernames and profile URLs to a third party through a direct pipeline. The content is public, but a direct transfer is a different processing activity from being crawled. Check your privacy notice, your lawful basis and how erasure requests flow through. The form also says submitted content is eligible for “Google Search and other surfaces”, which is broader than the documentation page’s “various features on Search”; your privacy notice should be written with that in mind. We’re not lawyers – this belongs with your DPO or counsel.

Put the markup under change control. The shared builder is now part of an external integration. Changes go through review, with the validator in CI, so a template change can’t silently break every payload.

Keep a kill switch. If moderation falls behind, or a spam wave gets through, you need to pause submission in minutes without a deployment.

Report on the right measure. Report crawl lag, the share of new posts indexed within 24 hours, and visibility for fresh content against the baseline – not total traffic. Total traffic will move for a dozen reasons that have nothing to do with the pipeline.

5. What We’re Watching Next

What the developer documentation says. Above all: how deletions, edits and moderation retractions are handled, what rate limits apply, and whether historical content can be sent. These decide how hard the integration is, and they’re not public yet.

Who is in. Google hasn’t named participating platforms. If it does, or if the pattern becomes visible in results, it will show how high the volume bar really is.

Which surfaces the content reaches. The documentation says “various features on Search”; the form says “Google Search and other surfaces”. Whether that includes AI Overviews, AI Mode or Discover, and whether participating platforms can see where their content appeared, are open questions.

Whether Search Console reports it. A participating platform would want to see pipeline-submitted content separately from crawled content. Nothing has been announced.

Whether eligibility widens. Lower volume thresholds, brand communities on hosted software, or other fast-moving content types would bring many more businesses into scope.

Whether fast-tracked UGC changes what ranks for trending queries. If it does, the platforms in the programme will gain an advantage on breaking topics that everyone else – publishers and brands included – will need to plan around.

6. About Szymaniak Digital

Szymaniak Digital is an enterprise AI SEO consultancy. We help platforms and brands understand how Google’s systems actually take in content – through crawling, feeds and now pipelines – and build the technical, editorial and governance work that follows.

If you run a community or UGC platform and want an independent read of whether you’re ready to apply, the criteria table in 4A and the sample URL audit in 4C are where we’d start. If you’re a brand whose customers talk about you on other people’s platforms, we can help you build the listening and response programme in 4B. Need our support?

Contact Us!

Scroll to Top