The Trigger: Google Pauses Its OSS Bug Bounty Program
Google's Open Source Software Vulnerability Rewards Program went on pause as of October 1, 2026, with the company promising an update in the first quarter of 2027. Other Google bug bounty programs remain open. The stated cause: a flood of automated submissions that engineers and maintainers could not triage fast enough. According to Tom's Hardware, Google engineers were overwhelmed by reports that were invalid or contained hallucinations. This wasn't a theoretical risk — it was an operational denial of service against the human review layer.
The pattern matches what security researchers warned about in July 2025: LLMs make it trivial to generate plausible-looking but semantically empty vulnerability reports at scale. The incentive structure of bug bounties — cash for valid findings — creates a direct economic motive for spam. But the same dynamic applies anywhere LLM output enters a shared corpus without verification: package registries, documentation wikis, Stack Overflow answers, and the training data of the next model generation.
Why Bug Bounty Spam Is a Leading Indicator for RAG Poisoning
Bug bounty programs have a built-in filter: human maintainers who reject invalid reports. RAG pipelines often lack an equivalent gate. When you point a retriever at a GitHub organization, a language ecosystem index, or a documentation snapshot, you inherit every commit, every merged PR, every generated docstring — whether a human reviewed it or not.
The threat model shifts in three ways:
- Volume asymmetry. One engineer can review perhaps 20 vulnerability reports per day. An LLM can generate 20,000 in the same window. The same ratio applies to code commits, documentation updates, and example snippets.
- Plausibility without correctness. Hallucinated vulnerabilities mimic the structure of real reports — file paths, function names, CVE-style language. Hallucinated code mimics syntax, imports, and commenting style. Retrieval systems optimized for semantic similarity cannot distinguish "looks right" from "is right" without external signals.
- Adversarial poisoning. Beyond opportunistic spam, attackers can seed repositories with malicious patterns designed to be retrieved: backdoored utility functions, credential-leaking examples, or API usage that bypasses authorization checks. If your RAG system treats the corpus as authoritative, these become injected attack vectors.
Failure Modes: How AI Slop Breaks RAG Pipelines
Concrete failure modes observed in production pipelines ingesting open source corpora:
Hallucinated APIs retrieved as ground truth
A retriever surfaces a code snippet showing a method client.secure_transfer() that never existed in the library. The generator incorporates it into an answer. The developer copies it. The code fails at runtime — or worse, the hallucinated method name matches a real but dangerous method in a different context.
Duplicate and near-duplicate chunks inflate context windows
AI-generated documentation often repeats the same explanation across dozens of files with minor variations. Chunking strategies that don't deduplicate by semantic hash waste context budget on noise, pushing relevant chunks out of the window.
Poisoned few-shot examples embedded in retrieved snippets
Many repositories include example directories. If an attacker or spammer contributes examples containing subtle bugs — hardcoded secrets, disabled certificate validation, SQL concatenation — those examples become part of the retrieved context for "how do I use this library?" queries.
Maintainer burnout reduces human curation
The bug bounty freeze demonstrates a second-order effect: when maintainers drown in slop, they stop reviewing legitimate contributions. The historical quality signal — "a human merged this" — degrades. Your provenance metadata becomes less reliable over time.
Defense Layer 1: Pre-Ingestion Filtering and Provenance
The cheapest place to stop slop is before it enters your vector store. Armor Tech's RAG pipeline hardening toolkit implements these filters at the ingestion boundary:
- AI-text classifiers. Run a lightweight detector (fine-tuned DeBERTa or similar) on every document at ingest. Flag high-probability LLM output for quarantine. False positives are manageable; false negatives propagate.
- Burst-rate monitoring. Track submission velocity per contributor, per repository, per file path. A single account opening 50 PRs across 30 repos in 4 hours is a signal — whether the content is AI-generated or human copy-paste.
- Contributor reputation scoring. Weight trust signals: signed commits, CI pass history, prior merged PRs, organizational affiliation, GPG keys. New pseudonymous accounts with no history start at low trust.
- Provenance metadata capture. Store for every chunk: source repository, commit SHA, author identity, CI status, merge approvals, file path, and a content hash. This enables downstream rollback and audit.
- Automated quarantine. Submissions matching known slop patterns — repetitive phrasing, impossible API combinations, hallucinated imports — route to a holding index. Human review or automated re-evaluation decides promotion.
Filtering at ingest costs compute once. Re-indexing a polluted corpus costs compute repeatedly and risks downstream generation failures you discover only in production.
Defense Layer 2: Retrieval-Time Guardrails
Pre-filtering misses sophisticated or novel slop. Retrieval-time checks catch what enters the index:
- Cross-encoder re-ranking with quality signals. Beyond semantic similarity, score chunks by: provenance trust score, recency (older stable code > newer unvetted code), test coverage proximity (chunks near tested code paths), and deduplication penalty.
- Consistency checks against reference implementations. Maintain a curated set of known-good patterns per library. When retrieved chunks deviate — different parameter orders, missing error handling, extra calls — flag or downweight.
- Citation verification. Require the retriever to surface verifiable sources: commit URLs, documentation permalinks, test files. Chunks without traceable provenance get a confidence penalty.
- Confidence thresholds with human-in-the-loop escalation. If the top-k retrieved chunks have mean trust score below a configurable threshold, the pipeline refuses to generate and routes the query to a review queue with the suspect chunks highlighted.
Defense Layer 3: Continuous Corpus Hygiene
Corpus quality is not a one-time achievement. It decays as new slop enters and old code rots:
- Scheduled re-evaluation of high-churn repositories. Repos with frequent commits, many contributors, or recent ownership changes get re-scanned weekly. Stable core libraries get monthly.
- Automated rollback of chunks linked to reverted or flagged commits. When a PR is reverted, a security advisory is published, or a maintainer flags a contribution, the ingestion pipeline traces all chunks from that commit and marks them deprecated in the vector store.
- Feedback loops from user interactions. Explicit downvotes, "this didn't work" signals, and implicit signals (user reformulates query immediately) feed back into chunk weight adjustment. Persistent negative signal triggers re-review.
- Monitoring dashboard. Track slop influx rate (flagged submissions per day), retrieval precision drift (evaluation set performance over time), quarantine queue depth, and human review throughput. Alert on anomalies.
What This Means for Your Roadmap
Treat corpus hygiene as a recurring operational expense, not a one-time data cleaning project. Budget for:
- Classifier model updates as LLM output evolves
- Human reviewer time for quarantine triage
- Re-indexing compute when rollback triggers
- Monitoring infrastructure and alerting
Evaluate RAG vendors and frameworks on slop resilience, not just benchmark scores on clean datasets. Ask: how does this system behave when 30% of the corpus is AI-generated noise? What provenance metadata does it capture? Can I inject custom trust signals?
Build "trust but verify" into every automated ingestion pipeline. The bug bounty freeze proves that even well-resourced organizations with dedicated security teams can be overwhelmed. Your RAG pipeline has fewer human reviewers and higher query volume.
Armor Tech's agentic AI security services help teams threat-model their ingestion paths, deploy the three-layer filter stack, and establish the observability needed to detect drift before it corrupts production answers.
Frequently Asked Questions
How do I know if my RAG corpus already contains AI-generated slop?
Run a sample audit: pull 500 random chunks from your vector store, run them through an AI-text classifier, and manually review the high-confidence positives. Look for hallucinated imports, impossible API combinations, and repetitive documentation patterns. If you find more than a few percent, assume the problem is systemic.
Can't I just use a better embedding model to filter out noise?
Embedding models optimize for semantic similarity, not factual correctness. A hallucinated API call embedded in plausible context will have high cosine similarity to legitimate usage examples. You need orthogonal signals — provenance, testing, human review — not better vectors.
What's the minimum viable defense for a small team with limited review capacity?
Start with provenance capture on every ingested chunk (commit SHA, author, CI status) and a simple quarantine rule: any contribution from accounts with zero prior merged PRs in your target repositories goes to a holding index. Review the holding index weekly. This catches the highest-volume spam with minimal tooling investment.




