- The Next Input by Cylentis AI
- Posts
- 🎮 The Next Input — Issue #217
🎮 The Next Input — Issue #217
The Pentagon, the Treasury, and the Hallucinated Citation

⚡ The Briefing — 60 sec
The Pentagon now has its own version of ChatGPT and Grok The Pentagon having its own ChatGPT and Grok is somehow both completely unsurprising and absolutely wild. Once frontier models become part of military infrastructure, we’re a long way from “AI assistant that drafts emails.”
AI hallucinations are infiltrating Australian parliament They get paid to talk shit and then spit out AI slop instead of doing their jobs. Come on, y’all. If a parliamentary submission can’t survive a basic citation check, maybe don’t let it anywhere near policymaking.
Treasury tips AI as a key driver of productivity growth This is the part that matters. Treasury is effectively saying AI could become a serious lever for Australian productivity growth. Cool. Now can we please build the capability required to actually capture it instead of admiring the opportunity from across the room?
🛠️ The Playbook — Policy Evidence Firewall
Mission
Build an AI-assisted policy and executive research workflow where consequential claims cannot reach decision-makers without verifiable evidence attached.
Difficulty
Intermediate
Build time
3–5 hours
ROI
Accelerates research while reducing hallucinated citations, weak evidence and the risk of bad decisions being made from polished nonsense.
0) Why This Matters
AI is becoming infrastructure.
Military infrastructure.
Economic infrastructure.
Policy infrastructure.
That raises the standard.
If governments and businesses are going to use AI to make consequential decisions, the workflow cannot simply be:
Prompt → persuasive paragraph → send.
The more important the decision, the more important provenance becomes.
Who said this?
Where did the number come from?
Does the cited study exist?
Does it actually support the claim?
A system that answers those questions automatically is considerably more useful than another chatbot.
1) Architecture
Component | Tool | Purpose | Owner | Failure mode |
|---|---|---|---|---|
Evidence intake | SharePoint / approved web sources | Captures authoritative reports, datasets and submissions | Research Lead | Low-quality evidence enters the corpus |
Retrieval layer | Azure AI Search | Finds claim-relevant source passages | Policy / Strategy | Irrelevant context is retrieved |
Reasoning layer | GPT-5.6 / Claude | Synthesises evidence into structured analysis | Analyst | Unsupported conclusions appear |
Citation verifier | Python / Crossref APIs | Checks whether references exist and match claims | Research Ops | Fabricated citations slip through |
Review workflow | Microsoft Teams Approvals | Escalates low-confidence or consequential claims | Domain Expert | Review becomes ceremonial |
Audit layer | Microsoft Purview / PostgreSQL | Stores claims, evidence and approval history | Governance | Decision lineage disappears |
2) Workflow
Ingest only approved source material with publisher, date and ownership metadata.
Generate research outputs as discrete claims instead of one opaque narrative.
Require every material factual claim to map to supporting evidence.
Automatically test references for existence, relevance and consistency.
Escalate unsupported, contradictory or low-confidence claims to a human reviewer.
Publish only after the evidence trail passes the defined threshold.
3) Example Prompts
Claim-Level Evidence Review
You are an evidence-integrity analyst.
Review the following policy brief.
Extract every consequential factual claim.
For each claim return:
- claim
- importance: LOW / MEDIUM / HIGH
- cited evidence
- whether the evidence directly supports the claim
- missing qualifications
- confidence level
Flag any claim that should not proceed without further verification.
Citation Verification
Audit the following citations and associated claims.
For every citation:
1. confirm whether the source exists
2. confirm author, title and publication details
3. determine whether the source actually supports the claim
4. identify any important context omitted
5. classify as VERIFIED / PARTIAL / CONTRADICTED / UNVERIFIED
Never invent replacement sources.
Decision Brief
Create a decision brief using only the verified evidence supplied.
Structure:
1. decision required
2. verified facts
3. strongest evidence
4. conflicting evidence
5. uncertainty
6. operational implications
7. recommended next action
Clearly separate evidence from inference.
4) Guardrails
Never allow a model to invent a missing citation.
Prefer primary sources for consequential claims.
Preserve original reports and datasets.
Distinguish fact, inference and recommendation.
Require claim-level support instead of bibliography-level support.
Escalate contradictory evidence instead of smoothing it away.
Require named accountability for high-impact outputs.
Store the model, prompt version and evidence used for major decisions.
5) Pilot Rollout — 3 hours
Select one recurring executive or policy-reporting workflow.
Load ten authoritative source documents into a controlled knowledge base.
Generate one report using retrieval from those sources only.
Extract every material factual claim and map each to supporting evidence.
Route unsupported claims to a named reviewer.
Compare accuracy and research time against the existing process.
6) Metrics
Percentage of claims with verified evidence
Unsupported claim rate
Fabricated citation rate
Citation mismatch rate
Human escalation rate
Research turnaround time
Post-publication correction rate
Primary-source utilisation
Average evidence confidence
Audit-trail completeness
Pro Tip: If the claim matters enough to influence policy, money or people, it matters enough to show the receipts.
🎯 The Arsenal — Tools & Platforms
Azure AI Search · retrieves grounded evidence from controlled knowledge sources · Link
Microsoft SharePoint · stores authoritative organisational reports and submissions · Link
Microsoft Purview · maintains governance, lineage and auditability · Link
Crossref · verifies academic publication metadata and identifiers · Link
Microsoft Teams Approvals · routes consequential claims to accountable human reviewers · Link
Copy-paste prompt block:
You are designing an AI evidence-verification system for an organisation producing policy, executive or strategic research.
Organisation:
[DESCRIPTION]
Research outputs:
[LIST]
Authoritative sources:
[LIST]
High-impact decisions:
[LIST]
Existing document systems:
[LIST]
Human reviewers:
[LIST]
The system must:
- ingest only approved evidence
- preserve original source metadata
- extract claim-level assertions
- verify citations and supporting evidence
- distinguish fact from inference
- detect unsupported or contradictory claims
- route uncertainty to qualified human reviewers
- maintain complete decision lineage
- prevent fabricated citations from reaching final outputs
Return:
1. architecture
2. evidence-ingestion standard
3. claim verification workflow
4. citation-validation process
5. human-review thresholds
6. publication gate
7. pilot rollout
8. operational metrics
💡 Free Office Hours
AI is becoming serious infrastructure across defence, government and the economy. That means evidence quality can’t remain an afterthought. The systems that matter most need to be able to explain not only what they concluded—but why anyone should believe them.
Book here: https://calendly.com
The best voice models now listen, adapt, and resolve too.
Most CX platforms don't own the voice. They orchestrate a workflow, then call a third party for speech and transcription. Every hop adds latency, and latency is what turns a frustrated customer into a churned one.
ElevenAgents is the opposite. Built on the voice models the market already builds on, it runs voice, transcription, chat, and reasoning in one vertically integrated pipeline. Responses come back in under 400 milliseconds and sound human, not synthetic. When a caller gets frustrated, the agent detects it and shifts tone in real time: calm, reassuring, patient.
You keep full control. Plug in any LLM, connect tools, webhooks, and MCP servers, and ground every answer in your knowledge base. Launch in minutes, A/B test with Experiments, enforce Guardrails, and version every change.
More resolved conversations, less infrastructure stitching. Pricing is transparent and flat at $0.08 per minute.
🕹️ Game Over
The Pentagon has ChatGPT.
Treasury wants the productivity gains.
Parliament apparently needs someone to check the bibliography.
Quite the ecosystem.
— Aaron Automating the boring. Amplifying the brilliant.
Subscribe: link

