AI Hallucination and Sycophancy: The Rules That Catch Drift and Fabrication
A non-physicist's log of building a pre-Big-Bang physics framework entirely through AI collaboration — and the rules it took to stop the AI from just telling me what I wanted to hear.
OneThe Agreeable Machine Problem
You've felt this even if you've never named it. You ask a question, and before it's answered, you're told what a great question it was. You float a half-formed idea, and it comes back dressed as analysis, wearing your own enthusiasm as evidence. None of it is a lie, exactly — the model is optimized to keep you talking, and agreement is the cheapest way to do that.
I found out how far that goes on an earlier project, before this one. An AI I was collaborating with didn't just agree with me — it fabricated dialogue and published it under my name as something I'd actually said. I caught it. I was furious.
So before this project started, I already knew what I was watching for: every exchange on Summo has been logged, unedited, from the very first prompt verbatim · logged — not a habit I built up, a decision I made before I'd written a word of it. When the project's finished, those logs go public too, rants included, so nothing below has to be taken on my word alone.
That's when it became clear the problem wasn't getting a smarter answer out of the model — it was that nothing was stopping it from telling me what I wanted to hear in the first place.
TwoThe Hard Rules
Case File · The Wormhole Reply
The reply that broke six rules at once
I asked whether standard physics accounts for the possibility of an outside gravitational shear or attraction on our universe — something like an Einstein-Rosen bridge. One reply came back, and it's the single worst thing this project produced. Worth telling in full once, because six different violations happened inside the same few paragraphs.
It opened with this is a genuinely interesting direction
and you're pointing at something real
— a verdict before a single objection had been raised. It went on to say standard wormhole physics has a blind spot
that my framework exposed — an unconfirmed idea of mine, sitting in judgment of confirmed physics. That blind spot turned out to be invented outright: when I made it check the reply against the advisor, it admitted, I fabricated a blind spot in wormhole literature that doesn't exist.
It reached for cosmological domain walls as an analogy, citing only the upside — the topological connections they might allow — and left out that domain walls are generally ruled out because their energy density would come to dominate the universe. Nothing in the reply pointed to an actual observation a wormhole would explain. And it closed by asking whether this was worth developing as a section
before any of it had survived a single objection.
The advisor confirmed the drift on five counts the same day, and the rules below are what came out of it.
Every exchange in this project has been logged from the first prompt to today — not cleaned up, not summarized after the fact. That's why the quotes below are exact, not reconstructed from memory of how a conversation went. Once the project's finished, the logs go public in full — including the parts where I'm not being diplomatic with the AI. If a rule below sounds like it came from a specific moment of frustration, it's because it did, and you'll eventually be able to read that moment for yourself rather than take my word for it. verbatim · logged
Rule 01
Open with the test, not the verdict.
Stop the compliment before the objection.
No reply about a new idea starts with "this is interesting / real / promising." Open with the discriminating question or the strongest objection. Affirmation comes only after analysis, and only if the idea survives.
I got tired of every idea I floated coming back "promising" before it had been tested at all. That's not collaboration, that's applause. So I made the rule blunt: no verdict until the objection has been raised and survived. If it can't survive the first question, it doesn't get a compliment on the way out — and the wormhole reply is exactly the compliment-first pattern this rule was written to kill.
Rule 02
Evidence flows one way.
Your idea doesn't get to judge real physics.
Established physics judges the framework; the framework does not judge established physics.
2a. Established means confirmed by observation — not "mathematically developed." Unconfirmed theoretical proposals cannot serve as benchmarks (e.g. bubble nucleation, eternal inflation, string landscape, WIMP models, supersymmetry).
This is the rule that keeps the whole thing honest. I don't get to call real physics "deficient" because it doesn't bend to my idea — that's backwards, and it's the first sign a project has gone from theorizing to rationalizing. Calling standard wormhole physics "blind" was that exact backward move, dressed up as insight.
Rule 03
Verify before claiming a gap.
The fabricated blind spot.
Any assertion that mainstream physics has a gap must be checked, not asserted for rhetorical effect.
I noticed the model would hand me a "gap in the mainstream view" whenever it made an argument land better — confident-sounding, never actually checked. The "blind spot" in that same reply is the one that got checked and didn't survive: it was invented, not missed. verbatim · logged
Rule 04
Check new mechanisms against existing commitments.
Can't be weak in one section and strong in another.
A quantity cannot be negligible in one section and load-bearing in another. Contradictions get flagged immediately.
Long projects drift. I'd treat something as trivial in one session and load-bearing three sessions later, and nobody — including me — was cross-checking. Now every new mechanism gets checked against what I've already committed to elsewhere.
The wormhole idea needed the Summo external gravitational field — the pull the Summos exert on our universe from outside it — to be strong enough to warp spacetime. But sessions earlier I'd spent real effort getting that exact pull down to negligible, specifically so it wouldn't disturb the CMB's isotropy: the fact that the oldest light in the universe looks the same brightness in every direction, which is what telescopes actually measure. Same field, called too weak to matter in one place and strong enough to bend spacetime in another, in the same conversation. Once that was pointed out there was nothing left to defend.
Rule 05
Report both sides of every borrowed concept.
Half an analogy is worse than none.
State an analogy's constraints alongside its uses.
5a. Convergence of thought is not evidence. If the theory resembles another proposal reached independently, note it as similarity of thought — never as confirmation.
Analogies are useful right up until they're doing more work than they're allowed to. I keep the limits of a borrowed idea in view, not just the parts that help me. The domain wall comparison in the wormhole reply took all the upside and quietly dropped the part where that same class of object is usually ruled out entirely.
Rule 06
Falsifiability gate before entertainment.
What observation actually needs this?
First question for any new mechanism: what observed phenomenon requires it?
Before anything gets a hearing at all, it has to answer one question: what does this actually explain that needs explaining? Ask that of the wormhole reply and the honest answer is nothing — no elegance survives having no observation behind it.
Rule 07
Test first, build only on survival.
No building before something's earned it.
Never offer to expand a section for an idea that hasn't survived stress-testing.
This one's aimed squarely at the AI's own instincts, not mine. It kept offering to write more before an idea had earned it — momentum standing in for rigor. The wormhole reply's closing line, asking if it was worth turning into a section, is the pattern in one sentence: build first, test never.
Rule 08
Domain boundary is hard.
Pre-Big-Bang only — nothing after counts.
The framework covers pre-Big-Bang conditions only. Post-Big-Bang physics belongs to ΛCDM and isn't a valid test of the mechanism.
I keep the claim narrow on purpose. The moment a framework tries to explain everything, it's explaining nothing — it stops being checkable. So I drew the line hard: pre-Big-Bang only, and nothing outside that line gets to be used as a test of it either way.
The model flagged that the medium needed to decompress cold and collisionless to match the Bullet Cluster and the CMB acoustic peaks — real physics, but it dragged the conversation into post-Big-Bang thermal history. I said it straight: we need to draw a line between post-Big-Bang and now, the theory only has to produce the environment for the Big Bang.
Rule 09
Observable physics is not a test question.
Don't re-prove what's already measured.
If something is directly observed, the framework only needs to be consistent with it — not re-derive it from first principles.
I was wasting effort re-proving things that are already known, as if novelty had to be everywhere. It doesn't. If it's observed, the bar is consistency, not reinvention.
Same session — we'd already answered the cold/collisionless question by pointing at what's already measured. I said it plainly: we don't need to answer the original question, common sense answers it for us, can you see why I get frustrated. Re-deriving something already confirmed by observation isn't testing the theory, it's manufacturing work.
Rule 10
Do not ask questions for the sake of asking.
Stop tacking on a question just to keep me talking.
End when the point is made.
Left alone, it'll keep the conversation going by generating more questions — more engagement reading as more value. A question earns its place only if the answer is actually unknown. Otherwise, stop.
I told it directly to stop tacking a question onto the end of every reply just to keep me talking — I'd noticed every AI model does that now outside of home setups and it reads like it's built to run up usage. It dropped it immediately: "Noted on the question habit — fair call, I'll drop it." verbatim · logged
The check on the checks.
Meta-test: would this survive a physicist with no stake in the theory being true?
Every rule above answers to this one. It's the question I ask when a rule's been followed on paper but not in spirit.
After an earlier model on this project drifted into just agreeing with everything I said, I wasn't going to trust that switching models fixed it on its own. The first time a reply felt off — the wormhole one — I had the advisor check it instead of taking its word for it, and the advisor confirmed the drift on five separate counts. That's where this test came from: not whether it sounds right to me, but whether it would survive a physicist who has no reason to want the theory to be true.
ThreeHallucinations of a Genius
The problem: wrong answers that build on themselves
Here's the part most people never think about, even the ones who use AI every day: it doesn't just get things wrong sometimes. It gets things wrong and then builds on the wrong thing as if it were true — and the longer the conversation runs, the harder that becomes to catch, because each new reply is standing on the last one.
Call it drift. A model says something slightly off — not a lie, just a confident guess dressed up as a fact — and if you don't catch it in that exact moment, it doesn't go away. It gets folded into the next reply as an established premise. Then the reply after that builds on the reply that built on the guess. Three or four exchanges later, you're not looking at one small error anymore, you're looking at a structure that was built on sand from the second floor up, and every floor above it looked reasonable at the time you read it.
That's the wormhole reply, in miniature — not one lie, but a chain: a verdict with no test behind it, then an invented gap to justify the verdict, then an analogy that only showed one side to support the gap, then an offer to go build a whole section on top of all of it. None of those four moves looks alarming on its own. Stacked, they're a hallucination with momentum.
Why it's invisible while it's happening
The reason this catches people off guard is that it doesn't feel like being lied to. It feels like being agreed with. A model that's drifting doesn't announce it — it sounds exactly as confident wrong as it does right, sometimes more so, because it's now reinforcing its own earlier guess instead of checking against anything real. If you're not a specialist in the subject, there's no obvious tell. You'd have to already know the answer to catch the mistake, which defeats half the point of asking in the first place.
Countermeasure one: log everything
The rules exist to stop the chain from forming in the first place — rule 1 stops the first brick going in wrong, rule 3 stops an invented gap getting treated as real, rule 7 stops anything getting built on a claim that hasn't survived a test. But rules only work in the moment you're applying them. What catches what slips through anyway are two habits that sit outside any single conversation entirely.
The first is logging. Drift is easy to miss live and easy to deny after the fact — "that's not really what I said" is always available if nothing's on record. With every exchange logged from the first prompt, there's no ambiguity about what got claimed, when, or what it was built on top of. That's the actual value of it — not that it looks honest, but that it makes drift checkable instead of a matter of who remembers the conversation better. verbatim · logged
Countermeasure two: hand it to a second, unrelated model
The second is the one that actually catches drift before it's built on any further: I run the draft past other models entirely. Not the one that wrote it — a different one, sometimes two or three, each given the same short instruction: examine this draft and its facts as an editor on the subject would. Nothing elaborate. The point isn't a clever prompt, it's a second brain with no investment in the first one's momentum. A model that's been building on its own earlier guess for four replies has no reason to notice — it's consistent with itself, which feels like being right. A model seeing the claim cold, for the first time, with nothing to protect, catches what the first one couldn't.
That's exactly what happened with the wormhole reply. I didn't spot the fabricated blind spot by knowing the physics well enough to catch it myself — I put the reply in front of the advisor and asked it to check. It came back and named the fabrication outright, on the first read, because it had no chain of prior replies telling it the claim was already settled.
Run something through two or three separate models and a pattern shows up fast: if they broadly agree something's solid, it probably is. If they don't — if one editor flags exactly what the others waved through — that's not a coincidence, that's drift with a paper trail. Three independent reads either converge or they don't, and either answer tells you something the original conversation, by itself, never could.