Last week I pointed an AI at a stock research report. I asked it to pull out the claims and tell me what would prove each one wrong.
Then I ran it again. Same file. Same settings. This was a simple LLM extraction job. Read a document, pull out the claims.
I got a different answer.
So I ran it four more times. Six runs in total. Here is what came back:
| Run | Claims found | Output tokens |
|---|---|---|
| 1 | 8 | 10,617 |
| 2 | 9 | 11,356 |
| 3 | 9 | 14,964 |
| 4 | 9 | 14,518 |
| 5 | 9 | 12,756 |
| 6 | 6 | 5,431 |
Same document every time. Temperature set to 0.2, which is low. I wrote about temperature settings a while back and described a low setting as making the model more deterministic. That is the standard explanation. It is what I expected here.
It did not hold.
It found between six and nine claims. And it used nearly three times more thinking on some runs than others.
That is the kind of thing you only notice if you run it more than once. Most people run it once, see output that looks fine, and ship.
Why I needed LLM extraction at all
Here is the problem I wanted to solve.
You read a research report. It convinces you. You buy the stock. Six months later you still own it. But do you still believe the reason you bought it?
Most of us cannot answer that. The reason was never written down. It sat in our heads as a feeling. And feelings drift.
So I am building a tool that writes the reason down. Not as a note. As a claim that reality can test.
This is module 1 of 7. It does two things. It reads a research document and pulls out the claims. Then it checks whether those claims are written in a way anyone could ever test.
The code is open source and it runs on your laptop. Nothing leaves your machine.
Kill criteria: the idea it all rests on
A kill criterion answers one question. What would have to happen for me to admit I was wrong?
That is the whole idea. Nothing more technical than that.
If you cannot answer it, you do not have a thesis. You have a hope. And hopes survive any amount of bad news.
Here is a bad one:
The stock drops 20%.
Useless. Stocks drop for a hundred reasons that have nothing to do with your claim. Wire your alarm to that and you will panic every few months for no reason.
Here is a good one:
Commercial segment revenue growth for the full year comes in below 45%.
A number. It has a source. And it has a date. In February you look it up and get a real answer.
A good kill criterion needs three things. A specific observation. Someone who publishes it on a schedule. And a date by which you would expect to know.
That middle one catches people out. The publisher does not have to be the company. A government agency counts, and so does a regulator. Even a rival company’s filing works.
Then I told it to be skeptical
Once the tool could pull out claims, I built a second tool to grade them. I wanted to know which criteria were actually testable.
In my first version I told the model to be tough. My exact words were “assume something is wrong until the criterion proves otherwise.”
It flagged 41 out of 45 criteria as broken.
That felt too high. So I changed one line. I swapped the tough instruction for a neutral one: “judge each criterion on its merits.” Nothing else changed. Same model, same criteria, same settings.
Now it flagged 36.
I expected a bigger drop. My theory was that the tough wording had caused most of the problems. It had not. The criteria really were weak.
But the breakdown is the interesting part.
| Type of flag | Tough wording | Neutral wording |
|---|---|---|
| “Nobody publishes this” | 11 | 12 |
| “Wrong number for the claim” | 6 | 7 |
| “Doesn’t settle anything” | 21 | 14 |
Look at the first two rows. They barely moved. Those are questions of fact. Does anyone publish this number? Is it the right number for the claim?
Now look at the third row. It dropped by a third. That one is a judgment call. Does this really settle the question?
So here is the finding. An AI’s factual judgments stay steady when you change your wording. Its opinions do not.
That is useful to know. When my tool says “nobody publishes this,” I believe it. When it says “this doesn’t settle anything,” I treat it as one opinion. It could have gone the other way.
The auditor made things up
This one bothered me most.
One of my claims was about US Navy shipbuilding. The tool killed it. Its reason:
It tests a non-existent program. There is no new nuclear-powered battleship in the Navy’s plan.
Clean logic. If the program does not exist, testing it is pointless.
The program exists. The Congressional Budget Office has costed it at roughly $275 billion for 15 ships, and it sits in the current shipbuilding plan.
It did it again on a different claim. It said the company does not report backlog by segment. The company reports backlog by segment every quarter.
Both times the reasoning was fine. The facts were invented.
And I fell for the first one. I read it, thought the reasoning was sharp, and told myself the tool had caught something I missed. It had not. It had made something up and argued well.
This is worse than a normal error. It came from a step I built to catch errors. When your error-checker sounds confident, you trust it more, not less.
The fix
Two changes. Both cheap.
First, I gave it a file of checked facts about each company. What segments they report. What they never break out. Now it looks things up instead of guessing.
Second, and this matters more, I made it declare its assumptions. Every fact it uses that I did not give it now gets printed on its own line:
ASSUMES: Project Janus is a real US government microreactor
contract whose award would be publicly announced.
You cannot check a hidden assumption. You can check one that is written down.
After that change, both fake facts vanished. The tool still killed both criteria. But now it gave real reasons instead of invented ones.
The known bug that still cost me an hour
My very first working call returned nothing at all.
tokens: 3790 in / 4000 out
Model did not return valid JSON: Expecting value: line 1 column 1
Status code 200. Tokens billed. Empty string.
The clue is that 4000. It was exactly my output limit. Suspiciously round numbers usually mean you got cut off.
Here is what happened. Modern models think before they answer. That thinking is hidden from you, but it eats the same budget as the answer. I gave it 4000 tokens. It spent all 4000 thinking. Then it stopped, before writing a single character.
This is a well-known problem. It is in the OpenAI forums, in a dozen GitHub issues, and in provider docs. I am not claiming to have found it.
I am telling you because of how it felt from the inside. Nothing failed. No error, no warning. Just a parsing error further down my own code, pointing at the wrong thing.
I only caught it because I was printing token counts. Without those numbers I would have spent the evening rewriting a prompt that was fine.
The lesson: when you cannot see inside a system, print every number it will give you. They are often the only evidence you get.
Twenty-two claims became six
Three research documents gave me 22 claims. Under a neutral audit, only 9 of 45 criteria passed. About 20%.
So I sat down and did it by hand.
Most of the 22 were the same uncertainty said twice. A bull case and a bear case arguing over one number is not two claims. It is one. If the number comes in at 40%, one of them dies. You do not need both.
I merged the duplicates. Anything untestable got cut. The vague ones got numbers.
Six claims left. The pass rate went from 20% to about 67%.
Then the audit caught two mistakes in my six. Both were mine.
One claim said cash flow should “land within the guided range.” My kill criterion only fired if cash flow came in low. So a great result would have made my claim false while my alarm stayed silent. Sloppy wording.
The other was worse. My claim said growth above 50% in each of two quarters. My criterion fired only if growth missed in both. Miss one quarter and nothing happens.
I fixed that. Then I found I had made the identical mistake in the very next claim. While I was actively hunting for it.
So here is the honest version. The human and the machine catch different things.
The machine cannot tell that a company publishes segment backlog. I can, in thirty seconds. But I will write “each” and “both” in the same breath and never see it. The machine catches that instantly.
Six ways a kill criterion breaks
This table is the most useful thing I built. Keep it if you find it handy.
| Problem | What is wrong | Example |
|---|---|---|
| Wrong signal | Measures the share price, not the business | “Stock falls behind its peers.” A stock moves for a hundred reasons that have nothing to do with your claim. |
| Nobody publishes it | No document, from anyone, on any schedule | “Management reports training is behind schedule.” No company publishes this. It only comes up if they choose to mention it. And if the news is bad, they won’t. |
| Wrong number | Real figure, wrong scope | Your claim is about one division. Your criterion measures the whole company. A weak division elsewhere can kill a claim that was right. |
| Settles nothing | Happens or not, you still don’t know | “The company wins Contract X.” It might win it and see no revenue for three years. |
| Wrong clock | Criterion is short, claim is long | Two quarters cannot prove a company has an advantage lasting a decade. |
| Too many claims | Several predictions in one sentence | “Growth slows and margins shrink and cash flow suffers.” No single test can kill it. |
The second one is the sneakiest. A missing criterion fails loudly and you notice. This one passes every check, looks careful, and then quietly never resolves. Two years later the claim is still marked “active.” Not because it held up. Because nothing could ever test it.
What it cost
About $1.50 in API charges. That covers every run, including all the failed ones and the six-times experiment.
Roughly one cent per document.
Money was never the hard part. Trusting the output was.
The honest limitation
Look at what survived my cut. Almost every claim boils down to the same shape: does the company hit its own forecast?
That is not an accident. Company guidance is the only forward-looking number published on a fixed schedule. So it becomes the anchor for anything testable.
Which means the big claims got cut. Does this company have a lasting advantage? Is management any good? Is this a decade-long trend? Those are the questions that actually make you money. None of them reduce to a number in a filing.
The tool tests what is testable. Not what matters most.
That is a real cost. I am not going to pretend otherwise.
If you are building something similar
Four things I would tell you.
Run it more than once. Same input, five times. Compare. If you had asked me before this project whether that was worth doing, I would have said no.
Print every number the API gives you. Tokens in, tokens out, why it stopped. It costs nothing and it is how you find the invisible failures.
Make it declare what it assumed. A fact you can see is a fact you can check.
Watch your own wording. Being told to be critical made my model more critical. That sounds obvious written down. It did not feel obvious while I was reading its very reasonable-sounding output.
What is next
Right now this tool never touches the internet. It reads a file you give it, and that is all. Because of that, it never checks a filing, never reads the news, and never tells you whether your claim is actually true.
It only tells you whether your claim is the kind of thing that could be tested. That is a smaller job than it sounds, and it has to come first. There is no point fetching evidence for a claim nobody can ever measure.

Update: Module 2 is now live. It adds the other half: a job that runs every morning, reads new SEC filings, and writes me a short note about what changed.
Most mornings it should say nothing. That is the point.
Get the code
Everything here is open source under an MIT license. It runs on your laptop. Your documents never leave your machine.
github.com/psachdev/thesis-drift-monitor
There is demo data included, built from public sources, so you can try it without a paid subscription. Two of the four demo claims are broken on purpose. See if the tool catches them.
Questions people have asked
Do I need to know Python?
A little. You need to install Python, run a few commands, and edit a text file. If you have never opened a terminal, this will be a stretch. There is also a version that works inside a chat window with no code at all, included in the repo.
Does it only work with DeepSeek?
No. DeepSeek is the default because it is cheap. Anything that speaks the OpenAI API format works by changing four settings. No code changes needed.
Why not just set the temperature to zero?
It helps a bit. It does not fix it. Even at zero, most providers do not promise identical output. The math behind the scenes has small wobbles, and on a long task those wobbles compound into different answers. If you want the background on what temperature actually does, I covered it in an earlier post. Running your job several times is a more honest test than trusting one setting.
Can I use it for any company?
Yes. There is a command that drafts a profile for any ticker. But it drafts from memory, so you have to check it against a real filing before you trust it. That takes about five minutes and it is the single best thing you can do for output quality.
Is this investment advice?
No. It is a piece of software and a write-up of building it. The companies mentioned are ones I follow. The point is the tool, not the picks. Do your own research and speak to a licensed professional before making decisions with your money.
About the author: I’m a software engineer in San Francisco and the creator of AIBuddy. I write about AI tools, agents, and practical experiments where technology meets everyday decisions. This is module 1 of a 7-part series building an agentic AI system in the open. Next: Module 2, Segment Revenue From SEC Filings. The full series is at funaibuddy.com/agentic-ai.







Leave a Reply