Nine feeling detectors, all above 0.8. Four days of grinding and three numbers I had to kill.
Four days ago the Feel Layer had three validated emotion detectors and one number I was proud of that turned out to be fabricated. Tonight all nine are locked: joy, interest, affection, desire, sadness, anger, fear, disgust and surprise, each at a measured precision between 0.86 and 0.97, each judged blind by a human with trap texts mixed in. This is the log of what that took, and which parts were mine.
The rule that made it slow, and real
Early in the week I set one rule: nothing locks below 80 percent measured precision, and the measurement is a blind read. Not the model's confidence, not a crowd vote. A person reads the detector's proposals without knowing which answers it believed, traps included, and the number is whatever the person says.
That rule cost us. A paid crowd had certified our desire detector at 97 percent. My own ten minute blind read showed the raters saying yes to nearly everything, including the traps. The real number was 38. We killed our best result. Three days and three retraining rounds later, desire locked at 93 on texts neither of us had seen. The rule also revoked a lock I had already celebrated: fear passed at home and then failed its proof reading in new territory, so its approval was withdrawn and re-earned with a narrower, published scope.
What was mine
The definitions. Every detector that converged did so after a one sentence boundary ruling: desire is wanting what you do not yet have. Blame is anger, missing out is sadness. A scary topic is not fear unless the writer lived it. My eleven year old contributed the desire rule's core insight last month. The machines could not learn a target that kept moving, and the target was mine to fix.
The judging. About 700 blind swipes on my phone over four days, roughly ten minutes at a time. Every locked number traces back to those reads.
The kills. Deciding the 97 was fake, deciding fear's lock had to be revoked, deciding a corpus was empty of the feeling we were fishing for. Agents surfaced the evidence; the calls were mine.
What was not mine
Everything else. The training runs, the scanning, the deck building, the watchers that pinged my phone when something needed my eyes, the pages that update themselves. An agent ran the whole loop while I slept, and the honest accounting is that I supplied maybe an hour a day of judgment to its twenty three hours of grinding.
The test I did not expect to run
Once all nine were locked we put the newest models from the big labs on the same 48 text benchmark, same prompt, same scoring: Gemini 3.1 Pro scored 0.50, GPT 5.6 scored 0.50, Kimi K3 0.48, Claude Opus 0.42. Our fine tuned 14B, running on the Mac on my desk, scored 0.55. The gap is not intelligence. It is calibration: the boundaries live in labeled data that nobody can zero-shot their way into.
What it is for
Feed the nine detectors every review of a game, a product or an artist and you get its feeling fingerprint: what it makes people feel, at what rate, with receipts. Spiritfarer: affection in 80 percent of reviews. And the first commercial signal is in: within the same game, reviews showing affection are 19 points more likely to recommend it. Ten games out of ten.
Next: fingerprinting the top thousand games on Steam, then the app niches everyone actually cares about. The instrument is built. Now it gets pointed at things.