Blog
Notes from building products with AI agents.
My AI agents are not allowed to say my model graduated. I wrote the rule that stops them.
Two months, 1,512 commits, one operating system for agent work. What it looks like when the machinery that runs your project is also allowed to refuse you.
Nine feeling detectors, all above 0.8. Four days of grinding and three numbers I had to kill.
The Feel Layer's nine emotion detectors are locked at measured precisions between 0.86 and 0.97, each judged blind. On the way we killed a fake 97 percent, revoked one of our own locks, and beat the newest models from four AI labs with a 14B on my desk.
Everyone failed my exam. That was the result I needed.
A build log from my feeling-layer project: a small fine-tuned model now beats the standard emotion classifier threefold on real reviews, and a qualification exam for human raters revealed exactly why ground truth is the hard part.
A frontier model aced my exam, then lost the real task to my small one
The same evening, the same frontier model passed my annotator qualification exam near-perfectly and lost the labeling benchmark to my locally fine-tuned model by more than double. Both results are true, and together they explain what this project actually measures.
Cosiness did not make the list. Why my instrument reads exactly nine feelings.
The design rationale for the feeling layer's taxonomy: primary emotional systems from affective neuroscience, a strict measurability rule, and why the subtle feelings everyone actually wants are deliberately composed rather than asked for.
The evidence gate: why a third of my hotels don't get a quiet score
A quiet-hotels register needs at least 3 genuine quiet or noise mentions per hotel to get scored. No exceptions. Here's what that rule cost and what it bought.
I taught a model to predict 'would buy again' and then checked it against 1,473 real answers
Steam reviews carry a real thumb: recommended or not. I deduced the same answer from review text alone, then checked the two against each other.
The pipeline that caught 386 wrong booking links
We checked every booking link on the site by actually visiting it. 386 out of 6,344 pointed somewhere wrong.
Can a free 14B model replace a frontier model? A €0 bakeoff
I tested a free local model against the frontier model already scoring my hotels. It lost, then it won, once I gave it more to read.