Research
Original Data
Benchmarks and findings on AI visibility, WebMCP, and agentic readiness — measured, not asserted.
The Teacher Sets the Ceiling: What Distilling an LLM Classifier Into a Small Model Actually Buys You
An experiment replacing a mid-size generalist LLM classifier with a fine-tuned small model. Model size, data volume and constrained decoding made no measurable difference. Changing who wrote the training labels made the largest one. And without ground truth, every metric turned out to measure agreement with a judge, not correctness.
Published 2026-09-28
WebMCP Readiness Is a Property of the Site, Not the Widget
A controlled three-site benchmark of agentic readiness. The same AI commerce widget was run against porsync.com and two companies that sell WebMCP and agentic-commerce tooling. The site with declarative WebMCP attributes succeeded; the two without them failed — including the vendor's own homepage.
Published 2026-05-28
Before You Fine-Tune: A Quant-Retention Framework for Edge SLM Feasibility
Most edge-AI projects start by training a model and end by discovering it never fit the target device. This is the inverted method: a measurement-first framework that answers 'can a small model do this job on this hardware at all?' before a single training run — using quant-retention ratios, behavioural adherence rates, and a hardware envelope as the go/no-go gate.
Published 2026-07-05