ResearchWorkSystemsAboutTalk shop
Research/Serving Choices Change Speed, Not Accuracy: Measuring a Local 9B Model on One Consumer GPU
Local LLMsInference ServingQuantizationCost AnalysisBenchmarkingOriginal Research

Serving Choices Change Speed, Not Accuracy: Measuring a Local 9B Model on One Consumer GPU

A pre-registered study of what serving decisions (quantization, inference stack, concurrency, request shape, structured output, prompt caching, KV cache type) cost and buy when a 9B model classifies 500 support tickets on a single 16 GB GPU. Speed varied by 2.2x between the default and the best setup and electricity cost stayed under a cent per thousand tickets, while accuracy did not measurably move. Individual answers did.

Porsync Research · Published 2026-10-06

Finding

On one 16 GB GPU classifying 500 tickets with Qwen3.5-9B, the best serving configuration (llama-server, 4-bit, 4 concurrent tickets) ran 2.2x faster than the Ollama default (178 vs 80 tickets/min) at $0.0032 per 1,000 tickets in electricity, against $22.43 per 1,000 for Claude on the same task, with no measurable change in accuracy; but 4-bit quantization changed 8% of queue answers relative to 8-bit.

Local LLM serving-configuration study, RTX 5060 Ti 16 GB (Oct 2026)

26 serving configurations, each run 3 times on 500 support tickets (3 questions each), Qwen3.5-9B, temperature 0, thinking off, one fixed prompt. Throughput, GPU board energy, answer flips against an 8-bit reference, and cost per 1,000 tickets. Scored by a fixed script against a pre-registered protocol.

MeasurementValueNote
Ollama default (Q4, c=1)80.3 tickets/min65.7 J per ticket
Best config (llama-server, Q4, c=4)178.4 tickets/min2.22x the default; 44.6 J per ticket
Same file, c=1: llama-server vs Ollama142.1 vs 92.1 tickets/min1.54x
Answers changed by 4-bit vs 8-bit (queue / priority)8.0% / 5.2%accuracy differences within 1.4 points, intervals spanning zero
One call returning all 3 answers vs 3 calls99.1 vs 142.1 tickets/min30% slower, and changed 16-19% of queue and priority answers
Marginal electricity cost, best config$0.0032 per 1,000 ticketsPHP 15/kWh; all-in $0.005-0.009 with the GPU amortised (assumed 4-year life)
Claude, same task, 1 call per ticket$22.43 per 1,000 ticketsAPI-equivalent cost of a plan-billed run

The question

A small local model can replace a paid API for high-volume classification, but the serving choices around it are usually picked by default or folklore: which quantization, which server, how many requests in flight, one call or three, whether to constrain the output. On a single consumer GPU I could not find a measurement of what each one costs in accuracy or buys in throughput on a real workload. So I wrote the protocol down first and measured.

The task is deliberately plain: classify 500 customer-support tickets on three questions (a 10-way queue, a 3-way priority, and whether it is an incident) with Qwen3.5-9B on an RTX 5060 Ti with 16 GB. The model, prompt, sampling and sample are fixed; only the serving configuration changes. This is not a model-quality contest. Every configuration, and Claude too, lands below a plain TF-IDF classifier on the 10-way queue question (0.29-0.34 against 0.418), so the labels are noisy and absolute accuracy says little about the model.

What moved, and what did not

Across 26 configurations, each run three times, accuracy stayed in a narrow band: queue 0.29-0.34, priority 0.40-0.41, incident 0.75-0.76. Paired intervals against the 8-bit reference span zero for every configuration except the vLLM reference rows. Speed varied by a factor of more than ten.

What did move is **individual answers**. Going from 8-bit to 4-bit changed 8.0% of queue answers and 5.2% of priority answers, without a measurable change in accuracy. Two configurations can have the same accuracy and disagree on many tickets, which matters if anyone downstream audits or diffs results.

ConfigurationTickets/minJ per ticketQueue / priority answers changed vs 8-bit
Ollama default (4-bit, one at a time)80.365.712.4% / 6.6%
llama-server 8-bit (reference)128.658.8-
llama-server 4-bit142.155.08.0% / 5.2%
llama-server 4-bit, 4 tickets in flight178.444.67.8% / 5.4%
Ollama, same file as the 4-bit row92.164.49.4% / 5.8%
llama-server BF16 (spills to CPU)48.694.70.8% / 1.8%
One call per ticket instead of three99.190.119.2% / 16.4%

Findings, one factor at a time

- **Quantization.** Each step up in bits costs 4-6% in speed and removes answer drift: 5-bit flips 3.4% / 3.8%, BF16 under 2%. BF16 does not fit in 16 GB, spills to the CPU, and runs 2.6x slower than 8-bit. The same trap showed up in an earlier experiment, where a model that did not fit ran 3.6x slower.
- **Stack.** On an identical model file at one request at a time, llama-server ran 1.54x faster than Ollama.
- **Concurrency.** llama-server gains 20% at two tickets in flight and 5% more at four, then slips. Ollama was flat from two upward, and its log says why: for this model architecture it reports that parallel requests are not supported and loads with a single slot. That is a property of one Ollama version and one architecture, not of Ollama in general.
- **One call instead of three.** Asking for all three answers as JSON in a single call was 30% slower per ticket, not faster, and changed 16-19% of answers. Output tokens per second is higher, but the single call decodes the whole reply, so the ticket takes longer.
- **Structured output and prompt caching.** A grammar constraint was free and produced identical output to the unconstrained run row for row, since the model already emitted valid options at temperature 0. Prompt caching made no difference: nothing was reused across the three calls on this model.
- **KV cache type.** An 8-bit KV cache changed neither speed nor answers beyond noise.
- **A reference point.** vLLM under WSL2 with FP8 reached 522 tickets/min at 17 J per ticket. It is a different quantization on a different stack and OS, and it shifted queue accuracy by -1.8 points (interval -3.2 to -0.4), so it is a headroom marker, not a like-for-like comparison.

Cost

Electricity for the best configuration is **$0.0032 per 1,000 tickets** (PHP 15/kWh), and all-in $0.005-0.009 once the GPU is amortised over an assumed four years at the two duty cycles I tried. Claude on the same task, one call per ticket, is **$22.43 per 1,000**, the API-equivalent cost of a plan-billed run. Buying the GPU for this job pays back at about 430 tickets a month, and that figure comes from the GPU price against Claude's per-ticket cost, because electricity is negligible. It is cost per decision, not per correct decision: both land below the simple baseline on this task.

A caveat in the decision rule

I pre-registered a "no answer change" rule: a configuration passes if its flip rate against the reference is at most twice its run-to-run noise floor (minimum 2%) and the accuracy interval does not fall below -3 points. The best configuration passed, but an identical earlier run of the same settings did not, because the allowance scales with measured noise and batching adds nondeterminism (2-5% run to run, against exactly zero when one request is in flight). The rule rewards noise. I report it as written; the sturdier reading is that 4-bit quantization, not concurrency, causes most of the drift.

Limits

- One GPU, one model family, one task, one 500-ticket sample, one seed.
- Outputs are a few tokens, so this workload is dominated by prompt processing and request overhead; generation-heavy jobs would rank configurations differently.
- Closed-loop load at fixed concurrency. Tail latency under bursty traffic was not measured.
- Energy is GPU board power only; the BF16 CPU spill energy is real and unmeasured.
- The exchange rate, four-year life and duty cycles in the cost section are my assumptions.

What I would take from it

1. **Measure before accepting defaults.** The default setup was the slowest configuration that fit on the card and was not offloading, and the fix was a server swap, not new hardware.
2. **Do not trust accuracy alone to say a change is safe.** Same accuracy, 8% different answers.
3. **Check what a feature actually does on your model.** Caching, parallel slots and single-call prompting all behaved differently from the folklore on this architecture.
4. **Pre-register the rule, then read where it is odd.** The pass/fail gate was gamed by noise, and writing that down is more useful than hiding it.

The protocol, raw per-run files, scorer and figures are in the project repository.