0.9998
Qwen3-1.7B in TQ4. The top answer matched on five of five prompts.
Reference: 1.0000 bf16 cosineInternal model measurements, published technical notes, and methods that are not public yet.
ScopeEvery result names its model and comparison. None has been independently reproduced.
Five readings from current Torad runs.
Qwen3-1.7B in TQ4. The top answer matched on five of five prompts.
Reference: 1.0000 bf16 cosineQwen3-1.7B in TQ4.
Reference: 2.7 GB bf16Qwen3-1.7B.
Reference: 267 tok/s on vLLMQwen3-8B and Qwen3-14B on 16 GB.
Tokens per secondGemma 4 E2B on an RTX 5080.
16 GB availableThe blog carries experiment history, reasoning, and later corrections.
Why imitation and reinforcement produce different kinds of learning.
→Refusal as part of the system design.
→Transfer between structurally matched systems.
→Training a model to find an answer.
→These artifacts exist internally. Their full methods or evaluations are not public.