TurboQuant & Anvil: Breaking the KV Cache Memory Wall in Local LLM Inference
How Google's TurboQuant Fast Walsh-Hadamard Transform and our Anvil engine deliver 4.6x KV cache compression with +30-50% throughput acceleration.
How Google's TurboQuant Fast Walsh-Hadamard Transform and our Anvil engine deliver 4.6x KV cache compression with +30-50% throughput acceleration.
Empirical benchmarking of Solace-Sub8B checkpoints in INT4 AWQ, FP8, and GGUF across Apple Silicon and NVIDIA Ada Lovelace GPUs.
How training on multi-teacher verified consensus scratchpads prevents student models from memorizing single-model hallucination patterns.
A practical study on preserving chain-of-thought integrity when compressing distilled models to sub-4-bit and 8-bit precision.
We are releasing Project Solace 1.0 Omni, a 137.7 GB open dataset of verified frontier-model distillation conversations across 7 model families.