Empirical Limits of Cross-Architecture Reasoning Transfer in Sub-8B Student LLMs
A comprehensive investigation into attention map divergence, reasoning token entropy, and layer-to-layer distillation across disparate transformer architectures.
Frontier models are locked behind proprietary APIs. Datacenters cannot sustain brute-force scale. Solstice-AI builds open distillation infrastructure, verified multi-teacher datasets, and efficient sub-8B weights to solve both.
The largest verified frontier-model distillation corpus ever released. 12,586,893 unique conversations synthesized from 60 audited source datasets across 7 frontier architectures. Ingesting 16.03M rows, exact SHA256 deduplication removed 3.44M redundant copies (21.5% overlap) into a single drop-in JSONL archive.
12,586,893 unique conversations synthesized from 60 audited sources across 7 frontier architectures
Eliminating cross-dataset contamination and scraping redundancy
Memory footprint at 32k context on NVIDIA RTX 4090 / Apple M4 Max
1,489+ community downloads across MLX and GGUF quantizations
We provide every layer of the post-training distillation stack as open, reproducible artifacts.
Uncensored chain-of-thought, code refactoring transcripts, and agentic sandbox tool invocations. All datasets are partitioned in Apache Parquet with full metadata attribution.
Activation-calibrated student weights available in FP8, INT4 AWQ, and GGUF formats. Engineered to deliver near-lossless reasoning recovery on consumer devices.
High-performance inference engines integrating Google TurboQuant FWHT-rotated KV cache compression (4.6x reduction) with Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).
Peer-reviewed evaluation methodologies measuring reasoning degradation, KV cache optimization, and cross-architecture distillation transfer.
A comprehensive investigation into attention map divergence, reasoning token entropy, and layer-to-layer distillation across disparate transformer architectures.
How Google's TurboQuant Fast Walsh-Hadamard Transform and our Anvil engine deliver 4.6x KV cache compression with +30-50% throughput acceleration.
Empirical benchmarking of Solace-Sub8B checkpoints in INT4 AWQ, FP8, and GGUF across Apple Silicon and NVIDIA Ada Lovelace GPUs.