Skip to main content

Distilling
intelligence

Frontier models are locked behind proprietary APIs. Datacenters cannot sustain brute-force scale. Solstice-AI builds open distillation infrastructure, verified multi-teacher datasets, and efficient sub-8B weights to solve both.

12.59MUnique Conversations
137.7 GBCorpus Size
60 Sources7 Frontier Families
AGPL-3.0Open Access
FLAGSHIP DATASET • MAY – AUG 202612,586,893 EXAMPLES

Project Solace 1.0 Omni

The largest verified frontier-model distillation corpus ever released. 12,586,893 unique conversations synthesized from 60 audited source datasets across 7 frontier architectures. Ingesting 16.03M rows, exact SHA256 deduplication removed 3.44M redundant copies (21.5% overlap) into a single drop-in JSONL archive.

MULTI-TEACHER CONSENSUS PIPELINE• 60 SOURCES AUDITED
01
GLM 5.2Zhipu AI
OpenHands Agent Rollouts & Regen
VERIFIED
02
Claude Fable 5Anthropic
Agentic Coding & CoT Traces
VERIFIED
03
GPT-5.6 SolOpenAI
Coding & ARC-AGI3 Challenges
VERIFIED
04
GPT-5.5 CodexOpenAI
Code Specialization Streams
VERIFIED
05
DeepSeek V4 ProDeepSeek
200K Math/STEM & SWE-bench
VERIFIED
06
Qwen 3.8 MaxAlibaba Cloud
Cross-Architecture Distillation
VERIFIED
07
Kimi K3 / Opus 4.7Moonshot / Anthropic
Multi-Teacher Reasoning
VERIFIED
08
Manus AgentsManus AI
Tool Sandbox Execution Logs
VERIFIED
DISTILLATION OUTPUT • PROJECT SOLACE (12.59M ROWS)
Exact SHA256 Dedup (3.45M copies purged) • 137.7 GB JSONL (34.3 GB gz) • Weighted Round-Robin Shuffled
Supported Frameworks & Environments:
Anvil (llama.cpp TurboQuant)TurboQuant MLX (Metal)vLLM (FP8)Hugging Face TRLAxolotl

Project Solace 1.0 Omni Composition

12,586,893 unique conversations synthesized from 60 audited sources across 7 frontier architectures

Qwen 3.8-Max & Multi-Teacher Reasoning
4,342,000 (34.5%)
GLM 5.2 (15 Sources • Agent Rollouts & Regen)
2,380,000 (18.9%)
GPT-5.6 Sol & GPT-5.5 Codex (Code & ARC-AGI3)
2,120,000 (16.8%)
Claude Fable 5 & Mythos (Clean Agent Traces)
2,006,487 (16.0%)
DeepSeek V4 Pro 0813 (Math/STEM & SWE)
1,738,406 (13.8%)

Exact SHA256 Deduplication Yield

Eliminating cross-dataset contamination and scraping redundancy

Raw Ingested Dataset Rows (60 Sources)
16,032,427 (100%)
Exact SHA256 Duplicates Purged
3,445,534 (21.5% Redundant)
Final Clean Solace Yield
12,586,893 Unique Conversations (78.5%)

Google TurboQuant KV Cache Compression (Anvil)

Memory footprint at 32k context on NVIDIA RTX 4090 / Apple M4 Max

Standard FP16 KV Cache (Vanilla llama.cpp)
16.4 GB (42.1 tok/s)
Naive INT4 RTN (Perplexity Degraded)
4.8 GB (Degraded Perplexity: 14.22)
TurboQuant 4-bit FWHT (Anvil Engine)
4.2 GB (-74% VRAM • 61.8 tok/s / +46.8% Throughput)

Published Checkpoints on Hugging Face Hub

1,489+ community downloads across MLX and GGUF quantizations

Qwen3.8-27B-Uncensored-mlx-6Bit
790 Downloads
Qwen3.8-27B-Cold-Fusion-GAIN-mlx-6Bit
324 Downloads
Huihui-Qwen3.6-35B-Opus-GGUF
100 Downloads
ThinkingCap-Qwen3.6-27B-mlx-6Bit
79 Downloads
Qwopus3.6-27B-Coder-mlx-6Bit
48 Downloads
Architecture & Tooling

Open Infrastructure Stack

We provide every layer of the post-training distillation stack as open, reproducible artifacts.

01

Multi-Teacher Reasoning Corpora

Uncensored chain-of-thought, code refactoring transcripts, and agentic sandbox tool invocations. All datasets are partitioned in Apache Parquet with full metadata attribution.

02

Sub-8B Quantized Student Models

Activation-calibrated student weights available in FP8, INT4 AWQ, and GGUF formats. Engineered to deliver near-lossless reasoning recovery on consumer devices.

03

Anvil & TurboQuant Inference Engines

High-performance inference engines integrating Google TurboQuant FWHT-rotated KV cache compression (4.6x reduction) with Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).

04

Empirical Research & Technical Reports

Peer-reviewed evaluation methodologies measuring reasoning degradation, KV cache optimization, and cross-architecture distillation transfer.

Publications & Writing

Latest Research & Notes

All Writing