Multiverse Computing Says CompactifAI Nearly Doubles Llama 3.3 Performance on Intel Xeon 6
SAN SEBASTIÁN, Spain, July 23, 2026 — Multiverse Computing, a leader in AI model compression, today announced that its CompactifAI-compressed version of the Llama 3.3 70B model now runs seamlessly on Intel Xeon 6 processors with Performance-cores, utilizing vLLM CPU and Intel Advanced Matrix Extensions (AMX), marking a major advancement in more energy-efficient and scalable AI.
With CompactifAI, Multiverse Computing’s proprietary AI model compression technology, the company compresses and offers optimized versions of leading open-source AI models. This breakthrough delivers improved throughput and latency while maintaining compatibility with mainstream open‑source tooling.
Performance and Efficiency
In controlled benchmarks using an Intel Xeon 6737P processor, the CompactifAI-compressed Llama 3.3 70B model demonstrated significant improvements over the uncompressed baseline. For a request comprising 1,024 input and 1,024 output tokens, the uncompressed baseline achieved an output throughput of 2.00 tokens/second and token throughput of 4.02 tokens/second. The CompactifAI-compressed model delivered an output throughput of 3.86 tokens/second and a total token throughput of 7.81 tokens/second, marking improvements of 93.6% and 94.1%, respectively.
Latency improvements were observed across all concurrency levels. For example, at one concurrent user, the uncompressed baseline required 5,056.34 seconds to process, while the compressed model reduced this to 2,598.22 seconds, resulting in a latency reduction of 48.6%. This trend continued across higher concurrency levels, with the compressed model consistently delivering faster response times.
Additional performance metrics including ITL (Inter-Token Latency), TPOT (Time Per Output Token), and TTFT (Time to First Token) showed substantial reductions:
- ITL (mean): 493.36 ms to 252.08 ms (48.9% reduction).
- TPOT (mean): 494.71 ms to 255.76 ms (48.3% reduction).
- TTFT (mean): 6939.20 ms to 3705.64 ms (46.6% reduction).
Across all concurrency levels tested, the compressed model showcased consistent latency reductions and throughput improvements, with the largest gains at the highest concurrency level (256 concurrent users), where throughput increased by 107.0% and latency decreased by 51.7%.
Accuracy
When measured against standard benchmarks, the CompactifAI-compressed model retained strong accuracy relative to the baseline model, with only minor variations observed:
- BoolQ: Baseline score of 0.904587 vs. compressed score of 0.896024 (0.95% decrease).
- GSM8K: Baseline score of 0.934799 vs. compressed score of 0.924185 (1.14% decrease).
- HellaSwag: Baseline score of 0.590520 vs. compressed score of 0.579068 (1.94% decrease).
- MMLU: Baseline score of 0.776528 vs. compressed model score of 0.757300 (2.48% decrease).
- WinoGrande (Template): Baseline score of 0.678769 vs. compressed score of 0.725335 (6.86% increase).
While it may seem counterintuitive, the accuracy improvement observed in some benchmarks can be attributed to the re-training, or “healing” process that follows compression. During compression, Multiverse Computing reduces the model’s parameter count while simultaneously applying a targeted re-training phase to optimize performance. Although improvement across all benchmarks cannot be guaranteed, the compressed model retains a significant share of the original model’s core information, which can in select cases yield performance gains, as demonstrated by the 6.86% improvement observed in the WinoGrande benchmark.
Overall, the compressed model retained over 97% of the baseline accuracy across these benchmarks, confirming that the compression did not meaningfully affect performance for real-world applications.
Model Size on Disk
The compressed model’s disk size is reduced by approximately 50%, from ~130 GiB to ~65 GiB, significantly reducing storage requirements and provisioning times for large-scale deployments.
According to Enrique Lizaso, Co‑founder & CEO of Multiverse Computing, “This isn’t just a technical milestone — it’s a game‑changer for AI builders.”
Built for Real‑World Deployment
CompactifAI integrates with PyTorch software, Hugging Face platform, and other open‑source frameworks, providing a straightforward path into existing developer pipelines. On the infrastructure side, the stack supports enterprise‑grade AI workloads such as retrieval‑augmented generation (RAG), multimodal reasoning, and domain‑specific applications. Supported models include Llama 4 Scout, Llama 3.3 70B, Llama 3.1 8B, Mistral Small 3.1, and DeepSeek R1, along with ‘Slim’ variants—compressed versions of the original AI models optimized by CompactifAI for enhanced efficiency.
Availability
CompactifAI is available across leading cloud platforms, and now efficiently supports all instances powered by Intel Xeon 6 processors with Performance-cores running vLLM CPU, complementing on-premises deployments and providing a scalable, secure path to production. The solution is optimized for demanding enterprise applications in industries such as finance, healthcare, and manufacturing, offering both high performance and energy efficiency.
About Multiverse Computing
Multiverse Computing is the leader in AI model compression. The company’s deep expertise in quantum software led to the development of CompactifAI, a revolutionary compressor that reduces computing requirements and unleashes new use cases for AI across industries. Headquartered in Donostia, Spain, with offices in the United States, Canada, and across Europe, Multiverse serves more than 100 global customers, including Iberdrola, Bosch, and the Bank of Canada. For more information, visit www.multiversecomputing.com.
Source: Multiverse Computing
Related

