02/03/2026

Distributed LLM Inference on Deucalion’s ARM Partition with EPICURE Support

By Alícia Oliveira (INESC TEC / Deucalion)   Most Large Language Model (LLM) inference systems are designed for GPU clusters, especially in multi-node deployments. Still, ARM-based […]