LANGUAGE AND TRANSFORMER MODELS WITH ESP32-CLASS MICROCONTROLLERS: A SYSTEMATIC REVIEW OF HYBRID DEPLOYMENT AND ON-DEVICE OPTIMIZATION
Abstract
Deploying language and transformer models with ultra-low-power microcontrollers creates a fundamental tension between model capability and the memory, energy, and processing limits of embedded hardware. This systematic review examines how ESP32-class devices participate in language-model systems and when transformer inference can be executed directly on the microcontroller. Following a PRISMA-guided search completed in June 2026, ten studies were included. To avoid conflating an LLM-enabled ESP32 system with LLM inference on the ESP32 itself, the evidence is organized into three deployment classes: cloud-assisted systems, local-host/edge-assisted systems, and fully microcontroller-resident transformer or small-language-model inference. The first two classes preserve access to larger models but introduce host or network dependencies. Fully on-device studies rely on aggressive compression and task specialization: reported language/foundation-model implementations operate within approximately 256–412 KB of RAM, while a fully ternarized Vision Transformer demonstrates a 5.81 MB model on an ESP32-S3 configuration with external PSRAM. Across the evidence base, model-size reductions of up to 63× are reported, but accuracy, latency, and energy values are not directly comparable because tasks and evaluation protocols differ substantially. No included study demonstrates a general-purpose conversational LLM running entirely on an ESP32. The review therefore identifies standardized benchmarking, transparent reporting of inference location and memory hierarchy, scalable compression, and hardware-software co-design as priorities for future research
Full Text:
PDFRefbacks
- There are currently no refbacks.
Publisher:
Department of Electrical Engineering
Universitas Padjadjaran
Jl. Ir. Soekarno km.21, Jatinangor, Sumedang, Jawa Barat 45363