cfaed Publications

A Cross-Stack Approach to Efficient and Scalable Generative AI (Special Session Paper)

Reference

Weitian Wang, Dhingra Pratyush, Asif Ali Khan, Arkapravo Ghosh, Partha Pratim Pande, Jeronimo Castrillon, Priyadarshini Panda, Cecilia De La Parra, Shubham Rai, "A Cross-Stack Approach to Efficient and Scalable Generative AI (Special Session Paper)" (to appear), Proceedings of the 2026 International Conference on Compilers, Architecture, and Synthesis of Embedded Systems (CASES), Oct 2026.

Abstract

This paper presents a holistic, cross-stack vision for efficient and scalable generative AI, motivated by the view that efficiency is most readily improved when hardware, compiler, and algorithmic techniques are pursued together rather than in isolation. We connect four complementary perspectives: (i) foundational hardware innovations, including heterogeneous 2.5D/3D integration and in-memory computing, that reduce the compute, memory, and energy footprint of billion-
parameter models; (ii) extensible compiler infrastructures, built around MLIR, that bridge emerging in-memory hardware to real applications and automate design-space exploration; (iii) cross-layer and algorithmic techniques –quantization, sparsification, and dataflow optimization– that compress and accelerate large language models from the cloud to the edge; and (iv) token reduction methods, such as token merging and clustering, that attack the quadratic cost of the transformer attention mechanism at the model level. By integrating hardware, compiler, and algorithm viewpoints into a single narrative, we outline how co-designed solutions across the stack can compound to enable affordable, scalable, and democratized access to generative intelligence.

Bibtex

@InProceedings{wang_cases26,
author = {Weitian Wang and Dhingra Pratyush and Asif Ali Khan and Arkapravo Ghosh and Partha Pratim Pande and Jeronimo Castrillon and Priyadarshini Panda and Cecilia De La Parra and Shubham Rai},
booktitle = {Proceedings of the 2026 International Conference on Compilers, Architecture, and Synthesis of Embedded Systems (CASES)},
title = {A Cross-Stack Approach to Efficient and Scalable Generative AI (Special Session Paper)},
location = {Barcelona, Spain},
series = {CASES '26 Companion},
abstract = {This paper presents a holistic, cross-stack vision for efficient and scalable generative AI, motivated by the view that efficiency is most readily improved when hardware, compiler, and algorithmic techniques are pursued together rather than in isolation. We connect four complementary perspectives: (i) foundational hardware innovations, including heterogeneous 2.5D/3D integration and in-memory computing, that reduce the compute, memory, and energy footprint of billion-
parameter models; (ii) extensible compiler infrastructures, built around MLIR, that bridge emerging in-memory hardware to real applications and automate design-space exploration; (iii) cross-layer and algorithmic techniques --quantization, sparsification, and dataflow optimization-- that compress and accelerate large language models from the cloud to the edge; and (iv) token reduction methods, such as token merging and clustering, that attack the quadratic cost of the transformer attention mechanism at the model level. By integrating hardware, compiler, and algorithm viewpoints into a single narrative, we outline how co-designed solutions across the stack can compound to enable affordable, scalable, and democratized access to generative intelligence.},
month = oct,
numpages = {10},
year = {2026},
}

Downloads

No Downloads available for this publication

Permalink

https://cfaed.tu-dresden.de/publications?pubId=3913


Go back to publications list