September 29, 2026 Program SC Awards Test of Time Award Share this page: Twitter Facebook LinkedIn Email By Lori Diachin, SC26 Test of Time Award Chair When Nathan Bell and Michael Garland presented their work on sparse matrix-vector multiplication at SC09, general-purpose GPU computing was still relatively young. CUDA had been introduced only a few years earlier, and researchers were still learning which problems could take advantage of the highly parallel architecture of GPUs, and how. Seventeen years later, their paper, “Implementing Sparse Matrix-Vector Multiplication on Throughput-Oriented Processors,” has been selected for the SC26 Test of Time Award. The award recognizes research published at SC that has continued to influence the field well beyond its original appearance. Bell and Garland’s work addressed a problem that remains familiar today: how do you get highly parallel hardware to perform efficiently when the problem itself is irregular? Nathan Bell Google Michael Garland NVIDIA Finding Structure in Sparsity Sparse matrix-vector multiplication, or SpMV, is an important operation in computational science. It appears repeatedly in methods for solving large-scale linear systems and eigenvalue problems across scientific and engineering applications. However, sparse matrices present a challenge for parallel hardware. Unlike dense matrices, where data and computation tend to follow predictable patterns, sparse matrices can range from highly regular to highly irregular. For GPUs at the time, extracting performance meant exposing enough fine-grained parallelism while keeping execution and memory access sufficiently regular. Bell and Garland explored how different sparse matrix representations could help achieve that balance. Rather than treating all sparse matrices in the same way, their work considered different sparsity patterns and the data formats best suited to them. The paper examined formats including DIA, ELL, CSR and COO, and showed how computation could be organized to reduce the execution and memory divergence caused by irregular matrix structures. The results demonstrated that sparse computation could be mapped successfully onto GPU architectures. On the NVIDIA GeForce GTX 285 used in the study, the proposed techniques achieved high levels of memory bandwidth utilization and substantial performance improvements over the comparison systems available at the time. But the lasting contribution went beyond a set of benchmark results. The work demonstrated something more fundamental: understanding the problem’s structure could be just as important as the hardware used to solve it. Twenty Years of Sparsity on GPUs At SC26, Garland will look back at that work and what has happened since in “Twenty Years of Sparsity on GPUs.” Garland joined NVIDIA in 2006 as one of the founding members of NVIDIA Research. He is now Senior Director of Programming Research, leading a group whose work spans parallel algorithms, programming languages, compilers and runtime systems, and low-level hardware/software interfaces. Sparse matrix-vector multiplication was among the first CUDA kernels he attempted to write. The SIMT execution model made it relatively straightforward to produce a working implementation. Achieving genuinely high performance, however, was a different challenge. A few years later, Garland and Bell set out to develop kernels that delivered high throughput while making effective use of the memory bandwidth GPUs could provide. That work became their SC09 paper. Looking back, Garland highlights two ideas that emerged from the research: selecting the right data structure to take advantage of known matrix structure can make high performance much easier to achieve, and irregularity can be managed effectively when it is relatively rare. Those ideas would continue to matter as GPU computing matured. In the years since SC09, sparse computation and GPU linear algebra have developed considerably, alongside dramatic changes in GPU hardware and the surrounding software ecosystem. Garland’s SC26 talk will revisit lessons from the original work, explore notable developments along the way and consider the challenges that remain. From Early CUDA to Today Both authors went on to contribute to GPU computing well beyond the original paper. Bell is now a Principal Engineer at Google, focused on Search. Previously a Research Scientist at NVIDIA Research, he specialized in sparse linear algebra and parallel programming models and co-developed the Thrust and Cusp libraries. Garland and his team have worked across the GPU software stack, producing high-performance sparse matrix methods; algorithm libraries, including Thrust and CUB; Python frameworks for programming large numbers of GPUs; new approaches to tensor layouts; and methods for scheduling AI kernels. Many of these innovations have become foundational components of the CUDA ecosystem. The hardware Bell and Garland used in 2009 belongs to another generation of computing. The questions raised by their work do not. How should data be represented? Where is the regularity in an irregular problem? How should work be organized to make effective use of increasingly parallel hardware? These questions have followed GPU computing from its early years into today’s HPC and AI systems. And that is what makes research stand the test of time. Congratulations to Nathan Bell and Michael Garland, recipients of the SC26 Test of Time Award. Read the Original Paper N. Bell and M. Garland, “Implementing sparse matrix-vector multiplication on throughput-oriented processors,” Proceedings of the Conference on High Performance Computing Networking, Storage and Analysis, Portland, OR, USA, 2009, pp. 1-11, doi: 10.1145/1654059.1654078