Improved Algorithms for Kernel Matrix-Vector Multiplication Under Sparsity Assumptions

2citations

arXiv:2507.23539

citations

#2790

in ICLR 2025

of 3827 papers

Top Authors

Data Points

Top Authors

Piotr Indyk Michael Kapralov Kshiteej Jitesh Sheth Tal Wagner

Topics

kernel matrix-vector multiplication attention matrices gaussian kernel matrices fast attention computation subquadratic-time algorithms sparsity assumptions

Abstract

Motivated by the problem of fast processing of attention matrices, we study fast algorithms for computing matrix-vector products for asymmetric Gaussian Kernel matrices $K\in \mathbb{R}^{n\times n}$. $K$'s columns are indexed by a set of $n$ keys $k_1,k_2\ldots, k_n\in \mathbb{R}^d$, rows by a set of $n$ queries $q_1,q_2,\ldots,q_n\in \mathbb{R}^d $, and its $i,j$ entry is $K_{ij} = e^{-\|q_i-k_j\|_2^2/2σ^2}$ for some bandwidth parameter $σ>0$. Given a vector $x\in \mathbb{R}^n$ and error parameter $ε>0$, our task is to output a $y\in \mathbb{R}^n$ such that $\|Kx-y\|_2\leq ε\|x\|_2$ in time subquadratic in $n$ and linear in $d$. Our algorithms rely on the following modelling assumption about the matrices $K$: the sum of the entries of $K$ scales linearly in $n$, as opposed to worst case quadratic growth. We validate this assumption experimentally, for Gaussian kernel matrices encountered in various settings such as fast attention computation in LLMs. We obtain the first subquadratic-time algorithm that works under this assumption, for unrestricted vectors.

Citation History

Jan 26, 2026

Jan 27, 2026

1+1

Feb 3, 2026

Feb 13, 2026

2+1

Feb 13, 2026