### Sidebar

Abstract: Adaptive gradient methods such as AdaGrad and its variants update the stepsize in stochastic gradient descent on the fly according to the gradients received along the way; such methods have gained widespread use in large-scale optimization for their ability to converge robustly, without the need to fine-tune parameters such as the stepsize schedule. Yet, the theoretical guarantees to date for AdaGrad are for online and convex optimization. We bridge this gap by providing strong theoretical guarantees for the convergence of AdaGrad over smooth, nonconvex landscapes. We show that the norm version of AdaGrad (AdaGrad-Norm) converges to a stationary point at the $O(\log(N)/\sqrt{N})$ rate in the stochastic setting, and at the optimal $O(1/N)$ rate in the batch (non-stochastic) setting – in this sense, our convergence guarantees are “sharp”. In particular, both our theoretical results and extensive numerical experiments imply that AdaGrad-Norm is robust to the unknown Lipschitz constant and level of stochastic noise on the gradient.

Rachel Ward, Xiaoxia Wu and Leon Bottou: AdaGrad Stepsizes: Sharp Convergence Over Nonconvex Landscapes, Proceedings of the 36th International Conference on Machine Learning, 97:6677–6686, Edited by Kamalika Chaudhuri and Ruslan Salakhutdinov, Proceedings of Machine Learning Research, PMLR, Long Beach, California, USA, 09–15 Jun 2019.
@inproceedings{ward-2019,
title = {{A}da{G}rad Stepsizes: Sharp Convergence Over Nonconvex Landscapes},
author = {Ward, Rachel and Wu, Xiaoxia and Bottou, Leon},
booktitle = {Proceedings of the 36th International Conference on Machine Learning},
pages = {6677--6686},
year = {2019},
editor = {Chaudhuri, Kamalika and Salakhutdinov, Ruslan},
volume = {97},
series = {Proceedings of Machine Learning Research},
address = {Long Beach, California, USA},
month = {09--15 Jun},
publisher = {PMLR},
url = {http://leon.bottou.org/papers/ward-2019},
}