Memory Savings at What Cost? A Study of Alternatives to Backpropagation"/> Memory Savings at What Cost? A Study of Alternatives to Backpropagation"/> Memory Savings at What Cost? A Study of Alternatives to Backpropagation (bibtex)
Memory Savings at What Cost? A Study of Alternatives to Backpropagation (bibtex)
by Kunjal Panchal, Sunav Choudhary, Yuriy Brun and Hui Guan
Abstract:

Forward-mode automatic differentiation (FmAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagation-free alternatives for large language model (LLM) fine-tuning, yet their benefits are typically evaluated only against standard backpropagation (BP), omitting memory-efficient variants such as activation checkpointing. We present a unified theoretical and empirical comparison of BP, checkpointed BP, FmAD, and ZO for LLM and vision-language model training, showing that while FmAD and ZO reduce activation memory, they trade memory for higher computational cost and longer wall-clock time to convergence, resulting in lower accuracy and slower training, especially under constrained perturbation budgets. Across models, BP with checkpointing outperforms FmAD and ZO variants, including variance-reduced methods, achieving up to 31.1% higher accuracy, 34.8% faster convergence, and 3.8 fewer computations at comparable memory usage, while also revealing instability-related failure modes in FmAD and ZO. Overall, our results correct a one-sided benchmarking narrative by showing that memory-efficient methods entail fundamentally different trade-offs, and that ignoring these distinctions has led to misleading conclusions about LLM optimization in prior work.

Citation:
Kunjal Panchal, Sunav Choudhary, Yuriy Brun, and Hui Guan, Memory Savings at What Cost? A Study of Alternatives to Backpropagation, in Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026.
Bibtex Entry:
@inproceedings{Panchal26icml,
  author = {Kunjal Panchal and Sunav Choudhary and Yuriy Brun and Hui Guan},
  title =
  {Memory Savings at What Cost? A Study of Alternatives to Backpropagation},
  booktitle = {Proceedings of the 43rd International Conference on Machine Learning (ICML)},
  venue = {ICML},
  address = {Seoul, South Korea},
  month = {July},
  date = {6--11},
  year = {2026},
  accept = {$\frac{6,352}{23,918} \approx 27\%$},

  abstract = {Forward-mode automatic differentiation (FmAD) and zero-order
  (ZO) optimization are increasingly proposed as memory-efficient,
  backpropagation-free alternatives for large language model (LLM)
  fine-tuning, yet their benefits are typically evaluated only against
  standard backpropagation (BP), omitting memory-efficient variants such as
  activation checkpointing. We present a unified theoretical and empirical
  comparison of BP, checkpointed BP, FmAD, and ZO for LLM and vision-language
  model training, showing that while FmAD and ZO reduce activation memory,
  they trade memory for higher computational cost and longer wall-clock time
  to convergence, resulting in lower accuracy and slower training, especially
  under constrained perturbation budgets. Across models, BP with
  checkpointing outperforms FmAD and ZO variants, including variance-reduced
  methods, achieving up to 31.1% higher accuracy, 34.8% faster convergence,
  and 3.8 fewer computations at comparable memory usage, while also revealing
  instability-related failure modes in FmAD and ZO. Overall, our results
  correct a one-sided benchmarking narrative by showing that memory-efficient
  methods entail fundamentally different trade-offs, and that ignoring these
  distinctions has led to misleading conclusions about LLM optimization in
  prior work.},
}
Powered by bibtexbrowser