$52.99
Optimization Algorithms for Distributed Machine Learning
Overview
This book discusses state-of-the-art stochastic optimization algorithms for distributed machine learning and analyzes their convergence speed.
Introduction to Stochastic Gradient Descent
The book first introduces stochastic gradient descent (SGD) and its distributed version, synchronous SGD, where the task of computing gradients is divided across several worker nodes.
Algorithms for Scalability and Efficiency
The author discusses several algorithms that improve the scalability and communication efficiency of synchronous SGD, such as asynchronous SGD, local-update SGD, quantized and sparsified SGD, and decentralized SGD.
Analysis of Algorithms
For each of these algorithms, the book analyzes its error versus iterations convergence, and the runtime spent per iteration.
Trade-offs
The author shows that each of these strategies to reduce communication or synchronization delays encounters a fundamental trade-off between error and runtime.