Adam#
Short for Adaptive Moment Estimation, the Adam optimizer pairs a learning-rate schedule with a momentum and rms-normalized step. In addition to storing an exponentially decaying average of past squared gradients like RMSprop, Adam also keeps an exponentially decaying average of past gradients, similar to Momentum. Whereas Momentum can be seen as a ball running down a slope, Adam behaves like a heavy ball with friction.
Parameters#
| # | Name | Default | Type | Description |
|---|---|---|---|---|
| 1 | scheduler | Scheduler | The learning-rate schedule that supplies the step size each batch. | |
| 2 | momentumDecay | 0.1 | float | The decay rate of the accumulated velocity. |
| 3 | normDecay | 0.001 | float | The decay rate of the rms property. |
Example#
use Rubix\ML\NeuralNet\Optimizers\Adam;
use Rubix\ML\NeuralNet\Optimizers\Schedulers\Constant;
$optimizer = new Adam(scheduler: new Constant(0.0001), momentumDecay: 0.1, normDecay: 0.001);
References#
-
D. P. Kingma et al. (2014). Adam: A Method for Stochastic Optimization. ↩