AdaMax#
A version of the Adam optimizer that replaces the RMS property with the infinity norm of the past gradients, using a learning-rate schedule to set the step size each batch. As such, AdaMax is generally more suitable for sparse parameter updates and noisy gradients.
Parameters#
| # | Name | Default | Type | Description |
|---|---|---|---|---|
| 1 | scheduler | Scheduler | The learning-rate schedule that supplies the step size each batch. | |
| 2 | momentumDecay | 0.1 | float | The decay rate of the accumulated velocity. |
| 3 | normDecay | 0.001 | float | The decay rate of the infinity norm. |
Example#
use Rubix\ML\NeuralNet\Optimizers\AdaMax;
use Rubix\ML\NeuralNet\Optimizers\Schedulers\Constant;
$optimizer = new AdaMax(scheduler: new Constant(0.0001), momentumDecay: 0.1, normDecay: 0.001);
References#
-
D. P. Kingma et al. (2014). Adam: A Method for Stochastic Optimization. ↩