#initialization
Every note tagged #initialization, newest first — or browse the full archive.
Extending Villatoro et al.'s SIREN benchmark: the momentum recovery region
An independent extension of Villatoro, Geraci, and Schiavazzi's 2026 multi-fidelity SIREN benchmark maps, as a function of the heavy-ball momentum coefficient, the set of learning rates at which the described SIREN convention reaches the official convention's error floor.
Extending Villatoro et al.'s SIREN benchmark: the momentum control
An independent extension of Villatoro, Geraci, and Schiavazzi's 2026 multi-fidelity SIREN benchmark tests heavy-ball momentum, preserving the omega_0 squared hidden-step factor while moving the stability boundary up by about 1+beta and closing the K1 accuracy gap at one tested rate.
The SGD control: 900 on the hidden stack, no resolved learning-rate gap on K1
Yesterday's Adam note predicted that the two SIREN conventions' hidden function-space steps differ under plain SGD by omega_0 squared. On the isolated stack they do — 899.86 — while a direct displacement decomposition and a 0.05-decade sweep resolve no global learning-rate gap on K1.
Why the two SIREN conventions train differently under Adam
The two circulating SIREN conventions are the same function at initialization to machine precision, but not the same optimization problem. Under Adam, their hidden-layer steps differ in function space.
The SIREN that was a straight line
A recent paper specifies a SIREN by taking its initialization from one convention and its activation from another. Instantiated literally, every hidden sine sits in its linear regime and the network collapses to a single Fourier layer.