Inherit the spectrum, optimize the frames.
1The University of Texas at Austin 2UIUC 3Emory University 4Together AI 5Recursive Superintelligence Inc 6ELLIS Institute Tübingen
Write each weight matrix as W = U Σ V⊤: the spectrum Σ sets the scale of each singular mode, the frames (U, V) set its output/input directions. Across RLVR runs — including a public reasoning run trained for 3,000+ updates — the learned checkpoint sits strikingly close to the fixed-spectrum family of its base weights: the spectral residual is only ≈3% of the total weight displacement, versus ≈35% for a comparable SFT transition. SFT reshapes the spectrum; RLVR doesn't need to.
Proximity alone could be a high-dimensional accident, so we test it functionally, with interventions in both directions: restore the base spectrum after training, or never let it move at all.
Take an RLVR checkpoint (DeepSeek-R1-Distill-Qwen-1.5B → Nemotron-Research-Reasoning-Qwen-1.5B), keep its learned frames (URL, VRL) fixed, and interpolate only the spectrum between the base model's Σ0 and the RL-trained ΣRL. At α = 0 the RL model's spectrum is fully replaced by the base spectrum. Drag the slider — the accuracy doesn't care:
| Benchmark | Base | α=0.0 | α=0.2 | α=0.4 | α=0.6 | α=0.8 | α=1.0 (RL) |
|---|
Restoring the base spectrum preserves RLVR's gains at every α. A spectrum-only parameterization (frames frozen, singular values trainable) likewise fails to learn comparably — what RLVR needs to change is the frames.
Run the same interpolation in the opposite direction: fix the frames of the SFT model (USFT, VSFT) and blend the spectrum toward the RL-trained ΣRL. If RLVR's gains lived in the spectrum, scores should climb toward the RL checkpoint as α grows. They don't move at all:
| Benchmark | RL ckpt | α=0.0 | α=0.2 | α=0.4 | α=0.6 | α=0.8 | α=1.0 |
|---|
The RL spectrum in SFT frames buys nothing — at any mixing ratio. All 7 metrics stay pinned at the SFT level (α-sweep range ≤ 3.2pp, within Mean@n sampling noise; the α=0 endpoint reproduces the original SFT model to ≤ 0.8pp). The spectra themselves barely differ — relative drift ‖ΔΣ‖/‖Σ‖ ≤ 0.04% across all 198 matrices. Together with the panel above, the evidence runs in both directions: useful RLVR progress is carried by singular-frame motion, not by spectral rescaling.
If the spectrum can be inherited, which variables must remain free? We reconstruct RL endpoints under four structural restrictions and measure how much of the checkpoint update each one fails to explain (two-stage VLA RL pipeline, transition W0 → W2, median across layers):
Inherit the spectrum, optimize the frames: W(U, V) = U Σ0 V⊤, U, V on the Stiefel manifold.
The tangent space TUt rides along the trajectory: at each step the base optimizer takes a tentative update in the tangent plane, and the polar retraction puts the iterate back on the manifold — so every Wt = Ut Σ0 Vt⊤ keeps the base spectrum exactly.
Every represented matrix shares the base spectrum Σ0 exactly; all learning and composition happen in the frame coordinates, with a polar retraction restoring orthonormality. ISO is a stack, not a single optimizer — one principle, two instantiations spanning the RLVR workflow:
Take specialists trained from a shared base, express each expert as a displacement of the singular frames (project → mask trailing modes → solve unit-retention coefficients → retract), and rebuild one fixed-spectrum model.
Apply your existing optimizer — AdamW or Muon, states and all — to (U, V) instead of W, keep Σ0 frozen, and retract after each step — a strictly more constrained model class.
Merging RLVR specialists in fixed-spectrum frame coordinates recovers their complementary capabilities — with zero post-merge compute. Recovery of each specialist's own score by the single merged model:
| Method | LiveBench | LCB v2 | Tool Use (BFCL) | HotpotQA | SQuAD | Avg | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ACC | UT | ACC | UT | Live | Non-Live | 32K | 64K | 32K | 64K | ||
| Base (Qwen2.5-7B-Instruct) | 36.25 | 48.27 | 28.31 | 43.08 | 62.96 | 69.26 | 49.87 | 45.54 | 58.90 | 56.77 | 49.92 |
| RLVR-Coder | 37.42 | 49.42 | 30.17 | 44.24 | 63.37 | 68.07 | 50.21 | 45.61 | 60.51 | 56.84 | 50.59 |
| RLVR-Tool | 36.87 | 48.08 | 28.20 | 42.43 | 71.83 | 81.82 | 50.39 | 45.33 | 59.73 | 57.86 | 52.25 |
| RLVR-Memory | 37.32 | 50.05 | 27.41 | 41.70 | 63.82 | 72.61 | 78.60 | 77.95 | 79.98 | 78.52 | 60.80 |
| Task Arithmetic | 38.85 | 52.18 | 31.55 | 47.06 | 73.73 | 82.86 | 72.22 | 70.49 | 75.03 | 74.95 | 61.89 |
| TIES | 40.30 | 52.04 | 31.46 | 47.33 | 68.24 | 76.54 | 76.42 | 75.24 | 78.27 | 77.86 | 62.37 |
| TSV | 36.96 | 49.23 | 28.92 | 43.37 | 73.64 | 82.69 | 75.37 | 74.10 | 78.30 | 77.26 | 61.98 |
| RAM | 38.59 | 50.38 | 31.31 | 46.33 | 72.56 | 81.89 | 76.17 | 75.29 | 78.42 | 77.83 | 62.88 |
| OrthoMerge-G-TIES | 40.74 | 53.22 | 31.98 | 47.48 | 66.98 | 78.17 | 74.79 | 74.20 | 77.75 | 76.32 | 62.16 |
| ISO-Merger (ours) | 41.81 | 52.73 | 31.06 | 45.69 | 72.32 | 81.07 | 79.46 | 76.77 | 79.10 | 78.01 | 63.80 |
Means over 3 independent stochastic evaluation runs (standard deviations in the paper). Expert rows shaded blue; bold marks the best expert and the best merged result per column.
| Method | LiveBench | LCB v5 | AIME24 | AIME25 | AMC23 | Minerva | Olympiad | Avg | ||
|---|---|---|---|---|---|---|---|---|---|---|
| ACC | UT | ACC | UT | @32 | @32 | @8 | @4 | @4 | ||
| Base (DS-R1-Distill-1.5B) | 17.99 | 26.24 | 16.82 | 20.85 | 30.90 | 23.40 | 63.15 | 27.33 | 43.11 | 29.98 |
| Archer2.0 (coding) | 26.66 | 37.26 | 26.90 | 38.99 | 42.05 | 28.02 | 73.44 | 30.09 | 49.99 | 39.27 |
| JustRL (math) | 22.23 | 31.85 | 23.05 | 30.18 | 53.16 | 36.84 | 82.78 | 34.71 | 55.67 | 41.16 |
| Task Arithmetic | 26.22 | 35.82 | 26.70 | 36.58 | 50.83 | 33.09 | 80.82 | 33.00 | 54.31 | 41.93 |
| TIES | 26.37 | 36.75 | 27.62 | 39.45 | 54.51 | 35.76 | 82.13 | 33.76 | 55.36 | 43.52 |
| TSV | 26.40 | 36.83 | 27.61 | 40.01 | 53.16 | 35.17 | 80.97 | 33.88 | 54.68 | 43.19 |
| RAM | 26.71 | 36.99 | 27.18 | 38.83 | 54.20 | 35.80 | 82.28 | 34.10 | 55.54 | 43.51 |
| OrthoMerge-G-TIES | 26.46 | 36.73 | 27.74 | 39.86 | 54.55 | 35.14 | 81.88 | 33.55 | 55.47 | 43.49 |
| ISO-Merger (ours) | 25.52 | 36.42 | 27.84 | 41.84 | 55.00 | 37.81 | 83.53 | 34.77 | 56.64 | 44.38 |
ISO-Merger achieves the strongest aggregate performance among the compared data-free merging methods on both backbones, and improves worst@4 by 1.62 / 1.36 points — more consistent capability recovery, not just a better mean.
ISO-AdamW uses a single learning rate everywhere (7.5×10−7), while the weight-space baselines get a full learning-rate sweep. On Qwen3-8B-Base, ISO-AdamW reaches AdamW's 270-step accuracy by step 100 — ≈2.7× fewer training steps — and keeps improving to 0.509 vs. 0.495:
Qwen3-8B-Base, DeepMath-103K. Aggregate average@16 over AIME24/25, AMC23, Minerva, OlympiadBench. Extending AdamW 60 more steps yields no further gain — the gap is not under-training.
Qwen3-1.7B-Base — highest final aggregate accuracy vs. the full AdamW sweep.
Qwen3-4B-Base — matches the best AdamW with ≈2.2× fewer steps, then keeps improving.
ISO-Muon — the construction transfers across base optimizers: ≈1.4× fewer steps than tuned Muon.
Coding (DS-1.5B, ArcherCodeR) — leads tuned AdamW on LiveCodeBench; extended baselines never match its peak.
| Model | Method | AIME24 | AIME25 | AMC23 | Minerva | Olympiad | Avg |
|---|---|---|---|---|---|---|---|
| Qwen3-1.7B-Base | Base | 3.75 | 1.25 | 23.04 | 18.49 | 18.43 | 12.99 |
| AdamW | 13.96 | 12.92 | 43.45 | 32.17 | 39.25 | 28.35 | |
| Muon | 16.25 | 8.33 | 43.60 | 32.24 | 38.06 | 27.70 | |
| ISO-AdamW (ours) | 16.04 | 11.04 | 44.50 | 33.11 | 39.01 | 28.74 | |
| Qwen3-4B-Base | Base | 6.88 | 5.00 | 25.45 | 20.97 | 22.43 | 16.15 |
| AdamW | 25.21 | 20.42 | 63.55 | 45.40 | 53.87 | 41.69 | |
| Muon | 25.42 | 19.58 | 63.63 | 46.51 | 56.02 | 42.23 | |
| ISO-AdamW (ours) | 27.29 | 23.13 | 65.74 | 45.15 | 55.99 | 43.46 |
Checkpoint with the highest aggregate score among the final three evaluations, per run. Baselines use the best learning rate from their sweep.
| Method | LCB v5 | LCB v6 | Avg |
|---|---|---|---|
| DS-R1-Distill-1.5B | 17.43 | 18.41 | 17.92 |
| AdamW | 24.01 | 26.24 | 25.13 |
| ISO-AdamW (ours) | 25.49 | 26.91 | 26.20 |
Overhead: the FP64 SVD polar retraction adds ≈86 s per actor update on Qwen3-4B — only ≈7% of end-to-end RL step time, and amenable to overlap with rollout generation in asynchronous RL systems.
@article{zhu2026iso,
title = {ISO: An RLVR-Native Optimization Stack},
author = {Zhu, Hanqing and Cong, Wenyan and Sha, Zhizhou and Mukherjee, Sagnik and
Song, Xinyuan and Gonz{\'a}lez-Mart{\'i}nez, David and Wu, Xiaoxia and
Tian, Yuandong and Liu, Shiwei and Pan, David Z. and Wang, Zhangyang},
journal = {arXiv preprint arXiv:2607.19331},
year = {2026},
eprint = {2607.19331},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2607.19331}
}