# Judge-facing evidence scorecard

- Paper ID: `BXE3Z0EHCs`
- Registered claims: 6
- Assessments: 4 verified, 2 falsified as literally registered
- Source matrix: `EVIDENCE_MATRIX.json`
- Prose-local artifact references: validated

## Claim summary

| # | Literal claim | Assessment | Decisive quantitative result |
| ---: | --- | --- | --- |
| 1 | Theorem 3.1 shows that for targets expressible as polynomials in parameters and Kronecker deltas, the coefficients of the formal power series expansion of the loss's time derivatives are themselves polynomials in the width H, the parameterization exponent p, and the initialization variance sigma^2 (Theorem 3.1). | VERIFIED | The pinned Theorem 3.1 states the all-order result; exact Fraction-valued diagram execution yields polynomial Y_s through s=4 and rejects an H^-1 target that violates the premise. |
| 2 | Theorem 4.1 characterizes the leading and Pareto-optimal terms of the loss expansion for identity-tensor targets, organizing gradient-flow learning regimes into a polygon structure defined by scaling conditions on H, p, and sigma^2 (Section 4, Theorem 4.1). | VERIFIED | Exact fronts match every printed Theorem 4.1 triple in 14 SYM/ASYM, nu=2/4, finite-order cells; the source gives the general theorem. |
| 3 | An NTK-like regime, in which features do not evolve during training, appears only at points/edges B-C of the Pareto polygon in the asymmetric parameterization case, while Proposition 8.2 proves the NTK kernel stays static throughout training in the symmetric matrix case (nu=2), explaining the absence of an NTK limit there (Section 4, Proposition 8.2). | FALSIFIED AS LITERALLY REGISTERED | The source says a learning SYM nu=2 model cannot have a stable NTK; a literal GF run has NTK relative drift 0.655313 while its loss decreases. |
| 4 | Mean-field, feature-evolving regimes require the initialization variance to scale as sigma^2 proportional to 1/H (symmetric case) or sigma^2 proportional to 1/H^(2/nu) (asymmetric case), corresponding to edges B-E and C-D of the Pareto polygon respectively (Section 4). | VERIFIED | Pinned source states the B-E and C-D scalings; exact substitution satisfies all printed dominance balances and a wrong ASYM exponent is rejected. |
| 5 | For the canonical polyadic (symmetric, nu=2) identity-tensor target, Section 9 derives a complete closed-form solution to the gradient flow (Equations 19-20), valid across all parameter scalings. | FALSIFIED AS LITERALLY REGISTERED | The source marks Eqs. (19)-(20) asymptotic/formal. At p=H=1, sigma2=1/2, t=0, the exact finite expectation is 3/8 while the registered universal reading gives 1/4. |
| 6 | For the nu=4 symmetric case, Section 10 derives a gradient-ascent solution identifying convergent low-noise and divergent high-noise regimes separated by an explicit threshold given in Equation 23. | VERIFIED | At the paper figure's p=64,H=256, direct Eq. (23) quadrature gives rho*=1.7497342724; 0.95 rho* has no root and 1.05 rho* has a finite critical root. |

## Claim 1 — VERIFIED

> Theorem 3.1 shows that for targets expressible as polynomials in parameters and Kronecker deltas, the coefficients of the formal power series expansion of the loss's time derivatives are themselves polynomials in the width H, the parameterization exponent p, and the initialization variance sigma^2 (Theorem 3.1).

- Decisive quantitative result: The pinned Theorem 3.1 states the all-order result; exact Fraction-valued diagram execution yields polynomial Y_s through s=4 and rejects an H^-1 target that violates the premise.
- Native scale: The theorem quantifies finite H, p, and sigma^2. Two independently seeded direct CP batches are run at H=8,16,32,64, rather than substituting an asserted bound.
- Source locator: camera_ready.tex:240-243; Appendix proof at 1033-1039
- Upstream pin:
  - authored_tex_sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - pdf_sha256: `b7fa6a39cdac27182f67e28ef43b2f70d931d4e686c86d5cabe46526b60e52a7`
  - sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - tar_sha256: `6c470ad469118a8bd3b61f82b3456d95169c0581bce9284a0d68b13b8e37ca9b`
  - version: `2602.04548v2`
- Independent evidence:
  - `outputs/claim1_native_scale.json`
  - `outputs/claim1_polynomial.json`
- Executed outputs:
  - `outputs/claim1_native_scale.json`
- Independent oracle paths:
  - `outputs/claim1_polynomial.json`
- Control paths:
  - `outputs/destructive_controls.json`
- Destructive or boundary control: The executed H^-1 target control has pure-target term p/(2H^2), violates the polynomial-target premise, and is rejected by the same checker.
- Rate relation: claim-consistent; mode `empirical_scaling`; measured slope `3.0466136124045367`.
  - Horizons: 8, 16, 32, 64
  - Repetitions per horizon: 2
  - Measurement: Four increasing-width direct CP gradient measurements give a fitted absolute Y_1 slope 3.0466136124, with two seeded batches per H and independent exact oracle checks.
  - Rate artifact: `outputs/claim1_native_scale.json`
- Limitation: Finite widths test a covered specialization and cannot replace the authored all-order proof.
- Scope boundary: Native execution covers the literal identity-target specialization through s=4; the all-s result is supplied by the pinned authored theorem, not inferred from finite interpolation.

## Claim 2 — VERIFIED

> Theorem 4.1 characterizes the leading and Pareto-optimal terms of the loss expansion for identity-tensor targets, organizing gradient-flow learning regimes into a polygon structure defined by scaling conditions on H, p, and sigma^2 (Section 4, Theorem 4.1).

- Decisive quantitative result: Exact fronts match every printed Theorem 4.1 triple in 14 SYM/ASYM, nu=2/4, finite-order cells; the source gives the general theorem.
- Native scale: Both source parameterizations are executed at H=8,16,32,64 with two independent Gaussian batches per width; the observations are raw finite-system derivatives.
- Source locator: camera_ready.tex:431-443; 479-511
- Upstream pin:
  - authored_tex_sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - pdf_sha256: `b7fa6a39cdac27182f67e28ef43b2f70d931d4e686c86d5cabe46526b60e52a7`
  - sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - tar_sha256: `6c470ad469118a8bd3b61f82b3456d95169c0581bce9284a0d68b13b8e37ca9b`
  - version: `2602.04548v2`
- Independent evidence:
  - `outputs/claim2_native_scale.json`
  - `outputs/claim2_pareto.json`
- Executed outputs:
  - `outputs/claim2_native_scale.json`
- Independent oracle paths:
  - `outputs/claim2_pareto.json`
- Control paths:
  - `outputs/destructive_controls.json`
- Destructive or boundary control: The executed one-power mutation of every predicted q exponent has nonempty symmetric difference in every tested Pareto family.
- Rate relation: claim-consistent; mode `empirical_scaling`; measured slope `3.038234025755577`.
  - Horizons: 8, 16, 32, 64
  - Repetitions per horizon: 2
  - Measurement: Four increasing-width direct SYM CP derivative measurements fit slope 3.03823402576; the same artifact separately records ASYM slope 1.43472199869 and independent diagram-oracle agreement.
  - Rate artifact: `outputs/claim2_native_scale.json`
- Limitation: Finite derivative orders corroborate but do not independently prove the complete all-order polygon.
- Scope boundary: Exact finite-order cells corroborate, rather than replace, the all-order source theorem.

## Claim 3 — FALSIFIED AS LITERALLY REGISTERED

> An NTK-like regime, in which features do not evolve during training, appears only at points/edges B-C of the Pareto polygon in the asymmetric parameterization case, while Proposition 8.2 proves the NTK kernel stays static throughout training in the symmetric matrix case (nu=2), explaining the absence of an NTK limit there (Section 4, Proposition 8.2).

- Decisive quantitative result: The source says a learning SYM nu=2 model cannot have a stable NTK; a literal GF run has NTK relative drift 0.655313 while its loss decreases.
- Native scale: The falsifying p=12,H=24 run is supplemented by two independent runs at each H=16,32,64,128; a single genuine counterexample is sufficient against static-throughout wording.
- Source locator: camera_ready.tex:2422-2452
- Upstream pin:
  - authored_tex_sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - pdf_sha256: `b7fa6a39cdac27182f67e28ef43b2f70d931d4e686c86d5cabe46526b60e52a7`
  - sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - tar_sha256: `6c470ad469118a8bd3b61f82b3456d95169c0581bce9284a0d68b13b8e37ca9b`
  - version: `2602.04548v2`
- Independent evidence:
  - `outputs/claim3_ntk_falsification.json`
  - `outputs/source_audit.json`
- Executed outputs:
  - `outputs/claim3_ntk_falsification.json`
- Independent oracle paths:
  - `outputs/claim3_ntk_falsification.json`
- Control paths:
  - `outputs/destructive_controls.json`
- Destructive or boundary control: The executed unchanged-U control has zero model and NTK drift, but its loss does not decrease, so it is not a learning trajectory.
- Rate relation: no rate-evidence fields are present in the matrix.
- Limitation: The result falsifies the registered static-kernel clause; it does not contest the source proposition that learning makes the kernel non-static.
- Scope boundary: The falsification targets the registered static-kernel wording only; it agrees with the paper's actual non-stability proposition.

## Claim 4 — VERIFIED

> Mean-field, feature-evolving regimes require the initialization variance to scale as sigma^2 proportional to 1/H (symmetric case) or sigma^2 proportional to 1/H^(2/nu) (asymmetric case), corresponding to edges B-E and C-D of the Pareto polygon respectively (Section 4).

- Decisive quantitative result: Pinned source states the B-E and C-D scalings; exact substitution satisfies all printed dominance balances and a wrong ASYM exponent is rejected.
- Native scale: For H=64,128,256,512, two direct trajectories per scenario measure realized variance, loss decrease, and parameter drift under sigma^2=1/H; ASYM nu=2 is the literal 1/H^(2/nu) case.
- Source locator: camera_ready.tex:509-511 and the B-E/C-D rows at 1567-1587
- Upstream pin:
  - authored_tex_sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - pdf_sha256: `b7fa6a39cdac27182f67e28ef43b2f70d931d4e686c86d5cabe46526b60e52a7`
  - sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - tar_sha256: `6c470ad469118a8bd3b61f82b3456d95169c0581bce9284a0d68b13b8e37ca9b`
  - version: `2602.04548v2`
- Independent evidence:
  - `outputs/claim4_native_scale.json`
  - `outputs/claim4_scalings.json`
- Executed outputs:
  - `outputs/claim4_native_scale.json`
- Independent oracle paths:
  - `outputs/claim4_scalings.json`
- Control paths:
  - `outputs/claim4_native_scale.json`
  - `outputs/destructive_controls.json`
- Destructive or boundary control: A wrong sigma^2=H^-1/2 schedule is integrated at every width as a stable short-horizon boundary control, alongside the separately rejected ASYM balance substitution.
- Rate relation: claim-consistent; mode `empirical_scaling`; measured slope `-0.9986695481699445`.
  - Horizons: 64, 128, 256, 512
  - Repetitions per horizon: 2
  - Measurement: Across four increasing widths, direct realized CP initialization variances have fitted slope -0.99866954817 under the claimed schedule, with loss decrease and feature drift recorded for each seeded trajectory.
  - Rate artifact: `outputs/claim4_native_scale.json`
- Limitation: These finite trajectories corroborate the schedules but do not independently prove asymptotic necessity.
- Scope boundary: This verifies the stated scaling relations, not a new independent proof that no alternate feature-learning regime exists.

## Claim 5 — FALSIFIED AS LITERALLY REGISTERED

> For the canonical polyadic (symmetric, nu=2) identity-tensor target, Section 9 derives a complete closed-form solution to the gradient flow (Equations 19-20), valid across all parameter scalings.

- Decisive quantitative result: The source marks Eqs. (19)-(20) asymptotic/formal. At p=H=1, sigma2=1/2, t=0, the exact finite expectation is 3/8 while the registered universal reading gives 1/4.
- Native scale: At increasing p=H=1,2,4,8 with fixed p sigma^2=1/2, two 40,000-sample CP batches per horizon exclude the registered equation value at every horizon.
- Source locator: camera_ready.tex:655-664 and 2473-2603
- Upstream pin:
  - authored_tex_sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - pdf_sha256: `b7fa6a39cdac27182f67e28ef43b2f70d931d4e686c86d5cabe46526b60e52a7`
  - sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - tar_sha256: `6c470ad469118a8bd3b61f82b3456d95169c0581bce9284a0d68b13b8e37ca9b`
  - version: `2602.04548v2`
- Independent evidence:
  - `outputs/claim5_native_scale_falsification.json`
  - `outputs/claim5_nu2_scope_falsification.json`
- Executed outputs:
  - `outputs/claim5_native_scale_falsification.json`
- Independent oracle paths:
  - `outputs/claim5_nu2_scope_falsification.json`
- Control paths:
  - `outputs/destructive_controls.json`
- Destructive or boundary control: The executed Gaussian-fourth-moment control rejects the non-Gaussian replacement that would incorrectly remove the exact finite gap.
- Rate relation: claim-consistent; mode `literal_falsification`; measured slope `5.136910474708504e-16`.
  - Horizons: 1, 2, 4, 8
  - Repetitions per horizon: 2
  - Measurement: At four increasing p=H horizons, two direct finite CP batches per horizon exclude the registered value and the exact gap remains 0.125; fitted gap slope is 5.13691047471e-16.
  - Rate artifact: `outputs/claim5_native_scale_falsification.json`
- Limitation: This disproves only the registered universal finite-scaling extension, not the source's formal or asymptotic use of its formula.
- Scope boundary: The mathematical paper formula may be useful asymptotically; only the registered claim's all-finite-scaling extension is falsified.

## Claim 6 — VERIFIED

> For the nu=4 symmetric case, Section 10 derives a gradient-ascent solution identifying convergent low-noise and divergent high-noise regimes separated by an explicit threshold given in Equation 23.

- Decisive quantitative result: At the paper figure's p=64,H=256, direct Eq. (23) quadrature gives rho*=1.7497342724; 0.95 rho* has no root and 1.05 rho* has a finite critical root.
- Native scale: The evidence uses the p=64,H=256 setting explicitly named by the source figure, with 0.95 rho* and 1.05 rho* as discriminating low- and high-noise runs.
- Source locator: camera_ready.tex:673-708
- Upstream pin:
  - authored_tex_sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - pdf_sha256: `b7fa6a39cdac27182f67e28ef43b2f70d931d4e686c86d5cabe46526b60e52a7`
  - sha256: `359ad3efe4ba8078910aa9841a9a1e6030f1160048ad24e74d4db1ce6cfef505`
  - tar_sha256: `6c470ad469118a8bd3b61f82b3456d95169c0581bce9284a0d68b13b8e37ca9b`
  - version: `2602.04548v2`
- Independent evidence:
  - `outputs/claim6_nu4_threshold.json`
  - `outputs/source_audit.json`
- Executed outputs:
  - `outputs/claim6_nu4_threshold.json`
- Independent oracle paths:
  - `outputs/claim6_nu4_threshold.json`
- Control paths:
  - `outputs/destructive_controls.json`
- Destructive or boundary control: The executed reversed-classification control conflicts with the first-integral signs and the high-noise finite root, so it is rejected.
- Rate relation: no rate-evidence fields are present in the matrix.
- Limitation: This executes the derived reduced solution, not an assertion that a raw finite-p ascent rerun has been performed.
- Scope boundary: The executed object is Eq. (23)'s reduced threshold equation at the paper figure's p=64,H=256, not a falsely labeled raw finite-p ascent simulation.
