← Proof Notes
Counterexample / Generative models

An omitted skew-part scale in the consensus damping bound

The Numerics of GANs

490 citations ↗Semantic Scholar · 2026-09-09

2017 · arXiv:1705.10461v3, 11 June 2018; NeurIPS 2017 main paper and author-hosted supplement · Reviewed 09 September 2026

An explicit example contradicts the selected statement as written.

Paper context

Overview

The Numerics of GANs studies training as a dynamical system in which two networks continually respond to one another. It connects poor local convergence to properties of the combined gradient's Jacobian and proposes consensus optimization, which adds a penalty based on the size of that gradient. The aim is to reduce unstable or strongly oscillatory behavior near an equilibrium.

Role of the theoretical result

Lemma 9 quantifies how consensus regularization changes the ratio of oscillation to decay in the local linear dynamics. This numerical estimate supports the paper's explanation of improved damping; it is distinct from the preceding qualitative local-convergence argument.

Original paper ↗

01 / Summary

Summary of the result

Lemma 9 bounds the imaginary-to-real eigenvalue ratio after replacing A by A−γAᵀA. The derivation drops the denominator associated with the skew quadratic form when introducing the singular-value lower bound. For A=[(−1,−1),(1,−1)] and γ=1, the transformed eigenvalues are −3±i: the ratio is 1/3, exceeding the printed upper bound 1/5. Retaining ‖A−Aᵀ‖₂ yields a valid sufficient bound and preserves the qualitative damping conclusion.

02 / Statement

Statement under review

Lemma 9, Eqs. (17)–(18), PDF p. 6; proof, Eq. (32), PDF p. 14 of the combined arXiv version. · paraphrased

Lemma 9 bounds the largest imaginary-to-real eigenvalue ratio of A−γAᵀA using a ratio c formed from the symmetric and skew parts of A, the smallest-singular-value lower bound ρ, and γ.

q(γ)=maxλspec(AγAA)ImλReλ  1c+2ρ2γ,c=infv=1v(A+A)vv(AA)v.q(\gamma)=\max_{\lambda\in\operatorname{spec}(A-\gamma A^\top A)}\frac{|\operatorname{Im}\lambda|}{|\operatorname{Re}\lambda|}\ \le\ \frac{1}{c+2\rho^2\gamma},\qquad c=\inf_{\|v\|=1}\frac{|v^*(A+A^\top)v|}{|v^*(A-A^\top)v|}.
Relevant assumptions
  • A is a real, invertible matrix whose symmetric part is negative semidefinite; the paper explicitly permits nonsymmetric A.
  • ||Av||≥ρ||v|| for a positive ρ, and γ>0.
  • The ratio defining c is evaluated on complex unit vectors; zero denominators are interpreted as an infinite ratio when the numerator is positive.
  • The counterexample satisfies the stronger condition A+Aᵀ=−2I.

03 / Derivation

Counterexample and derivation

6 steps · complete derivation
  1. 01

    Choose an admissible matrix

    Use A=−I+J, where J rotates by ninety degrees. Its symmetric part is strictly negative and AᵀA=2I.

    A=(1111),J=(0110),AA=2I,ρ=2.A=\begin{pmatrix}-1&-1\\1&-1\end{pmatrix},\quad J=\begin{pmatrix}0&-1\\1&0\end{pmatrix},\quad A^\top A=2I,\quad\rho=\sqrt2.
  2. 02

    Compute c

    For every complex unit vector the numerator is 2 and the denominator is at most 2. A complex eigenvector of J attains the upper denominator, so c=1.

    v(A+A)v=2,v(AA)v2,v=12(1,i)c=1.|v^*(A+A^\top)v|=2,\quad |v^*(A-A^\top)v|\le2,\quad v=\tfrac1{\sqrt2}(1,-i)^\top\Longrightarrow c=1.
  3. 03

    Compute the damped eigenvalues exactly

    At γ=1 the transformed matrix is −3I+J, with eigenvalues −3±i.

    B=AAA=(3113),spec(B)={3+i,3i}.B=A-A^\top A=\begin{pmatrix}-3&-1\\1&-3\end{pmatrix},\qquad \operatorname{spec}(B)=\{-3+i,-3-i\}.
  4. 04

    Compare the two sides

    The actual ratio is 1/3. The printed right-hand side is 1/(1+4)=1/5.

    q(1)=13>15=1c+2ρ2.q(1)=\tfrac13>\tfrac15=\frac{1}{c+2\rho^2}.
  5. 05

    Locate the missing factor

    For a unit eigenvector, write a=|v*(A+Aᵀ)v| and b=|v*(A−Aᵀ)v|. The exact ratio is b/(a+2γ||Av||²). Dividing by b retains the factor 1/b on the regularization term.

    ImλReλ=ba+2γAv2=1a/b+2γAv2/b.\frac{|\operatorname{Im}\lambda|}{|\operatorname{Re}\lambda|}=\frac{b}{a+2\gamma\|Av\|^2}=\frac{1}{a/b+2\gamma\|Av\|^2/b}.
  6. 06

    Check compatibility with a smooth game

    The same A is the Jacobian of the ascent–descent field for a smooth concave–convex quadratic payoff, so the example is compatible with the paper's game framework.

    f(x,y)=12x2xy+12y2,(xf,yf)=A(x,y).f(x,y)=-\tfrac12x^2-xy+\tfrac12y^2,\qquad(\partial_xf,-\partial_yf)=A(x,y)^\top.

Counterexample

The matrix has AᵀA=2I, c=1 and ρ=√2. With γ=1, its regularized eigenvalues are −3±i, so the observed ratio 1/3 exceeds the claimed bound 1/5.

A=I+J,γ=1,c=1,ρ2=2:q(1)=13>15.A=-I+J,\quad \gamma=1,\quad c=1,\quad\rho^2=2:\qquad q(1)=\tfrac13>\tfrac15.

04 / Implications

Implications and proposed correction

Theoretical implications

Affected result

Lemma 9's quantitative bound is false. The example and the repaired bound both retain the qualitative conclusion that sufficiently strong consensus damping reduces this eigenvalue ratio.

Empirical scope

Relation to reported experiments

This does not disprove the reported experiments or establish that consensus optimization diverges. It is a failure of the stated numerical guarantee.

Proposed correction

Sufficient conditions and revised bound

Write s=‖A−Aᵀ‖₂. For s>0, bounding the skew quadratic form by s gives the following sufficient estimate. If s=0, the regularized matrix is symmetric and the imaginary-to-real eigenvalue ratio is zero.

q(γ)1c+2γρ2/AA2(AA2>0).q(\gamma)\le\frac{1}{c+2\gamma\rho^2/\|A-A^\top\|_2}\qquad(\|A-A^\top\|_2>0).

Implementation implications

Any step-size or damping certificate derived from this particular quantitative bound must retain the skew scale or compute the relevant spectrum directly. The consensus update A−γAᵀA itself does not need to change.

Limits of this review

  • The negative-semidefinite convention is the one explicitly used by the paper for nonsymmetric game Jacobians.
  • The repaired estimate is a sufficient general bound and is exact for the displayed witness.
  • Only Lemma 9's quantitative claim is evaluated here.

05 / References

Sources and correction history

  1. 01
    The Numerics of GANs, exact combined version

    Lemma 9, Eqs. (17)–(18), p. 6; proof Eq. (32), p. 14

  2. 02
  3. 03
  4. 04
    Version history

    Latest listed revision v3, 11 June 2018

  5. 05
    Official reviews

    Three public review reports

  6. 06
    Author code issues

    Open and closed issue records checked

Download the arithmetic witnesses · Python, no dependencies ↓
Correction search · 09 September 2026

2026-09-09 The lemma, matrix assumptions, and proof were reread and the theorem page visually checked. Current arXiv history, the NeurIPS main paper, the author-hosted supplement, public reviews, and available author repository issues were checked. The relevant qualitative discussion in the author's dissertation was also inspected during the audit. "The Numerics of GANs" "Lemma 9" correction "Numerics of GANs" erratum correction No source repairing the missing skew scale in Lemma 9 was found in the bounded search. The inspected later qualitative convergence discussion does not give this quantitative inequality. Previously discussed minibatch-gradient bias concerns a different mechanism. The public issue inventory contained two records; absence of a correction there is not evidence about private correspondence. This search does not establish novelty or exhaust every citing paper.

A bounded search is not evidence of priority or proof that no correction exists.

Entirely AI-generated analysis, including cross-checks by separate AI agents; no independent human verification. Authors have not been contacted. Review standard.

Suggest a correction with a source ↗
Next analysisIncorrect Gaussian noise scaling in the DPS Jensen-gap bound