Last Update

OPML feed of all feeds.

Subscribe to the Atom feed, RSS feed to stay up to date.

Thank you to arXiv for use of its open access interoperability.

Note: the date of arXiv entries announced right after publication holidays might incorrectly show up as the date of the publication holiday itself. This is due to our ad hoc method of inferring announcement dates, which are not returned by the arXiv API.

Powered by Pluto.

Source on GitHub.

Maintained by Nima Anari, Arnab Bhattacharyya, Gautam Kamath.

Theory of Computing Report

Friday, September 11

Jesús A. De Loera, Ethan X. Fang, Shengtao Guo, Junwei Lu, and Hailun Zheng Proved the Simplex–Cube Conjecture for Simple Polytopes.

from Gil Kalai

Shana Tova Shana Tova (happy new Jewish year) to all our readers! We have just returned to Tel Aviv from  the beautiful city of Tiberias, on the Sea of Galilee. The Simplex cube conjecture The simplex–cube conjecture was posed in … Continue reading →
Shana Tova

Shana Tova (happy new Jewish year) to all our readers! We have just returned to Tel Aviv from  the beautiful city of Tiberias, on the Sea of Galilee.

The Simplex cube conjecture

The simplex–cube conjecture was posed in my 1990 paper and was among the five problems on convex polytopes discussed in this 2008 post.

Conjecture A. For every k there exists an integer d(k)  such that if P is a d-polytope with d\ge d(k), then P has a k-face which is either a simplex or (combinatorially) a cube.

We denote by d(k) the smallest such integer, if it exists, and otherwise set d(k) = \infty.

A weaker conjecture, which remains open in general, is the following.

Conjecture B. For every positive integer k, there exist an integer d'(k) and a finite collection \mathcal F_k of k-dimensional polytopes such that every d-polytope with d\ge d'(k) has a k-face combinatorially equivalent to a member of \mathcal F_k.

As with d(k), we let d'(k) denote the smallest possible threshold, and set d'(k)=\infty if no such threshold exists.

Euler’s theorem implies that $latex d′(2)=3$: every 3-polytope has a 2-face that is a triangle, quadrilateral, or pentagon. I proved that $latex d(2)=5$, namely, every 5-polytope has a 2-face that is either a triangle or a quadrilateral. This answered a question of Perles and Shephard from 1967. The bound is sharp: the regular 120-cell is a 4-polytope all of whose 2-faces are pentagons.

On Unavoidable Faces of High-Dimensional Polytopes

I was very happy to learn that Jesús A. De Loera, Ethan X. Fang, Shengtao Guo, Junwei Lu, and Hailun Zheng  in their paper On Unavoidable Faces of High Dimensional Polytopes  proved Conjecture A for simple polytopes, along with other remarkable results. (A d-polytope is simple if exactly d edges meet at each vertex.)

In my 1990 paper I considered the asymmetric version of the conjecture. Let \ell,k be positive integers, and let f(\ell,k) be the smallest integer d such that every polytope of dimension at least d contains either an \ell-dimensional simplex face or a k-dimensional face combinatorially equivalent to a cube.  If no such integer exists, set f(\ell,k)=\infty. The simplex–cube conjecture asserts that

f(\ell,k)<\infty

for all \ell,k. The diagonal case is d(k)=f(k,k).

If f_s(\ell,k) denotes the corresponding threshold restricted to simple polytopes, De Loera, Fang, Guo, Lu, and Zheng proved that f_s(\ell,k)<\infty for every \ell\ge 2 and k\ge 3. Moreover, they obtained the explicit bounds:

Theorem (De Loera, Fang, Guo, Lu, and Zheng). For every integer k\ge3:

(i) f_s(2,k)\le 2k^2-1.

(ii) For every integer \ell\ge3,

\displaystyle f_s(\ell,k)\le \frac{1}{2}k^2\ell,2^k.

The first bound slightly improves my old bound f_s(2,k)\le 2k^2.

The paper contains several other developments. First, the authors obtain substantial new lower bounds for both the general and simple versions of the problem. Second, they make remarkable progress on a related question concerning unavoidable small 3-dimensional faces. Earlier work of Meisinger, Kleinschmidt, and me showed that every rational d-polytope with d\ge9 has a 3-face with fewer than 78 vertices or fewer than 78 facets. For dimensions at least 15, the new paper substantially improves the size bound and removes the rationality assumption: every convex polytope in these dimensions has a 3-face with at most 13 facets. The proof uses the nonnegativity of toric g-numbers, inequalities due to Billera and Ehrenborg for the cd-index, and convolution operations. An exact rational certificate involving flag numbers is obtained using linear programming.

By Gil Kalai

The Computational Complexity of Holant Problems on 4-regular Graphs from the Stable Subgroup Sequence of $SL(2,\mathbb{C})$

from arXiv: Computational Complexity

Authors: Yuan Huang, Zhiguo Fu

The Holant framework provides a general setting for studying counting problems and includes graph homomorphisms (\#GH) and counting constraint satisfaction problems (\#CSP) as special cases. Over the past twenty years, a series of computational complexity dichotomies have been established for Holant problems, but the classification for complex-valued signatures is still open. The main obstacle is the case in which all signatures have even arity. In this paper, we establish a dichotomy for Holant problems with a complex-valued 4-ary signature, which is a key base case for the full classification of Holant problems. We present a new strategy by introducing Schur's theorem, the classification of finite subgroups of $\mathrm{SL}(2,\mathbb{C})$ and stable subgroup sequences into the proof. These new techniques are of independent interest.

Authors: Yuan Huang, Zhiguo Fu

The Holant framework provides a general setting for studying counting problems and includes graph homomorphisms (\#GH) and counting constraint satisfaction problems (\#CSP) as special cases. Over the past twenty years, a series of computational complexity dichotomies have been established for Holant problems, but the classification for complex-valued signatures is still open. The main obstacle is the case in which all signatures have even arity. In this paper, we establish a dichotomy for Holant problems with a complex-valued 4-ary signature, which is a key base case for the full classification of Holant problems. We present a new strategy by introducing Schur's theorem, the classification of finite subgroups of $\mathrm{SL}(2,\mathbb{C})$ and stable subgroup sequences into the proof. These new techniques are of independent interest.

Online Treasure Hunt in Vertex-Permuted Dynamic Rings

from arXiv: Computational Complexity

Authors: Kamran Ayoubi, Bernard Mans, Lata Narayanan

We study the problem of treasure hunt by a group of $k \geq 1$ agents in vertex-permuted dynamic rings (VP). In this model, the $n$ vertices remain on a ring but are permuted at each time step. We first show that treasure hunt is impossible for any $k \leq n-3$ agents, if there are no restrictions on the sequence of permutations used in the dynamic ring. We then study the $VP(δ)$ setting, in which for every pair $i, j$ of vertices, the edge $(i, j)$ is guaranteed to appear within $δ$ steps. We show that the class $VP(δ)$ is feasible only for $δ\geq \left\lceil \frac{n-1}{2}\right\rceil$. For the one-agent case, we show a tight bound of $Θ(δn)$ on the worst-case search time as well as competitive ratio of any online algorithm for treasure hunt, provided $δ\geq 2n$. We then give an optimal algorithm for $k$ agents, thereby showing that $k$ agents can obtain a speedup of $k$ on the worst-case search time. Finally, in the R-VP setting, in which in every step, the vertices are arranged as a ring according to a random permutation, we show that treasure hunt takes expected $Θ(n)$ steps against an oblivious adversary and $Θ(n \log n)$ steps against an adaptive adversary.

Authors: Kamran Ayoubi, Bernard Mans, Lata Narayanan

We study the problem of treasure hunt by a group of $k \geq 1$ agents in vertex-permuted dynamic rings (VP). In this model, the $n$ vertices remain on a ring but are permuted at each time step. We first show that treasure hunt is impossible for any $k \leq n-3$ agents, if there are no restrictions on the sequence of permutations used in the dynamic ring. We then study the $VP(δ)$ setting, in which for every pair $i, j$ of vertices, the edge $(i, j)$ is guaranteed to appear within $δ$ steps. We show that the class $VP(δ)$ is feasible only for $δ\geq \left\lceil \frac{n-1}{2}\right\rceil$. For the one-agent case, we show a tight bound of $Θ(δn)$ on the worst-case search time as well as competitive ratio of any online algorithm for treasure hunt, provided $δ\geq 2n$. We then give an optimal algorithm for $k$ agents, thereby showing that $k$ agents can obtain a speedup of $k$ on the worst-case search time. Finally, in the R-VP setting, in which in every step, the vertices are arranged as a ring according to a random permutation, we show that treasure hunt takes expected $Θ(n)$ steps against an oblivious adversary and $Θ(n \log n)$ steps against an adaptive adversary.

The Quantum Overlap Gap Property and Algorithmic Hardness for the Quantum Hypergraph Max-Cut Problem

from arXiv: Computational Complexity

Authors: Mikhail Mints, Eric R. Anschuetz

In this work, we analyze the average-case hardness of approximation for the Quantum Hypergraph Max-Cut problem using the theoretical framework of the Quantum Overlap Gap Property (QOGP). We establish two main results. Our first result applies to a wide class of stable quantum algorithms, satisfying a Lipschitz property with respect to the quantum Wasserstein distance of order $2$. We show a weak hardness result, demonstrating that for any Lipschitz constant $L$, there is some $k$ such that $L$-stable algorithms cannot approximate the optimal solution to Quantum Hypergraph Max-Cut on $k$-uniform hypergraphs in the average case. Additionally, we establish a strong hardness result where $k$ is independent of $L$, but only for a more restricted class of local quantum algorithms defined using the quantum Wasserstein distance of order $\infty$. We apply these results to establish concrete depth lower bounds for popular quantum algorithms for preparing near-optimal states for this problem.

Authors: Mikhail Mints, Eric R. Anschuetz

In this work, we analyze the average-case hardness of approximation for the Quantum Hypergraph Max-Cut problem using the theoretical framework of the Quantum Overlap Gap Property (QOGP). We establish two main results. Our first result applies to a wide class of stable quantum algorithms, satisfying a Lipschitz property with respect to the quantum Wasserstein distance of order $2$. We show a weak hardness result, demonstrating that for any Lipschitz constant $L$, there is some $k$ such that $L$-stable algorithms cannot approximate the optimal solution to Quantum Hypergraph Max-Cut on $k$-uniform hypergraphs in the average case. Additionally, we establish a strong hardness result where $k$ is independent of $L$, but only for a more restricted class of local quantum algorithms defined using the quantum Wasserstein distance of order $\infty$. We apply these results to establish concrete depth lower bounds for popular quantum algorithms for preparing near-optimal states for this problem.

PureSuperQMA(exp) = BellPureSymQMA(poly) = QMA via Dimension-Free Bosonic Argmax

from arXiv: Computational Complexity

Authors: William Gay, Fernando Granha Jeronimo, Lenny Liu, Itai Leigh, Pei Wu, Haochen Xu

Pure-state consistency problems naturally lead to quantum proof systems in which a single pure witness must satisfy many acceptance constraints. The corresponding class $\mathsf{PureSuperQMA}$ was previously known to lie between $\mathsf{QMA}$ and $\mathsf{QMA}(2)$, and Kamminga and Rudolph (ITCS'26) conjectured that both containments are strict. In this paper, we prove the following surprising complexity collapses $$ \mathsf{QMA} = \mathsf{PureSuperQMA} = \mathsf{PureSuperQMA}(\text{exp}) = \mathsf{BellPureSymQMA}(\text{poly}) $$ Here $\mathsf{PureSuperQMA}(\text{exp})$ allows exponentially many checks which are uniformly indexed and efficiently generated, while requiring an inverse-polynomial violation margin and an inverse-polynomial fraction of violated checks for the NO cases. $\mathsf{BellPureSymQMA}(\text{poly})$ is a related model that requires the prover to give the verifier polynomially many copies of a pure state, which the verifier measures separately with logarithmic output length for each local measurement, before processing the outcomes jointly. The main technical ingredient is a dimension-free stability bound for symmetric tensor states. Our simulations use polynomially many witness registers and combine a random-pair SWAP test with a permutation-invariant lift of the original verification procedure. The key step is to show that, on the symmetric subspace, the extremal verification value is close to that of some tensor-power witness with dimension-independent error. Applying this argument to the two verification models yields both simulations. As a consequence, exact $k$-local pure-state consistency is $\mathsf{QMA}$-complete for every fixed $k\ge2$, and so are the corresponding exact bosonic and fermionic pure $N$-representability problems.

Authors: William Gay, Fernando Granha Jeronimo, Lenny Liu, Itai Leigh, Pei Wu, Haochen Xu

Pure-state consistency problems naturally lead to quantum proof systems in which a single pure witness must satisfy many acceptance constraints. The corresponding class $\mathsf{PureSuperQMA}$ was previously known to lie between $\mathsf{QMA}$ and $\mathsf{QMA}(2)$, and Kamminga and Rudolph (ITCS'26) conjectured that both containments are strict. In this paper, we prove the following surprising complexity collapses $$ \mathsf{QMA} = \mathsf{PureSuperQMA} = \mathsf{PureSuperQMA}(\text{exp}) = \mathsf{BellPureSymQMA}(\text{poly}) $$ Here $\mathsf{PureSuperQMA}(\text{exp})$ allows exponentially many checks which are uniformly indexed and efficiently generated, while requiring an inverse-polynomial violation margin and an inverse-polynomial fraction of violated checks for the NO cases. $\mathsf{BellPureSymQMA}(\text{poly})$ is a related model that requires the prover to give the verifier polynomially many copies of a pure state, which the verifier measures separately with logarithmic output length for each local measurement, before processing the outcomes jointly. The main technical ingredient is a dimension-free stability bound for symmetric tensor states. Our simulations use polynomially many witness registers and combine a random-pair SWAP test with a permutation-invariant lift of the original verification procedure. The key step is to show that, on the symmetric subspace, the extremal verification value is close to that of some tensor-power witness with dimension-independent error. Applying this argument to the two verification models yields both simulations. As a consequence, exact $k$-local pure-state consistency is $\mathsf{QMA}$-complete for every fixed $k\ge2$, and so are the corresponding exact bosonic and fermionic pure $N$-representability problems.

Oracle Separations in the Fourier Hierarchy

from arXiv: Computational Complexity

Authors: Atul Mantri

The Fourier hierarchy $\mathrm{FH}_0\subseteq\mathrm{FH}_1\subseteq\mathrm{FH}_2\subseteq\cdots$, introduced by Shi (TCS 2005), measures a quantum computation by the number of Hadamard layers it uses. Between two layers the circuit may permute basis states and attach phases, but it may not create superposition; the layers are its only source of interference. The first level is exactly $\mathrm{BPP}$, while the second already solves Simon's problem and, through phase estimation, factors integers. Shi conjectured that every additional layer strictly increases computational power, and asked, as a first step, for oracle separations between consecutive levels. To our knowledge, the question was open at every level $k\ge2$. We prove that for every constant $k\ge2$ there is an oracle relative to which $\mathrm{FH}_k\subsetneq\mathrm{FH}_{k+1}$. The separating problem is built from Forrelation (Aaronson and Ambainis, STOC 2015): the level above solves it with a constant number of queries, whereas at level $k$ it stays hard even for circuits making exponentially many queries. This holds for both of the usual ways of giving a circuit access to an oracle, the phase oracle and the standard oracle, which writes its answer into a register. The two are not interchangeable: relative to an oracle, the standard oracle is strictly more powerful at the same number of layers. We also separate the union of all the levels from $\mathrm{BQP}$ relative to an oracle. The lower bounds rest on a structural property of the hierarchy: the number of Hadamard layers limits how adaptively a circuit can query its oracle. With a phase oracle, a circuit with $k$ layers is reproduced exactly by an algorithm making only $k-1$ rounds of parallel queries, which brings known lower bounds for such algorithms to bear. The standard oracle lets a circuit branch on earlier answers, and that case needs a separate argument.

Authors: Atul Mantri

The Fourier hierarchy $\mathrm{FH}_0\subseteq\mathrm{FH}_1\subseteq\mathrm{FH}_2\subseteq\cdots$, introduced by Shi (TCS 2005), measures a quantum computation by the number of Hadamard layers it uses. Between two layers the circuit may permute basis states and attach phases, but it may not create superposition; the layers are its only source of interference. The first level is exactly $\mathrm{BPP}$, while the second already solves Simon's problem and, through phase estimation, factors integers. Shi conjectured that every additional layer strictly increases computational power, and asked, as a first step, for oracle separations between consecutive levels. To our knowledge, the question was open at every level $k\ge2$. We prove that for every constant $k\ge2$ there is an oracle relative to which $\mathrm{FH}_k\subsetneq\mathrm{FH}_{k+1}$. The separating problem is built from Forrelation (Aaronson and Ambainis, STOC 2015): the level above solves it with a constant number of queries, whereas at level $k$ it stays hard even for circuits making exponentially many queries. This holds for both of the usual ways of giving a circuit access to an oracle, the phase oracle and the standard oracle, which writes its answer into a register. The two are not interchangeable: relative to an oracle, the standard oracle is strictly more powerful at the same number of layers. We also separate the union of all the levels from $\mathrm{BQP}$ relative to an oracle. The lower bounds rest on a structural property of the hierarchy: the number of Hadamard layers limits how adaptively a circuit can query its oracle. With a phase oracle, a circuit with $k$ layers is reproduced exactly by an algorithm making only $k-1$ rounds of parallel queries, which brings known lower bounds for such algorithms to bear. The standard oracle lets a circuit branch on earlier answers, and that case needs a separate argument.

Topology inside NC$^1$

from arXiv: Computational Complexity

Authors: Eric Allender, Samir Datta, Arsenii Karnaukhov, Sambuddha Roy, Alexander Shekhovstov

We show that ACC$^0$ is precisely what can be computed with constant-width circuits of polynomial size and polylogarithmic genus. This extends a characterization given by Hansen, showing that planar constant-width circuits also characterize ACC$^0$. Thus polylogarithmic genus provides no additional computational power in this model. We consider other generalizations of planarity, including crossing number and thickness. We show that constant-width circuits of polynomial size and thickness two already suffice to capture all of NC$^1$.

Authors: Eric Allender, Samir Datta, Arsenii Karnaukhov, Sambuddha Roy, Alexander Shekhovstov

We show that ACC$^0$ is precisely what can be computed with constant-width circuits of polynomial size and polylogarithmic genus. This extends a characterization given by Hansen, showing that planar constant-width circuits also characterize ACC$^0$. Thus polylogarithmic genus provides no additional computational power in this model. We consider other generalizations of planarity, including crossing number and thickness. We show that constant-width circuits of polynomial size and thickness two already suffice to capture all of NC$^1$.

Some results on Archdeacon's conjecture for rotation systems

from arXiv: Computational Geometry

Authors: Arahat Chikkatur, Ji Zeng

A rotation system on $n$ elements assigns to each element a cyclic order of the other $n-1$ elements. A four-element subset is non-planar if its induced rotation system cannot be realized by a crossing-free drawing of $K_4$. As a combinatorial strengthening of Hill's conjecture on the crossing number of the complete graph, Archdeacon conjectured that every rotation system on $n$ elements has at least $H(n)=\frac{1}{4} \lfloor\frac {n}{2}\rfloor \lfloor\frac{n-1}{2}\rfloor \lfloor\frac{n-2}{2}\rfloor \lfloor\frac{n-3}{2}\rfloor$ non-planar four-element subsets. We computationally verify Archdeacon's conjecture for $n\leq 10$ and show that every extremal rotation system in these orders is realizable by a simple drawing. With computer assistance, we prove that every rotation system on $n$ elements has at least $(8/9 - o(1)) H(n)$ non-planar four-element subsets. We also present a proof by hand for a weaker lower bound of $(2/3-o(1)) H(n)$. Finally, extending recent work of Felsner on antipodal pairs in drawings, we show that Archdeacon's conjecture holds for antipodally shellable rotation systems.

Authors: Arahat Chikkatur, Ji Zeng

A rotation system on $n$ elements assigns to each element a cyclic order of the other $n-1$ elements. A four-element subset is non-planar if its induced rotation system cannot be realized by a crossing-free drawing of $K_4$. As a combinatorial strengthening of Hill's conjecture on the crossing number of the complete graph, Archdeacon conjectured that every rotation system on $n$ elements has at least $H(n)=\frac{1}{4} \lfloor\frac {n}{2}\rfloor \lfloor\frac{n-1}{2}\rfloor \lfloor\frac{n-2}{2}\rfloor \lfloor\frac{n-3}{2}\rfloor$ non-planar four-element subsets. We computationally verify Archdeacon's conjecture for $n\leq 10$ and show that every extremal rotation system in these orders is realizable by a simple drawing. With computer assistance, we prove that every rotation system on $n$ elements has at least $(8/9 - o(1)) H(n)$ non-planar four-element subsets. We also present a proof by hand for a weaker lower bound of $(2/3-o(1)) H(n)$. Finally, extending recent work of Felsner on antipodal pairs in drawings, we show that Archdeacon's conjecture holds for antipodally shellable rotation systems.

Learn the Solid, Not the File: Canonical Inputs for Neural Networks on CAD Boundary Representations

from arXiv: Computational Geometry

Authors: Heinrich Jiang, Hager Yasser Mohamed, Alexander Hitt, Valeriia Lomakina, Henning Jiang, Jennifer Jang

Boundary representation (B-rep) is the standard format used by modern CAD systems for parametric 3D models. It turns out, the exact same solid can be represented by different B-reps: for example, two engineers using different operations, a geometry kernel rebuilding the file, and an export setting repartitioning faces will lead to different B-reps even though the underlying solid remains the same. We show that existing B-rep encoders are not robust to variation in the B-rep with the same solid on perturbations applied to standard benchmarks, naturally occurring variations inherent to CAD software, and differences in how designers model the same part via a human dataset we created in FreeCAD. The performance of popular B-rep encoders often collapses catastrophically. We propose the canonical region graph, an input representation whose nodes, features and coordinate frame are derived from the solid itself and show theoretical invariance guarantees on repartitioning and rigid motions. It matches the strongest baseline on standard benchmarks, and is stable under every perturbation we test.

Authors: Heinrich Jiang, Hager Yasser Mohamed, Alexander Hitt, Valeriia Lomakina, Henning Jiang, Jennifer Jang

Boundary representation (B-rep) is the standard format used by modern CAD systems for parametric 3D models. It turns out, the exact same solid can be represented by different B-reps: for example, two engineers using different operations, a geometry kernel rebuilding the file, and an export setting repartitioning faces will lead to different B-reps even though the underlying solid remains the same. We show that existing B-rep encoders are not robust to variation in the B-rep with the same solid on perturbations applied to standard benchmarks, naturally occurring variations inherent to CAD software, and differences in how designers model the same part via a human dataset we created in FreeCAD. The performance of popular B-rep encoders often collapses catastrophically. We propose the canonical region graph, an input representation whose nodes, features and coordinate frame are derived from the solid itself and show theoretical invariance guarantees on repartitioning and rigid motions. It matches the strongest baseline on standard benchmarks, and is stable under every perturbation we test.

Almost Linear Universal Point Sets for Planar Graphs

from arXiv: Computational Geometry

Authors: Taylor Gordon

A point set is universal for planar graphs on $n$ vertices if every such graph has a straight-line drawing without crossings whose vertices belong to the set. We construct universal point sets of size $n^{1+o(1)}$, improving the previous quadratic upper bound. Our construction uses the reduction of Bannister, Cheng, Devanny, and Eppstein from universal point sets to superpatterns for $213$-avoiding permutations. We represent these permutations by ordered rooted forests and construct a small family of intervals containing every such forest. The result follows from a straightforward bound on the size of the family of intervals. GPT-6 Astra assisted in developing the construction and proof.

Authors: Taylor Gordon

A point set is universal for planar graphs on $n$ vertices if every such graph has a straight-line drawing without crossings whose vertices belong to the set. We construct universal point sets of size $n^{1+o(1)}$, improving the previous quadratic upper bound. Our construction uses the reduction of Bannister, Cheng, Devanny, and Eppstein from universal point sets to superpatterns for $213$-avoiding permutations. We represent these permutations by ordered rooted forests and construct a small family of intervals containing every such forest. The result follows from a straightforward bound on the size of the family of intervals. GPT-6 Astra assisted in developing the construction and proof.

Max Independent Set Remains NP-hard when Excluding a Planar Induced Minor

from arXiv: Data Structures and Algorithms

Authors: Édouard Bonnet, Yeonsu Chang

We show that there is a fixed planar graph $H$, namely the $5 \times 5$ grid, such that Max Independent Set remains NP-hard in $H$-induced-minor-free graphs. This refutes the Dallard--Milanič--Štorgel conjecture and a weakening of it by Gartland and Lokshtanov, and by Korhonen.

Authors: Édouard Bonnet, Yeonsu Chang

We show that there is a fixed planar graph $H$, namely the $5 \times 5$ grid, such that Max Independent Set remains NP-hard in $H$-induced-minor-free graphs. This refutes the Dallard--Milanič--Štorgel conjecture and a weakening of it by Gartland and Lokshtanov, and by Korhonen.

Tight Time-Space Lower Bounds for Collision Finding and Element Distinctness under Label Symmetry

from arXiv: Data Structures and Algorithms

Authors: Frédéric Magniez, Sebastian Zur

How much memory is needed to retain the quantum speedup for collision finding? For a uniformly random function $f:[N]\to [N]$, the BHT algorithm finds a collision using $O(N^{1/3})$ queries and a quantumly accessible classical table containing $O(N^{1/3})$ input-output pairs, whereas a logarithmic-space Grover search uses $O(\sqrt N)$ queries. Determining the optimal query-space tradeoff between these extremes remains a major open problem. We resolve this equation within the class of label-symmetric algorithms, which treat the function $f$'s output labels as interchangeable. We prove that such algorithm that makes $T$ queries, uses $S$ qubits, and finds a collision in a uniformly random function $f:[M]\to [N]$ with constant probability satisfies $$T=Ω(N^{1/3}) \qquad\text{and}\qquad T^2S=Ω(N\log N).$$ For the setting where $M=N$, these bounds are matched by a space-efficient implementation of the BHT algorithm. As a consequence of our tradeoff, any label-symmetric algorithm for the search version of Element Distinctness on $f: [n] \to [n^2]$ must satisfy $$T=Ω(n^{2/3}) \qquad\text{and}\qquad T^2S=Ω(n^2\log n),$$ matching Ambainis's quantum walk. Thus, both tradeoffs are optimal within the class of label-symmetric algorithms. To prove these results, we develop a space-sensitive version of the compressed oracle technique. The compressed oracle records the information learned by the algorithm in an evolving superposition of databases. Using label symmetry and representation theory, we show that an algorithm using $S$ qubits can effectively retain information about only $O(S/\log N)$ collision-free database entries. Substituting this estimate into the compressed oracle technique yields the stated tradeoffs.

Authors: Frédéric Magniez, Sebastian Zur

How much memory is needed to retain the quantum speedup for collision finding? For a uniformly random function $f:[N]\to [N]$, the BHT algorithm finds a collision using $O(N^{1/3})$ queries and a quantumly accessible classical table containing $O(N^{1/3})$ input-output pairs, whereas a logarithmic-space Grover search uses $O(\sqrt N)$ queries. Determining the optimal query-space tradeoff between these extremes remains a major open problem. We resolve this equation within the class of label-symmetric algorithms, which treat the function $f$'s output labels as interchangeable. We prove that such algorithm that makes $T$ queries, uses $S$ qubits, and finds a collision in a uniformly random function $f:[M]\to [N]$ with constant probability satisfies $$T=Ω(N^{1/3}) \qquad\text{and}\qquad T^2S=Ω(N\log N).$$ For the setting where $M=N$, these bounds are matched by a space-efficient implementation of the BHT algorithm. As a consequence of our tradeoff, any label-symmetric algorithm for the search version of Element Distinctness on $f: [n] \to [n^2]$ must satisfy $$T=Ω(n^{2/3}) \qquad\text{and}\qquad T^2S=Ω(n^2\log n),$$ matching Ambainis's quantum walk. Thus, both tradeoffs are optimal within the class of label-symmetric algorithms. To prove these results, we develop a space-sensitive version of the compressed oracle technique. The compressed oracle records the information learned by the algorithm in an evolving superposition of databases. Using label symmetry and representation theory, we show that an algorithm using $S$ qubits can effectively retain information about only $O(S/\log N)$ collision-free database entries. Substituting this estimate into the compressed oracle technique yields the stated tradeoffs.

A Reusable Framework for Robust Approximation Algorithms in the Interval Uncertainty Model

from arXiv: Data Structures and Algorithms

Authors: Klasing Ralf, Mömke Tobias, Naquin Émile

Robust optimization under interval uncertainty aims to compute solutions that perform well on a range of scenarios that are described by interval-constrained costs. In this paper, we revisit a framework introduced by Ganesh, Maggs and Panigrahi in 2020 to study the robust optimization of NP-hard problems under interval uncertainty. We start by generalizing a result in the $\ell=0$ case, which transforms a category of approximation algorithms into a robust approximation algorithm. Furthermore, in the general case, we provide a theorem that turns any local search-based approximation algorithm into a robust approximation algorithm under three newly formalized conditions over the moves of the local search algorithm. We then use this result to present the first robust approximation algorithm for Weighted $k$-Set Cover, the third NP-hard problem known to admit a robust approximation, and the first since the publication of Ganesh, Maggs and Panigrahi's paper.

Authors: Klasing Ralf, Mömke Tobias, Naquin Émile

Robust optimization under interval uncertainty aims to compute solutions that perform well on a range of scenarios that are described by interval-constrained costs. In this paper, we revisit a framework introduced by Ganesh, Maggs and Panigrahi in 2020 to study the robust optimization of NP-hard problems under interval uncertainty. We start by generalizing a result in the $\ell=0$ case, which transforms a category of approximation algorithms into a robust approximation algorithm. Furthermore, in the general case, we provide a theorem that turns any local search-based approximation algorithm into a robust approximation algorithm under three newly formalized conditions over the moves of the local search algorithm. We then use this result to present the first robust approximation algorithm for Weighted $k$-Set Cover, the third NP-hard problem known to admit a robust approximation, and the first since the publication of Ganesh, Maggs and Panigrahi's paper.

Convex Optimization with Nested Evolving Feasible Sets (CONES) under Time-Varying Loss Functions

from arXiv: Data Structures and Algorithms

Authors: Rahul Vaze

Convex Optimization with Nested Evolving Feasible Sets (CONES)} was introduced in \cite{CONESVaze} where the objective function \(f\) remains fixed but the feasible region evolves over time as a nested sequence \(S_1 \supseteq S_2 \supseteq \cdots \supseteq S_T\). The goal of an online algorithm is to simultaneously minimize the regret with respect to hindsight static optimal benchmark and the total movement cost $M_\cA(T)$ while ensuring feasibility at all times. CONES is an optimization-oriented generalization of the well-known \emph{nested convex body chasing} (NCBC). In this paper, we extend CONES to allow for loss functions $f_t'$s to also change over time. When all loss functions are convex, we show that the projected proximal algorithm achieves $O(T^{1-β}), O(T^β)$ simultaneous regret and movement cost, respectively, for any $β\in [0,1)$, over a time horizon of $T$. We also show that any {\it weakly adaptive} online algorithm with $O(T^β)$ regret has a movement cost of $Ω\left(T^{\frac{1-β}{2}}\right)$ for any $β\in [0,1)$. When all loss functions are strongly convex, we show that the projected proximal algorithm simultaneously achieves $O(1)$ regret and a movement cost of $O(\log T)$. To complement this, we show that any online algorithm with sublinear {\it anytime} regret has a movement cost of $Ω\left(\log T\right)$.

Authors: Rahul Vaze

Convex Optimization with Nested Evolving Feasible Sets (CONES)} was introduced in \cite{CONESVaze} where the objective function \(f\) remains fixed but the feasible region evolves over time as a nested sequence \(S_1 \supseteq S_2 \supseteq \cdots \supseteq S_T\). The goal of an online algorithm is to simultaneously minimize the regret with respect to hindsight static optimal benchmark and the total movement cost $M_\cA(T)$ while ensuring feasibility at all times. CONES is an optimization-oriented generalization of the well-known \emph{nested convex body chasing} (NCBC). In this paper, we extend CONES to allow for loss functions $f_t'$s to also change over time. When all loss functions are convex, we show that the projected proximal algorithm achieves $O(T^{1-β}), O(T^β)$ simultaneous regret and movement cost, respectively, for any $β\in [0,1)$, over a time horizon of $T$. We also show that any {\it weakly adaptive} online algorithm with $O(T^β)$ regret has a movement cost of $Ω\left(T^{\frac{1-β}{2}}\right)$ for any $β\in [0,1)$. When all loss functions are strongly convex, we show that the projected proximal algorithm simultaneously achieves $O(1)$ regret and a movement cost of $O(\log T)$. To complement this, we show that any online algorithm with sublinear {\it anytime} regret has a movement cost of $Ω\left(\log T\right)$.

Single-Exponential Algorithms and a Polynomial Kernel for Strong Connectivity Augmentation

from arXiv: Data Structures and Algorithms

Authors: Tomohiro Koana, Soh Kumabe

Strong Connectivity Augmentation (SCA) asks whether a directed acyclic graph can be made strongly connected by adding at most $k$ prescribed links whose total weight is within a given budget. Klinkby, Misra, and Saurabh (SODA 2021) gave an $O^*(2^{O(k\log k)})$-time algorithm and asked whether the problem admits a single-exponential parameterized algorithm and a polynomial kernel. We answer both questions affirmatively: SCA can be solved in $O^*(9^k)$ time and admits a polynomial kernel with $O(k^4)$ vertices and $O(k^{16})$ bits. For unweighted SCA, we obtain $O^*(4^k)$ time and a kernel with $O(k^3)$ vertices. Our algorithms are based on a particularly simple reduction to Strongly Connected Spanning Subgraph with two edge costs.

Authors: Tomohiro Koana, Soh Kumabe

Strong Connectivity Augmentation (SCA) asks whether a directed acyclic graph can be made strongly connected by adding at most $k$ prescribed links whose total weight is within a given budget. Klinkby, Misra, and Saurabh (SODA 2021) gave an $O^*(2^{O(k\log k)})$-time algorithm and asked whether the problem admits a single-exponential parameterized algorithm and a polynomial kernel. We answer both questions affirmatively: SCA can be solved in $O^*(9^k)$ time and admits a polynomial kernel with $O(k^4)$ vertices and $O(k^{16})$ bits. For unweighted SCA, we obtain $O^*(4^k)$ time and a kernel with $O(k^3)$ vertices. Our algorithms are based on a particularly simple reduction to Strongly Connected Spanning Subgraph with two edge costs.

A deterministic $(1+\varepsilon)^n$ approximation for the permanent of a nonnegative matrix

from arXiv: Data Structures and Algorithms

Authors: Dingding Dong, Vishesh Jain

For every fixed $0<\varepsilon\le1$, we give a deterministic strongly polynomial algorithm that, given a nonnegative matrix $A\in\mathbb{R}_{\ge0}^{n\times n}$, returns $Q$ satisfying $\operatorname{per} A\le Q\le(1+\varepsilon)^n\operatorname{per} A$.

Authors: Dingding Dong, Vishesh Jain

For every fixed $0<\varepsilon\le1$, we give a deterministic strongly polynomial algorithm that, given a nonnegative matrix $A\in\mathbb{R}_{\ge0}^{n\times n}$, returns $Q$ satisfying $\operatorname{per} A\le Q\le(1+\varepsilon)^n\operatorname{per} A$.

Quasi-Monte Carlo Beyond Hardy-Krause II: $(1 + \varepsilon)n$ Samples Suffice

from arXiv: Data Structures and Algorithms

Authors: Ekene Ezeunala, Agastya Vibhuti Jha, Haotian Jiang

Numerical integration studies how well one can estimate the integral of a function $f$ over $[0,1)^d$ using $n$ sample points. The two classical methods, Monte Carlo (MC) and quasi-Monte Carlo (QMC), have complementary strengths and weaknesses, and a fundamental question is to design an approach that combines the benefits of both. Recently, building on the transference principle in discrepancy theory, Bansal and Jiang~\cite{BJ25a} gave a randomized QMC method that bridges MC and QMC guarantees using only i.i.d.\ samples. Their method also goes beyond the classical Koksma--Hlawka inequality: it achieves integration error $\widetilde{O}_d(σ_{\mathsf{SO}}(f)/n)$, where the smoothed-out variation $σ_{\mathsf{SO}}(f)$ can be substantially smaller than the Hardy--Krause variation that governs the classical bound. However, their algorithm requires $n^2$ i.i.d.\ samples as input, and this quadratic blowup is inherent to any method based on the transference principle. In this work, we bypass the quadratic blowup: for any constant $\varepsilon > 0$, we show that $(1+\varepsilon)n$ i.i.d.\ samples suffice to both obtain the beyond-Hardy--Krause guarantee of~\cite{BJ25a}, resolving an open problem posed there, and to produce low-discrepancy point sequences. Our algorithms are variants of the online Haar-thinning method of Dwivedi, Feldheim, Gurel-Gurevich, and Ramdas~\cite{DFG+19}.

Authors: Ekene Ezeunala, Agastya Vibhuti Jha, Haotian Jiang

Numerical integration studies how well one can estimate the integral of a function $f$ over $[0,1)^d$ using $n$ sample points. The two classical methods, Monte Carlo (MC) and quasi-Monte Carlo (QMC), have complementary strengths and weaknesses, and a fundamental question is to design an approach that combines the benefits of both. Recently, building on the transference principle in discrepancy theory, Bansal and Jiang~\cite{BJ25a} gave a randomized QMC method that bridges MC and QMC guarantees using only i.i.d.\ samples. Their method also goes beyond the classical Koksma--Hlawka inequality: it achieves integration error $\widetilde{O}_d(σ_{\mathsf{SO}}(f)/n)$, where the smoothed-out variation $σ_{\mathsf{SO}}(f)$ can be substantially smaller than the Hardy--Krause variation that governs the classical bound. However, their algorithm requires $n^2$ i.i.d.\ samples as input, and this quadratic blowup is inherent to any method based on the transference principle. In this work, we bypass the quadratic blowup: for any constant $\varepsilon > 0$, we show that $(1+\varepsilon)n$ i.i.d.\ samples suffice to both obtain the beyond-Hardy--Krause guarantee of~\cite{BJ25a}, resolving an open problem posed there, and to produce low-discrepancy point sequences. Our algorithms are variants of the online Haar-thinning method of Dwivedi, Feldheim, Gurel-Gurevich, and Ramdas~\cite{DFG+19}.

Navigating Small-World Networks with Distance Predictions

from arXiv: Data Structures and Algorithms

Authors: Ladan Kian, Ming Ming Tan, Dariusz Kowalski

The small-world phenomenon was given an algorithmic foundation by Kleinberg, who showed that in an augmented $k$-dimensional lattice a decentralized greedy algorithm delivers a message in $O(\log^2 n)$ expected steps. We study predicted-greedy routing, in which a mobile agent forwarding the message moves at each step to the neighbor minimizing a noisy $(\varepsilon,δ)$-prediction of its distance to the target, redrawn at every step from an oracle conditioned on the full routing history. Two cases arise from what this agent can observe. An agent with the coordinate awareness can still compute lattice distance exactly, but not graph distance in the shortcut-augmented network, since that depends on the shortcuts of nodes it has not yet visited; given an $(\varepsilon,δ)$-prediction of graph distance, information the classical model never supplies, it achieves expected delivery time $O(\log n/(1-4k\varepsilonδ))$, an asymptotic improvement over $Θ(\log^2 n)$. An agent with no coordinate awareness at all, the natural model for a privacy-preserving network whose nodes never disclose their coordinates, cannot compute even lattice distance; given an $(\varepsilon,δ)$-prediction of lattice distance instead, it still reaches the target in $O(n/(1-4k\varepsilonδ))$ expected steps. Together these results show that a modest amount of predicted information, of the right kind, is enough to accelerate decentralized routing well below Kleinberg's classical bound, and that even when nodes reveal no coordinates at all, reliable delivery remains achievable.

Authors: Ladan Kian, Ming Ming Tan, Dariusz Kowalski

The small-world phenomenon was given an algorithmic foundation by Kleinberg, who showed that in an augmented $k$-dimensional lattice a decentralized greedy algorithm delivers a message in $O(\log^2 n)$ expected steps. We study predicted-greedy routing, in which a mobile agent forwarding the message moves at each step to the neighbor minimizing a noisy $(\varepsilon,δ)$-prediction of its distance to the target, redrawn at every step from an oracle conditioned on the full routing history. Two cases arise from what this agent can observe. An agent with the coordinate awareness can still compute lattice distance exactly, but not graph distance in the shortcut-augmented network, since that depends on the shortcuts of nodes it has not yet visited; given an $(\varepsilon,δ)$-prediction of graph distance, information the classical model never supplies, it achieves expected delivery time $O(\log n/(1-4k\varepsilonδ))$, an asymptotic improvement over $Θ(\log^2 n)$. An agent with no coordinate awareness at all, the natural model for a privacy-preserving network whose nodes never disclose their coordinates, cannot compute even lattice distance; given an $(\varepsilon,δ)$-prediction of lattice distance instead, it still reaches the target in $O(n/(1-4k\varepsilonδ))$ expected steps. Together these results show that a modest amount of predicted information, of the right kind, is enough to accelerate decentralized routing well below Kleinberg's classical bound, and that even when nodes reveal no coordinates at all, reliable delivery remains achievable.

Lower Bounds for Private Graph Optimization Problems using Reconstruction Attacks

from arXiv: Data Structures and Algorithms

Authors: Jacob Imola, Rasmus Pagh, Lukas Retschmeier

This paper studies fundamental graph optimization problems under differential privacy (DP) and shows new, reconstruction-based lower bounds. We consider a graph $G = (V, E, \vec{w})$ where the vertex set $V$ and edges $E$ are public and the weights $\mathbf{w}:E\rightarrow \mathbb{R}$ must be kept differentially private under an $\ell_1$ neighboring relation. For the problems of releasing a minimum-weight spanning tree and a minimum-weight perfect matching, we show new, tight error bounds of $Ω(n\cdot\log(m/n)/ε)$ on worst-case graphs with $n$ vertices and $m>2n$ edges. The upper bounds are known pure DP algorithms while the new lower bound holds even under approximate $(\varepsilon,δ)$-DP as long as $δ\leq (n/m)^{Ω(1)}$. Our lower bounds improve the $Ω(n/ε)$ lower bounds of Sealfon (PODS~'16). The fact that approximate DP does not reduce error for MST under the $\ell_1$ neighboring relation contrasts with the recent upper bound of Pagh et al. (PODS~'25) which shows that approximate DP allows much better error under the $\ell_\infty$ neighboring relation. Going beyond worst-case graphs, we give lower bounds for large families of sparse graphs with expansion properties. We show a lower bound of $Ω(n / ε)$ for the minimum spanning tree for any graph where the minimum cut is at least $Ω(\log(n))$. Finally, we consider the problem of private hierarchical clustering under Dasgupta's cost function (STOC~'16) and show the first approximate DP lower bound parameterized by the minimum weight of a balanced cut. This extends lower bounds of Deng et al. (ICLR~'25) to general graphs and to approximate DP.

Authors: Jacob Imola, Rasmus Pagh, Lukas Retschmeier

This paper studies fundamental graph optimization problems under differential privacy (DP) and shows new, reconstruction-based lower bounds. We consider a graph $G = (V, E, \vec{w})$ where the vertex set $V$ and edges $E$ are public and the weights $\mathbf{w}:E\rightarrow \mathbb{R}$ must be kept differentially private under an $\ell_1$ neighboring relation. For the problems of releasing a minimum-weight spanning tree and a minimum-weight perfect matching, we show new, tight error bounds of $Ω(n\cdot\log(m/n)/ε)$ on worst-case graphs with $n$ vertices and $m>2n$ edges. The upper bounds are known pure DP algorithms while the new lower bound holds even under approximate $(\varepsilon,δ)$-DP as long as $δ\leq (n/m)^{Ω(1)}$. Our lower bounds improve the $Ω(n/ε)$ lower bounds of Sealfon (PODS~'16). The fact that approximate DP does not reduce error for MST under the $\ell_1$ neighboring relation contrasts with the recent upper bound of Pagh et al. (PODS~'25) which shows that approximate DP allows much better error under the $\ell_\infty$ neighboring relation. Going beyond worst-case graphs, we give lower bounds for large families of sparse graphs with expansion properties. We show a lower bound of $Ω(n / ε)$ for the minimum spanning tree for any graph where the minimum cut is at least $Ω(\log(n))$. Finally, we consider the problem of private hierarchical clustering under Dasgupta's cost function (STOC~'16) and show the first approximate DP lower bound parameterized by the minimum weight of a balanced cut. This extends lower bounds of Deng et al. (ICLR~'25) to general graphs and to approximate DP.

Free-Probabilistic State Evolution and Random Matrix Discrepancy

from arXiv: Data Structures and Algorithms

Authors: August Y. Chen, Ahmed El Alaoui

Let $A_1,\ldots,A_n$ be independent $d \times d$ real symmetric Gaussian random matrices, and consider the linear operator $A(x) = n^{-1/2}\sum_{i=1}^n x_i A_i$, $x\in \mathbb{R}^n$. We construct an iterative algorithm in the Approximate Message Passing family which iterates over $A$ and its adjoint $A^*$, and establish a state evolution result which characterizes its behavior in the limit $d\rightarrow \infty, 2n/d^2 \rightarrow α$ in terms of a correlated Gaussian-semicircular process in a free probability space, in the sense of strong convergence of operators. We then apply this iteration to the random matrix discrepancy problem which asks for a binary vector $x \in \{-1,+1\}^n$ such that $A(x)$ has a small operator norm. Our algorithm achieves an operator norm $2σ(α)$, for an explicit expression of the standard deviation $σ(α)<1$ for all $0<α<α_* \simeq 5.74$. This resolves the algorithmic question of Kunisky-Zhang (2023) and Maillard (2025) in this interval.

Authors: August Y. Chen, Ahmed El Alaoui

Let $A_1,\ldots,A_n$ be independent $d \times d$ real symmetric Gaussian random matrices, and consider the linear operator $A(x) = n^{-1/2}\sum_{i=1}^n x_i A_i$, $x\in \mathbb{R}^n$. We construct an iterative algorithm in the Approximate Message Passing family which iterates over $A$ and its adjoint $A^*$, and establish a state evolution result which characterizes its behavior in the limit $d\rightarrow \infty, 2n/d^2 \rightarrow α$ in terms of a correlated Gaussian-semicircular process in a free probability space, in the sense of strong convergence of operators. We then apply this iteration to the random matrix discrepancy problem which asks for a binary vector $x \in \{-1,+1\}^n$ such that $A(x)$ has a small operator norm. Our algorithm achieves an operator norm $2σ(α)$, for an explicit expression of the standard deviation $σ(α)<1$ for all $0<α<α_* \simeq 5.74$. This resolves the algorithmic question of Kunisky-Zhang (2023) and Maillard (2025) in this interval.

Thursday, September 10

Matters of Degree

from Ben Recht

Every forecasting course must engage with the dappled world of probability.

Hi there, argmin readers! As the fall semester picks up, posting volume will, too. So I’m going to commit to writing short descriptive headers to help you sort through the different threads. Today’s post is a live blog of Class 5 of my graduate seminar “Forecasting: A Critical Retrospective.” A table of contents is here.

The fun thing about teaching a new class is realizing in week three how your sequencing is already off. I do this every year, so I’m no longer surprised. At this point in my career, I’d put the probability that I’ll be disappointed with my syllabus before the official add-course deadline at 95%.

How did I get to that number?

So yes, though it’s not directly about forecasting, the first mathematical lecture of this class should have been about probability. Because probability is unavoidable in forecasting, and it is used far too casually for my tastes. We saw this in the weather example where outputs of data assimilation were interpreted as “probability distributions” over the state of the atmosphere. Samples from this distribution were used to create a sample of future possibilities. Frequencies of these future possibilities were turned into beliefs about whether it will rain on Sunday.

All of these probabilities are sort of different, right? Some are about counts, some are about beliefs, and somehow we move between the two as if there is a well-specified set of formal rules for doing so.

I think every class on applied probability, including the undergrad ones, should call out this slippery transmutation between frequency and belief, and I’m particularly fond of the development in Paul Meehl’s course on Philosophical Psychology.1 Meehl follows Rudolf Carnap, who explicitly distinguishes the two kinds of probability.

Probability 1, sometimes called logical probability, relates propositions with beliefs. It quantifies how much credence we should give a hypothesis based on the observations and facts laid before us. Probability 1 represents the certainty of a logical proposition being true, the doubt that something in the past happened, the likelihood that future events will occur, or a person’s internal beliefs about the world.

Probability 2 concerns frequencies of events. It more or less just amounts to counting, measuring the relative frequency of some property in a set of objects. Probability 2 might quantify the relative frequencies of occurrences that currently exist in the world, but it can also describe the relative frequencies of hypothetical infinite populations.

Probability 2 is usually the one we start with and teach and never problematize, and then we casually jump to using Probability 1 in our thinking, writing, and experimenting without realizing it. This is because they both obey the same axioms of probability.

Let’s say I have a bunch of marbles in an urn. I want to describe the frequencies with which certain properties of those marbles hold. The following things are true for any subset of marbles:

  1. Any subset has proportion greater than or equal to 0

  2. If I take the entire set, the proportion equals 1.

  3. If I take two nonoverlapping subsets, the proportion of their union is equal to the sum of their proportions.

These are Kolmogorov’s axioms of probability. Item 1 is nonnegativity, Item 2 is unit measure, Item 3 is additivity.

Probability 2 obviously obeys Kolmogorov’s axioms, but what about Probability 1? Well, we can sort of force it to have those properties too.

  1. Any syntactically valid statement has probability greater than or equal to zero.

  2. Any statement that is certain or tautological has probability 1.

  3. If two statements describe mutually exclusive outcomes, then the probability of either one or both of the outcomes is equal to the sum of the individual probabilities.

Lo and behold, statements of belief seem to also obey Kolmogorov’s Axioms. These assertions might feel a bit less obvious and certain than marble counting. You can’t get a tangible grip on why beliefs should obey Kolmogorov’s axioms because beliefs live in your head. You can check the properties of Probability 2 on a set of marbles. You can never check whether the axioms work for Probability 1.

But there are strategic reasons for using probabilities when quantifying beliefs. In last Thursday’s lecture, we saw that if you are scored using the Brier Score, then when your forecasts don’t obey the axioms of probability, there is always another set of forecasts that achieves a higher place on the leaderboard. If instead of being a forecaster, you’re a degenerate gambler, there are arguments about dealing with bookies that motivate the same probabilistic rules. As Dennis Lindley put it, probability is inevitable once you try to quantitatively evaluate belief.

I think we should also distinguish a third kind of probability, Probability 0, to describe the formal language of mathematical probability. This is probability whose referent is neither frequencies nor beliefs but mathematics itself. In this case, we’d consider the following to all be Probability 0:

  1. Any finite list of numbers that is nonnegative and sums to one

  2. Any infinite list of numbers that is nonnegative and whose infinite sum converges to one

  3. Any nonnegative function on the unit interval whose integral is equal to one

These sorts of mathematical structures arise a lot when we’re doing calculations. And whenever people find a convenient mathematical application of Probability 0, they tend to find a convenient application of that structure in Probability 1 or Probability 2. We see Probability 0 everywhere. It only becomes metaphysical when we attach some meaning to it off the chalkboard.

When we build computational forecasts, we have to play with all three kinds of probability. We saw this last time: We need Probability 2 because the best predictions correspond to the rates of outcomes in similar future events. We need Probability 1, or else our forecasts are incoherent. We need Probability 0 to write and reason about algorithms that analyze noisy data. But how we tie the three together is not given. There is no god-given algorithm of prognostication that we can derive from Kolmogorov’s axioms, and techniques and conventions vary between disciplines. Which scoring rule are we using? Are we insisting on building calibrated forecasts? What do our forecasts do? These questions determine how we work with probability.

Subscribe now

1

The videos are here, and my blog series engaging with the course is here.

By Ben Recht

TR26-173 | Matrix identities are hard: Fast blackbox PIT for noncommutative exponential-size constant-depth homogeneous circuits | Foram Lakhani, Nitin Saxena

from ECCC Papers

In this work we give a randomized blackbox polynomial identity testing (PIT) algorithm for constant-depth homogeneous noncommutative circuits, with poly-logarithmic time complexity in the circuit size. In fact, we show that the polynomial computable by such a depth-$(\Delta-1)$ circuit of size $s$ cannot be a polynomial identity for the $O(\log ^{\Delta}s)$-dimension matrix algebra $\mathcal{M}_{O(\log ^{\Delta}s)}(\mathbb{F})$, for a sufficiently large field $\mathbb{F}$. This represents progress toward answering the question raised by Arvind, Joglekar, Mukhopadhyay, and Raja (2019), and Bogdanov and Wee (2005) to find efficient randomized blackbox PIT algorithm for circuits that output polynomials with $doubly$ exponential sparsity. Arvind, Joglekar, Mukhopadhyay, and Raja (2019) and Bharadwaj and S. Raja (2025) defined constant-depth $+$-regular circuit and gave a PIT for it. Their model of size $S$ is simulated by our model of size $\exp(S)$; hence, we recover their result. Because of the regularity condition their model is highly restrictive than ours. We decompose a depth-$\Delta$ formula into its constituent depth-$(\Delta - 1)$ subformulas and analyze them concurrently using Hadamard algebra, subsequently translating these findings back to the depth-$\Delta$ setting. We reduce the multivariate case to the bivariate case involving the noncommuting variables $x$ and $y$. Finally, we apply a transformation to $x$ and $y$ to ensure that the nonzero witness coefficient is associated with a monomial of exponentially-lower $y$-degree. We call this the phenomenon of $noncommutative\ low-degree\ concentration$.
In this work we give a randomized blackbox polynomial identity testing (PIT) algorithm for constant-depth homogeneous noncommutative circuits, with poly-logarithmic time complexity in the circuit size. In fact, we show that the polynomial computable by such a depth-$(\Delta-1)$ circuit of size $s$ cannot be a polynomial identity for the $O(\log ^{\Delta}s)$-dimension matrix algebra $\mathcal{M}_{O(\log ^{\Delta}s)}(\mathbb{F})$, for a sufficiently large field $\mathbb{F}$. This represents progress toward answering the question raised by Arvind, Joglekar, Mukhopadhyay, and Raja (2019), and Bogdanov and Wee (2005) to find efficient randomized blackbox PIT algorithm for circuits that output polynomials with $doubly$ exponential sparsity. Arvind, Joglekar, Mukhopadhyay, and Raja (2019) and Bharadwaj and S. Raja (2025) defined constant-depth $+$-regular circuit and gave a PIT for it. Their model of size $S$ is simulated by our model of size $\exp(S)$; hence, we recover their result. Because of the regularity condition their model is highly restrictive than ours. We decompose a depth-$\Delta$ formula into its constituent depth-$(\Delta - 1)$ subformulas and analyze them concurrently using Hadamard algebra, subsequently translating these findings back to the depth-$\Delta$ setting. We reduce the multivariate case to the bivariate case involving the noncommuting variables $x$ and $y$. Finally, we apply a transformation to $x$ and $y$ to ensure that the nonzero witness coefficient is associated with a monomial of exponentially-lower $y$-degree. We call this the phenomenon of $noncommutative\ low-degree\ concentration$.

On the Limits of Quantum Multiparty Simultaneous Communication

from arXiv: Computational Complexity

Authors: Pedro Montealegre, Ivan Rapaport, Jorge Valenzuela

The Simultaneous Message Passing (SMP) model provides a fundamental framework for comparing classical and quantum communication. For two players, Gavinsky et al. (STOC 2006) established a separation underlying the incomparability of shared randomness and quantum communication: \textsc{Index Coordination} needs $O(\log n)$ public-coin bits but $Ω(n^{1/3})$ bounded-error qubits. In this work, we establish a multiparty exponential separation through $\operatorname{IC}_{k,n}$, a natural $k$-party generalization of \textsc{Index Coordination}. Public-coin protocols solve it unambiguously with maximum message length $O(\log n)$ bits. In contrast, quantum SMP protocols without shared entanglement or public coins require maximum message length $Ω(n^{1-1/k})$ qubits in the unambiguous regime and $Ω(n^{(k-1)/(k+1)})$ qubits in the bounded-error regime. A classical private-coin protocol matches the unambiguous bound, so quantum communication provides no asymptotic advantage over private randomness in this regime. For fixed error parameters, all constants are independent of $k$, establishing the exponential separation for every integer-valued function $k=k(n)\ge2$, without restricting its growth. Both quantum lower bounds become $Ω(n)$ when $k\ge c\log n$ for any fixed $c>0$, matching the full-input protocol and yielding tight linear complexity in both regimes. Our results demonstrate that quantum superposition cannot efficiently simulate the coordination afforded by public randomness, extending this separation to arbitrary $k$. To bound success probabilities for multiparty product states, we prove an exact factorization theorem for unambiguous quantum state identification, which may be of independent mathematical interest.

Authors: Pedro Montealegre, Ivan Rapaport, Jorge Valenzuela

The Simultaneous Message Passing (SMP) model provides a fundamental framework for comparing classical and quantum communication. For two players, Gavinsky et al. (STOC 2006) established a separation underlying the incomparability of shared randomness and quantum communication: \textsc{Index Coordination} needs $O(\log n)$ public-coin bits but $Ω(n^{1/3})$ bounded-error qubits. In this work, we establish a multiparty exponential separation through $\operatorname{IC}_{k,n}$, a natural $k$-party generalization of \textsc{Index Coordination}. Public-coin protocols solve it unambiguously with maximum message length $O(\log n)$ bits. In contrast, quantum SMP protocols without shared entanglement or public coins require maximum message length $Ω(n^{1-1/k})$ qubits in the unambiguous regime and $Ω(n^{(k-1)/(k+1)})$ qubits in the bounded-error regime. A classical private-coin protocol matches the unambiguous bound, so quantum communication provides no asymptotic advantage over private randomness in this regime. For fixed error parameters, all constants are independent of $k$, establishing the exponential separation for every integer-valued function $k=k(n)\ge2$, without restricting its growth. Both quantum lower bounds become $Ω(n)$ when $k\ge c\log n$ for any fixed $c>0$, matching the full-input protocol and yielding tight linear complexity in both regimes. Our results demonstrate that quantum superposition cannot efficiently simulate the coordination afforded by public randomness, extending this separation to arbitrary $k$. To bound success probabilities for multiparty product states, we prove an exact factorization theorem for unambiguous quantum state identification, which may be of independent mathematical interest.

On the Tightness of Standard Relaxations for Mixed-Integer Bilevel Linear Programs

from arXiv: Computational Complexity

Authors: Sergey S. Ketkov, Oleg A. Prokopyev

Exact algorithms for solving mixed-integer bilevel linear programs (MIBLPs) typically rely on sequences of lower and upper bounds that converge to the optimal value. These procedures are commonly initialized using the single-level relaxation (SLR), obtained by omitting the follower's optimality condition and solving the resulting single-level optimization problem. In this paper, we investigate whether, for broad classes of MIBLPs, the resulting standard bounds admit uniform improvements that can be computed within the same computational complexity regime. For pure continuous bilevel linear programs, we show that, unless $P = NP$, neither the SLR-based lower bound nor its associated upper bound can be uniformly improved in polynomial time, even for the class of min-max problems. We then extend this analysis to the class of pure integer min-max bilevel linear programs under the assumption that the polynomial hierarchy does not collapse. First, we show that the continuous relaxation of the SLR admits no uniform polynomial-time computable improvement. We then prove that neither the SLR itself nor its associated upper bound admits a uniform improvement by a polynomial-time algorithm with access to a mixed-integer linear programming (MILP) oracle. Importantly, this rules out uniform improvements by iterative MILP-based approaches, including cutting-plane-based and decomposition algorithms. Overall, our results demonstrate that the SLR-based bounds are, in a complexity-theoretic sense, unimprovable systematically within their natural computational regimes.

Authors: Sergey S. Ketkov, Oleg A. Prokopyev

Exact algorithms for solving mixed-integer bilevel linear programs (MIBLPs) typically rely on sequences of lower and upper bounds that converge to the optimal value. These procedures are commonly initialized using the single-level relaxation (SLR), obtained by omitting the follower's optimality condition and solving the resulting single-level optimization problem. In this paper, we investigate whether, for broad classes of MIBLPs, the resulting standard bounds admit uniform improvements that can be computed within the same computational complexity regime. For pure continuous bilevel linear programs, we show that, unless $P = NP$, neither the SLR-based lower bound nor its associated upper bound can be uniformly improved in polynomial time, even for the class of min-max problems. We then extend this analysis to the class of pure integer min-max bilevel linear programs under the assumption that the polynomial hierarchy does not collapse. First, we show that the continuous relaxation of the SLR admits no uniform polynomial-time computable improvement. We then prove that neither the SLR itself nor its associated upper bound admits a uniform improvement by a polynomial-time algorithm with access to a mixed-integer linear programming (MILP) oracle. Importantly, this rules out uniform improvements by iterative MILP-based approaches, including cutting-plane-based and decomposition algorithms. Overall, our results demonstrate that the SLR-based bounds are, in a complexity-theoretic sense, unimprovable systematically within their natural computational regimes.

When Does a Quantum Speedup Survive End-to-End?

from arXiv: Computational Complexity

Authors: Pablo Herrero Gómez, Antonio Jimeno Morenilla, David Muñoz Hernández, Higinio Mora Mora

Primitive quantum speedups are interface-relative: they depend on the input access used to run the primitive and on the output contract used to consume its state or samples. This paper introduces a transcript-level admissibility relation \(A_M\preceq_{\mathrm{int}}A_Q\), defined relative to the declared implementation package of the quantum interface. It identifies which adaptive classical access transcripts that same package licenses, with all setup, transcript-generation, and precision overheads charged. The main application is an operational audit for normalized-Betti estimation in clique-complex TDA, separating three declared-interface regimes. Reversible indexed simplex interfaces certify matched classical simplex sampling and local Laplacian row access by evaluating their reversible routines on single computational branches. Membership-based preparations induce a rejection route of overhead \(\binom{n}{k+1}/|S_k|\). Abstract spectral or block-encoding interfaces require an accompanying implementation package, transcript reduction, or shared representation. Under the indexed certificate and interface closure, the end-to-end cost is fixed by the imported estimator's spectral dependence on the gap \(γ\); the concretely realized bounded-treewidth family already admits exact \(\mathrm{poly}(n)\) classical Betti computation by rank over \(\mathbb{Q}\). A low-rank separation supports the role of access and output contracts.

Authors: Pablo Herrero Gómez, Antonio Jimeno Morenilla, David Muñoz Hernández, Higinio Mora Mora

Primitive quantum speedups are interface-relative: they depend on the input access used to run the primitive and on the output contract used to consume its state or samples. This paper introduces a transcript-level admissibility relation \(A_M\preceq_{\mathrm{int}}A_Q\), defined relative to the declared implementation package of the quantum interface. It identifies which adaptive classical access transcripts that same package licenses, with all setup, transcript-generation, and precision overheads charged. The main application is an operational audit for normalized-Betti estimation in clique-complex TDA, separating three declared-interface regimes. Reversible indexed simplex interfaces certify matched classical simplex sampling and local Laplacian row access by evaluating their reversible routines on single computational branches. Membership-based preparations induce a rejection route of overhead \(\binom{n}{k+1}/|S_k|\). Abstract spectral or block-encoding interfaces require an accompanying implementation package, transcript reduction, or shared representation. Under the indexed certificate and interface closure, the end-to-end cost is fixed by the imported estimator's spectral dependence on the gap \(γ\); the concretely realized bounded-treewidth family already admits exact \(\mathrm{poly}(n)\) classical Betti computation by rank over \(\mathbb{Q}\). A low-rank separation supports the role of access and output contracts.

Small-Bias Quantum Approximate Counting via the Multiplicative Adversary Method

from arXiv: Computational Complexity

Authors: Albert Lin, Han-Hsuan Lin

We study the two-weight decision version of quantum approximate counting: given oracle access to $x\in\{0,1\}^N$, distinguish $|x|=M$ from $|x|=M+Δ$ with success probability $1/2+ζ$. Using the multiplicative adversary method, we prove $Ω\left(\max\left\{ζ\sqrt{(N-M)(M+Δ)}/Δ,\sqrt{ζN/Δ}\right\}\right)$. The same parameter dependence follows from the polynomial-method characterization of the two-layer symmetric function by Podder, Yao, and Ye. Our contribution is a multiplicative-adversary derivation that tracks the progress produced by individual oracle queries. For the first term, after complementing the input if necessary, we assume $M+Δ\le N-M$. We use the Hamming-layer subspaces from the eigenspace method of Ambainis, Spalek, and de Wolf and compose their adjacent-layer unitary maps to relate the two nonadjacent promise layers. After fixing the queried coordinate, the analysis block-diagonalizes into four-dimensional subspaces. An exact calculation of the one-query progress ratio gives the first lower bound. The same estimate also implies $\left\|(I-\widehatΠ_{\mathrm{bad}})\lvertΨ^T\rangle\right\|^2=O\left(T^2Δ^2/((N-M)(M+Δ))\right)$ for the coherent input superposition used in the adversary argument. For the second term, we prove directly using a three-eigenvalue multiplicative adversary that unique OR on $n$ bits with success probability $1/2+ζ$ requires $Ω(\sqrt{ζn})$ queries, and then reduce unique OR to the two-weight counting problem.

Authors: Albert Lin, Han-Hsuan Lin

We study the two-weight decision version of quantum approximate counting: given oracle access to $x\in\{0,1\}^N$, distinguish $|x|=M$ from $|x|=M+Δ$ with success probability $1/2+ζ$. Using the multiplicative adversary method, we prove $Ω\left(\max\left\{ζ\sqrt{(N-M)(M+Δ)}/Δ,\sqrt{ζN/Δ}\right\}\right)$. The same parameter dependence follows from the polynomial-method characterization of the two-layer symmetric function by Podder, Yao, and Ye. Our contribution is a multiplicative-adversary derivation that tracks the progress produced by individual oracle queries. For the first term, after complementing the input if necessary, we assume $M+Δ\le N-M$. We use the Hamming-layer subspaces from the eigenspace method of Ambainis, Spalek, and de Wolf and compose their adjacent-layer unitary maps to relate the two nonadjacent promise layers. After fixing the queried coordinate, the analysis block-diagonalizes into four-dimensional subspaces. An exact calculation of the one-query progress ratio gives the first lower bound. The same estimate also implies $\left\|(I-\widehatΠ_{\mathrm{bad}})\lvertΨ^T\rangle\right\|^2=O\left(T^2Δ^2/((N-M)(M+Δ))\right)$ for the coherent input superposition used in the adversary argument. For the second term, we prove directly using a three-eigenvalue multiplicative adversary that unique OR on $n$ bits with success probability $1/2+ζ$ requires $Ω(\sqrt{ζn})$ queries, and then reduce unique OR to the two-weight counting problem.

NP-Hardness of the $H$-Free Edge-Deletion Problem

from arXiv: Computational Complexity

Authors: Lior Gishboliner, Ethan Honest

For a graph $H$, the $H$-freeness edge-deletion problem is the algorithmic problem of finding, for an input graph $G$, the minimum number of edges of $G$ whose deletion turns $G$ into an $H$-free graph. We show that for every graph $H$ containing a cycle, this problem is NP-hard. This proves a conjecture of Gishboliner, Levanzov and Shapira, and completes the characterization of the complexity of the $H$-freeness edge-deletion problem, answering a question of Alon, Shapira and Sudakov.

Authors: Lior Gishboliner, Ethan Honest

For a graph $H$, the $H$-freeness edge-deletion problem is the algorithmic problem of finding, for an input graph $G$, the minimum number of edges of $G$ whose deletion turns $G$ into an $H$-free graph. We show that for every graph $H$ containing a cycle, this problem is NP-hard. This proves a conjecture of Gishboliner, Levanzov and Shapira, and completes the characterization of the complexity of the $H$-freeness edge-deletion problem, answering a question of Alon, Shapira and Sudakov.

A Note on the Point-Clothoid Distance Algorithm

from arXiv: Computational Geometry

Authors: Haibin Ye, Hao Ge, Gong Cheng

Computing the closest point on a clothoid is a recurring task in geometric design, road and railway alignment, and path planning. The efficient algorithm of Frego and Bertolazzi addresses this problem, but its candidate-selection analysis assumes at most one local minimum per search interval. We exhibit admissible configurations with two local minima, raising the question of whether the existing strategy accounts for every possible minimum. Using the geometry of the clothoid evolute, we prove that, for any query point and any proper no-inflection planar clothoid segment with tangent-angle variation at most $2π$, the squared-distance function has at most three stationary points; if all three are local extrema, their order is min-max-min. This establishes the completeness of the original candidate-selection logic beyond the one-minimum premise. It also shows that no interior search is needed when neither endpoint derivative test is active, allowing unnecessary midpoint searches to be omitted while retaining numerical fallback. Numerical experiments demonstrate reductions in iteration count and evaluation time.

Authors: Haibin Ye, Hao Ge, Gong Cheng

Computing the closest point on a clothoid is a recurring task in geometric design, road and railway alignment, and path planning. The efficient algorithm of Frego and Bertolazzi addresses this problem, but its candidate-selection analysis assumes at most one local minimum per search interval. We exhibit admissible configurations with two local minima, raising the question of whether the existing strategy accounts for every possible minimum. Using the geometry of the clothoid evolute, we prove that, for any query point and any proper no-inflection planar clothoid segment with tangent-angle variation at most $2π$, the squared-distance function has at most three stationary points; if all three are local extrema, their order is min-max-min. This establishes the completeness of the original candidate-selection logic beyond the one-minimum premise. It also shows that no interior search is needed when neither endpoint derivative test is active, allowing unnecessary midpoint searches to be omitted while retaining numerical fallback. Numerical experiments demonstrate reductions in iteration count and evaluation time.

A 2.37332-Competitive Algorithm for Online Square Packing with Gravity

from arXiv: Computational Geometry

Authors: Nichlas Langhoff Rasmussen

We consider online packing of axis-parallel squares into a unit-width strip under the Tetris and gravity constraints: An incoming square must be lowered from above along a monotonic downwards path until it reaches support from below. Fekete, Kamphans, and Schweer [Algorithmica, 2014] gave an algorithm with asymptotic competitive ratio $34/13\approx2.6154$ in this model. We present $\mathrm{AsymmetricSlots}$, a recursive algorithm based on splitting each slot into a wide and narrow subslot. The proof uses a local charging argument: squares that are large relative to its associated slot pay for the height they create with their own area, while smaller squares are balanced between the two subslots and may use a bounded temporary credit. For a suitable split parameter $p^\star$, we prove $\mathrm{AsymmetricSlots}{p^\star}(σ)\le 2.37332 \operatorname{OPT} (σ)+O(1)$ for every input sequence $σ$. Additionally, we show that the same framework gives an algorithm with asymptotic competitive ratio $O(κ)$ for rectangles of aspect ratio at most $κ$, and a matching $ Ω(κ)$ lower bound shows that the dependence on $κ$ is asymptotically optimal. For the square algorithm, we give a lower bound of $2$ on its asymptotic competitive ratio.

Authors: Nichlas Langhoff Rasmussen

We consider online packing of axis-parallel squares into a unit-width strip under the Tetris and gravity constraints: An incoming square must be lowered from above along a monotonic downwards path until it reaches support from below. Fekete, Kamphans, and Schweer [Algorithmica, 2014] gave an algorithm with asymptotic competitive ratio $34/13\approx2.6154$ in this model. We present $\mathrm{AsymmetricSlots}$, a recursive algorithm based on splitting each slot into a wide and narrow subslot. The proof uses a local charging argument: squares that are large relative to its associated slot pay for the height they create with their own area, while smaller squares are balanced between the two subslots and may use a bounded temporary credit. For a suitable split parameter $p^\star$, we prove $\mathrm{AsymmetricSlots}{p^\star}(σ)\le 2.37332 \operatorname{OPT} (σ)+O(1)$ for every input sequence $σ$. Additionally, we show that the same framework gives an algorithm with asymptotic competitive ratio $O(κ)$ for rectangles of aspect ratio at most $κ$, and a matching $ Ω(κ)$ lower bound shows that the dependence on $κ$ is asymptotically optimal. For the square algorithm, we give a lower bound of $2$ on its asymptotic competitive ratio.

Overlap-Helly theorems

from arXiv: Computational Geometry

Authors: Andreas F. Holmsen, Alfredo Hubard

In this paper we introduce a generalization of Helly's theorem closely connected to Bárány-Gromov overlap theorems (also called selection lemmas). Our main result implies both the topological colorful Helly of Kalai and Meschulam and Karasev's topological centerpoint theorem. We further investigate the topological fractional Helly theorem from this overlap perspective, and show an overlap theorem for dense complexes (a continuous second selection lemma for tame maps).

Authors: Andreas F. Holmsen, Alfredo Hubard

In this paper we introduce a generalization of Helly's theorem closely connected to Bárány-Gromov overlap theorems (also called selection lemmas). Our main result implies both the topological colorful Helly of Kalai and Meschulam and Karasev's topological centerpoint theorem. We further investigate the topological fractional Helly theorem from this overlap perspective, and show an overlap theorem for dense complexes (a continuous second selection lemma for tame maps).

The Hyperbolic Surface Distance, Diameter, and Dirichlet Problems

from arXiv: Computational Geometry

Authors: Vincent Despre, Auguste Gezalyan, Marc Pouget

Despite the prominence of hyperbolic surfaces in mathematics, basic algorithmic questions about them, even computing the distance between two points, have remained open, leaving many features of these surfaces inaccessible. The classical machinery assumes a polyhedral structure absent on a smooth surface. We remove these obstacles. We begin with an efficient $O(g^2)$ algorithm for the distance between two points, where $g$ is the genus of the surface. Building on it, we obtain an $O(g^2 \log g)$ method for answering distance queries from a fixed source and, as a consequence, for recentering a Dirichlet domain around an arbitrary point. This understanding of distances on the surface then lets us approximate the diameter to within any $\eps$ in time $O(g^3 \log g / \eps^2)$. We further show that the diameter, a single real number encoding a great deal about the surface, is exactly computable. Its hyperbolic cosine is an algebraic number over the field encoding the coefficients of the hyperbolic isometries defining the surface.

Authors: Vincent Despre, Auguste Gezalyan, Marc Pouget

Despite the prominence of hyperbolic surfaces in mathematics, basic algorithmic questions about them, even computing the distance between two points, have remained open, leaving many features of these surfaces inaccessible. The classical machinery assumes a polyhedral structure absent on a smooth surface. We remove these obstacles. We begin with an efficient $O(g^2)$ algorithm for the distance between two points, where $g$ is the genus of the surface. Building on it, we obtain an $O(g^2 \log g)$ method for answering distance queries from a fixed source and, as a consequence, for recentering a Dirichlet domain around an arbitrary point. This understanding of distances on the surface then lets us approximate the diameter to within any $\eps$ in time $O(g^3 \log g / \eps^2)$. We further show that the diameter, a single real number encoding a great deal about the surface, is exactly computable. Its hyperbolic cosine is an algebraic number over the field encoding the coefficients of the hyperbolic isometries defining the surface.

iLogMap: Geodesic Polar Coordinates Parameterization with the Magnetic Laplacian

from arXiv: Computational Geometry

Authors: Tomás Banduc, Simone Pezzuto, Francisco Sahli Costabal

Geodesic polar coordinates (GPCs) provide an intrinsic parameterization over curved surfaces, but their accurate estimation remains challenging, particularly in the presence of anisotropic metrics, high curvature and complex topology. We introduce iLogMap, a method for computing GPCs in curved domains that recasts the angular component of the logarithmic map to a ground-state magnetic eigenproblem over the circumferential direction field of geodesic distance. Our method effortlessly extends to anisotropic metric tensors and solid volumes, enabling cylindrical and spherical parameterizations in tetrahedral meshes. Experiments on diverse shapes with varying genus confirm competitive angular accuracy and reduced metric distortion relative to heat-based methods, with improved performance on surfaces with boundary and domains with anisotropy. We demonstrate the utility of iLogMap in computational cardiology applications, where we use it to initialize spiral phases on atrial surfaces and estimate local activation patterns in ventricular models.

Authors: Tomás Banduc, Simone Pezzuto, Francisco Sahli Costabal

Geodesic polar coordinates (GPCs) provide an intrinsic parameterization over curved surfaces, but their accurate estimation remains challenging, particularly in the presence of anisotropic metrics, high curvature and complex topology. We introduce iLogMap, a method for computing GPCs in curved domains that recasts the angular component of the logarithmic map to a ground-state magnetic eigenproblem over the circumferential direction field of geodesic distance. Our method effortlessly extends to anisotropic metric tensors and solid volumes, enabling cylindrical and spherical parameterizations in tetrahedral meshes. Experiments on diverse shapes with varying genus confirm competitive angular accuracy and reduced metric distortion relative to heat-based methods, with improved performance on surfaces with boundary and domains with anisotropy. We demonstrate the utility of iLogMap in computational cardiology applications, where we use it to initialize spiral phases on atrial surfaces and estimate local activation patterns in ventricular models.

On the Parameterized Complexity of Coloring Discovery

from arXiv: Data Structures and Algorithms

Authors: Eric Decker, Sebastian Siebertz

Coloring Discovery asks whether a possibly improper initial coloring can be made proper within a prescribed number of allowed changes. We study the parameterized complexity of three modification step models that were studied previously in the literature: recoloring one vertex (color flipping), swapping the colors of arbitrary vertices (color swapping), and swapping colors only across an edge (color sliding). For color flipping, we give exact fixed-parameter algorithms for the parameters vertex cover and distance to complete. For color swapping, we obtain fixed-parameter tractability for the parameter vertex cover plus the number of colors. Our lower bounds show W[1]-hardness for treedepth plus feedback vertex set in the color flipping model and for the number of colors plus bandwidth or distance to disjoint paths in the swapping and sliding models. All three variants remain NP-complete with four colors on graphs of diameter two.

Authors: Eric Decker, Sebastian Siebertz

Coloring Discovery asks whether a possibly improper initial coloring can be made proper within a prescribed number of allowed changes. We study the parameterized complexity of three modification step models that were studied previously in the literature: recoloring one vertex (color flipping), swapping the colors of arbitrary vertices (color swapping), and swapping colors only across an edge (color sliding). For color flipping, we give exact fixed-parameter algorithms for the parameters vertex cover and distance to complete. For color swapping, we obtain fixed-parameter tractability for the parameter vertex cover plus the number of colors. Our lower bounds show W[1]-hardness for treedepth plus feedback vertex set in the color flipping model and for the number of colors plus bandwidth or distance to disjoint paths in the swapping and sliding models. All three variants remain NP-complete with four colors on graphs of diameter two.

Optimal Non-Adaptive Vantage Point Selection

from arXiv: Data Structures and Algorithms

Authors: Jie Gao, Nicole Wein, Chang Wu

We study the \emph{vantage point selection} problem, introduced by Ashvinkumar, Chowdhury, Gao, Goswami, Mitchell, and Polishchuk [WADS'25] to model the problem of estimating bottleneck capacities on the Internet. The input is a weighted undirected graph with unique shortest paths where every edge has a distinct unknown \emph{capacity}. When the algorithm \emph{queries} a vertex $v$, it reveals the minimum-capacity edge on the shortest path from $v$ to every other vertex reachable from $v$. The goal is to maximize the total number of revealed edges. The quality of an algorithm is measured by its competitive ratio against an optimal algorithm that knows all edge capacities a priori. We first consider the foundational single-query setting, where both the algorithm and the optimal algorithm are restricted to a single query. There is a trivial upper bound of $O(n)$ on the competitive ratio and the best known lower bound was $\tildeΩ(\sqrt{n})$. We provide an algorithm and matching lower bound (up to polylogarithmic factors) showing that the best possible competitive ratio is $\tildeΘ(n^{2/3})$. Furthermore, we extend our results to the general setting where the optimal algorithm is allowed $k$ queries and our algorithm is allowed $αk$ queries for $α\geq 1$. We present a randomized non-adaptive algorithm and matching lower bound (up to polylogarithmic factors) showing that the best possible expected competitive ratio for non-adaptive algorithms is the following surprisingly complex bound: $$ \tildeΘ\left( \min\left\{ \frac{n}{αk}, \max\left( \sqrt{\frac{n}α}, \frac{n^{2/3}}{αk^{1/3}} \right) \right\} \right). $$

Authors: Jie Gao, Nicole Wein, Chang Wu

We study the \emph{vantage point selection} problem, introduced by Ashvinkumar, Chowdhury, Gao, Goswami, Mitchell, and Polishchuk [WADS'25] to model the problem of estimating bottleneck capacities on the Internet. The input is a weighted undirected graph with unique shortest paths where every edge has a distinct unknown \emph{capacity}. When the algorithm \emph{queries} a vertex $v$, it reveals the minimum-capacity edge on the shortest path from $v$ to every other vertex reachable from $v$. The goal is to maximize the total number of revealed edges. The quality of an algorithm is measured by its competitive ratio against an optimal algorithm that knows all edge capacities a priori. We first consider the foundational single-query setting, where both the algorithm and the optimal algorithm are restricted to a single query. There is a trivial upper bound of $O(n)$ on the competitive ratio and the best known lower bound was $\tildeΩ(\sqrt{n})$. We provide an algorithm and matching lower bound (up to polylogarithmic factors) showing that the best possible competitive ratio is $\tildeΘ(n^{2/3})$. Furthermore, we extend our results to the general setting where the optimal algorithm is allowed $k$ queries and our algorithm is allowed $αk$ queries for $α\geq 1$. We present a randomized non-adaptive algorithm and matching lower bound (up to polylogarithmic factors) showing that the best possible expected competitive ratio for non-adaptive algorithms is the following surprisingly complex bound: $$ \tildeΘ\left( \min\left\{ \frac{n}{αk}, \max\left( \sqrt{\frac{n}α}, \frac{n^{2/3}}{αk^{1/3}} \right) \right\} \right). $$

Introvert Clustering for Distributed Graph Algorithms

from arXiv: Data Structures and Algorithms

Authors: Yi-Jun Chang, Nima Dolatabadi

We introduce a graph decomposition primitive called introvert clustering, which strengthens standard low-diameter clustering by guaranteeing that every clustered vertex keeps at least a $\left(\frac12-\varepsilon\right)$-fraction of its relevant neighbors in its own cluster. Repeatedly applying this primitive yields a layered introvert network decomposition with $O(\log n)$ layers and weak diameter $O(\log n)$. We give two applications in the $\mathsf{LOCAL}$ model. For every constant $\varepsilon>0$, we obtain a $\widetilde O(\log^2 n)$-round deterministic algorithm for list $\left(\frac32+\varepsilon\right)Δ$-edge coloring on graphs of maximum degree $Δ\geqΔ_0(\varepsilon)$; for bipartite graphs, the result holds for all $Δ$. For every constant $0<\varepsilon<1/4$, we also obtain a $\widetilde O(\log^2 n)$-round deterministic algorithm for a $\left(\frac14-\varepsilon\right)$-locally balanced cut, where every vertex has at least a $\left(\frac14-\varepsilon\right)$-fraction of its neighbors on the opposite side. The resulting algorithms are remarkably simple: edge coloring processes the layers in reverse order and colors each cluster, while locally balanced cut processes them forward and computes a locally maximum cut within each cluster. The introvert guarantee enables these procedures beyond the usual greedy regime of network decomposition. We construct the decomposition in $O(\log^2 n)$ randomized rounds using Miller--Peng--Xu low-diameter clustering and a simple trimming procedure, and deterministically in $\widetilde O(\log^2 n)$ rounds via a white-box adaptation of the recursive network decomposition algorithm of Ghaffari and Grunau [FOCS 2024].

Authors: Yi-Jun Chang, Nima Dolatabadi

We introduce a graph decomposition primitive called introvert clustering, which strengthens standard low-diameter clustering by guaranteeing that every clustered vertex keeps at least a $\left(\frac12-\varepsilon\right)$-fraction of its relevant neighbors in its own cluster. Repeatedly applying this primitive yields a layered introvert network decomposition with $O(\log n)$ layers and weak diameter $O(\log n)$. We give two applications in the $\mathsf{LOCAL}$ model. For every constant $\varepsilon>0$, we obtain a $\widetilde O(\log^2 n)$-round deterministic algorithm for list $\left(\frac32+\varepsilon\right)Δ$-edge coloring on graphs of maximum degree $Δ\geqΔ_0(\varepsilon)$; for bipartite graphs, the result holds for all $Δ$. For every constant $0<\varepsilon<1/4$, we also obtain a $\widetilde O(\log^2 n)$-round deterministic algorithm for a $\left(\frac14-\varepsilon\right)$-locally balanced cut, where every vertex has at least a $\left(\frac14-\varepsilon\right)$-fraction of its neighbors on the opposite side. The resulting algorithms are remarkably simple: edge coloring processes the layers in reverse order and colors each cluster, while locally balanced cut processes them forward and computes a locally maximum cut within each cluster. The introvert guarantee enables these procedures beyond the usual greedy regime of network decomposition. We construct the decomposition in $O(\log^2 n)$ randomized rounds using Miller--Peng--Xu low-diameter clustering and a simple trimming procedure, and deterministically in $\widetilde O(\log^2 n)$ rounds via a white-box adaptation of the recursive network decomposition algorithm of Ghaffari and Grunau [FOCS 2024].

Minimum-makespan completion and vertex selection leave the Wang-Sitters constant at 11/6

from arXiv: Data Structures and Algorithms

Authors: Adam Y. Shavit

The 11/6 worst-case constant of the Wang-Sitters rounding scheme, which a companion note establishes, can naturally be attributed to the freedom in Step 3, where an arbitrary valid slot matching is permitted. We show that eliminating that freedom does not improve the constant. A minimum-makespan completion oracle still has worst-case constant exactly 11/6 against the optimum; both natural 7/4 statements about it are false; and restricting Step 1 to vertices of the relaxation does not help. The loss therefore cannot be attributed solely to the freedom in Step 3. We also record what structure survives: a reduction confining every overload to two shapes, a seven-machine instance defeating the natural two-phase repair, and a strict 7/4 bound on the generalized three-path family.

Authors: Adam Y. Shavit

The 11/6 worst-case constant of the Wang-Sitters rounding scheme, which a companion note establishes, can naturally be attributed to the freedom in Step 3, where an arbitrary valid slot matching is permitted. We show that eliminating that freedom does not improve the constant. A minimum-makespan completion oracle still has worst-case constant exactly 11/6 against the optimum; both natural 7/4 statements about it are false; and restricting Step 1 to vertices of the relaxation does not help. The loss therefore cannot be attributed solely to the freedom in Step 3. We also record what structure survives: a reduction confining every overload to two shapes, a seven-machine instance defeating the natural two-phase repair, and a strict 7/4 bound on the generalized three-path family.

A Sharp Barrier for Consistent Submodular Maximization: Any Improvement over $2-\sqrt{2}$ Entails Exponential Queries or Linear Recourse

from arXiv: Data Structures and Algorithms

Authors: Shi Fu, Qixin Zhang, Dacheng Tao

Consistent submodular maximization studies the tradeoff between solution quality and stability when elements arrive over time. For a monotone submodular objective, which models diminishing returns, an algorithm maintains a set of at most $k$ available elements and changes only $O(1)$ elements after each insertion. Dütting et al. [2025] established a tight $2/3$ approximation with unrestricted computation and a polynomial-time $0.51$ approximation. They left open at STOC 2025 whether efficient algorithms can match the offline $1-1/e$ guarantee. We resolve this problem by proving that the supremum approximation achievable with polynomially many value queries and worst-case constant recourse is \[ β=2-\sqrt2\approx0.5858<1-1/e. \] For every $\varepsilon>0$, our randomized algorithm attains $β-\varepsilon$ with $O(\varepsilon^{-2})$ changes per insertion. Any fixed improvement requires exponentially many queries before one critical insertion or linear recourse of $Ω(k)$ changes at that insertion, even with unlimited queries afterwards. This gap quantifies the cost of consistency: the current oracle hides which elements will be needed after an arrival. We also determine the exact curvature-dependent threshold $1-(\sqrt2-1)\vartheta$, attain $1-1/e-\varepsilon$ for weighted coverage with $O(\varepsilon^{-1})$ recourse, and separate the existence of universal future-price certificates from their efficient computation. Our algorithm has a bounded-bit polynomial-time implementation for polynomial-bit rational oracle answers; the lower bound uses only logarithmic-bit rational answers.

Authors: Shi Fu, Qixin Zhang, Dacheng Tao

Consistent submodular maximization studies the tradeoff between solution quality and stability when elements arrive over time. For a monotone submodular objective, which models diminishing returns, an algorithm maintains a set of at most $k$ available elements and changes only $O(1)$ elements after each insertion. Dütting et al. [2025] established a tight $2/3$ approximation with unrestricted computation and a polynomial-time $0.51$ approximation. They left open at STOC 2025 whether efficient algorithms can match the offline $1-1/e$ guarantee. We resolve this problem by proving that the supremum approximation achievable with polynomially many value queries and worst-case constant recourse is \[ β=2-\sqrt2\approx0.5858<1-1/e. \] For every $\varepsilon>0$, our randomized algorithm attains $β-\varepsilon$ with $O(\varepsilon^{-2})$ changes per insertion. Any fixed improvement requires exponentially many queries before one critical insertion or linear recourse of $Ω(k)$ changes at that insertion, even with unlimited queries afterwards. This gap quantifies the cost of consistency: the current oracle hides which elements will be needed after an arrival. We also determine the exact curvature-dependent threshold $1-(\sqrt2-1)\vartheta$, attain $1-1/e-\varepsilon$ for weighted coverage with $O(\varepsilon^{-1})$ recourse, and separate the existence of universal future-price certificates from their efficient computation. Our algorithm has a bounded-bit polynomial-time implementation for polynomial-bit rational oracle answers; the lower bound uses only logarithmic-bit rational answers.

Online Inverse Integer Linear Optimization via Small-Gradient Skipping: Constant Regret and Finite Mistakes

from arXiv: Data Structures and Algorithms

Authors: Akira Kitaoka

In online inverse linear optimization, the learner predicts a weight at each round, observes the optimal action of the agent, and updates its prediction. In the general setting, the gap of $\log T$ between the regret upper bound $O(d \log T)$ and the lower bound $Ω(d)$ is unresolved (here $T$ is the total number of rounds and $d$ is the dimension). When the action set is M-convex, the regret is known to be bounded by $O(d \log d)$, but the method attaining it computes a center of gravity at every round. This paper therefore proposes Small-Gradient Skipping (SGS), a mechanism that skips the update at rounds without a mistake in the case where the correct action is uniformly separated from the other candidates, and applies it to online gradient descent, the online Newton step, and MetaGrad. The number of mistakes is then bounded, for all three, by a quantity independent of $T$; and for the online Newton step and for MetaGrad with SGS, the dimension dependence of the regret becomes $O(d^2)$ when the forward problem is an integer linear program, that is, the factor $\log T$ is removed. Moreover, when the action set is M-convex, the regret is bounded efficiently without computing a center of gravity.

Authors: Akira Kitaoka

In online inverse linear optimization, the learner predicts a weight at each round, observes the optimal action of the agent, and updates its prediction. In the general setting, the gap of $\log T$ between the regret upper bound $O(d \log T)$ and the lower bound $Ω(d)$ is unresolved (here $T$ is the total number of rounds and $d$ is the dimension). When the action set is M-convex, the regret is known to be bounded by $O(d \log d)$, but the method attaining it computes a center of gravity at every round. This paper therefore proposes Small-Gradient Skipping (SGS), a mechanism that skips the update at rounds without a mistake in the case where the correct action is uniformly separated from the other candidates, and applies it to online gradient descent, the online Newton step, and MetaGrad. The number of mistakes is then bounded, for all three, by a quantity independent of $T$; and for the online Newton step and for MetaGrad with SGS, the dimension dependence of the regret becomes $O(d^2)$ when the forward problem is an integer linear program, that is, the factor $\log T$ is removed. Moreover, when the action set is M-convex, the regret is bounded efficiently without computing a center of gravity.

Fast Algorithms for Sparse PCA and Robust Sparse Estimation

from arXiv: Data Structures and Algorithms

Authors: Giannis Iakovidis, Ankit Pensia

We study fast algorithms for sparse-PCA certification. Given a positive semidefinite matrix $M$, the problem asks either to rule out a large $k$-sparse quadratic form or to return a high-value (relaxed) witness. The standard semidefinite relaxation provides such certificates, but existing general-purpose solvers require $Ω(d^4)$ time. We give a bicriteria algorithm running in $O(d^2+d k^{O(\log k)})$ time: if some $k$-sparse unit vector has quadratic form greater than $2$, it returns either an $O(k^2)$-sparse unit vector or an SDP-feasible matrix of value at least $1$. For $k\leq\exp(O(\sqrt{\log d}))$, this running time is $O(d^2)$. We also go below the quadratic barrier in the sample-access model: Given $n=d^{o(1)}$ samples, our algorithm obtains a related one-sided certificate in $d^{2 - Ω(1)}$ time for $k=\mathrm{polylog}(d)$, without forming the empirical covariance matrix. As an application, these certificate routines yield the first quadratic and subquadratic-time algorithms for robust sparse estimation for broad families of distributions. Our sparse-PCA algorithm reduces a high-value sparse direction to a bounded-radius set in the graph of large correlations and searches the resulting candidate supports. The subquadratic implementation constructs this graph using fast correlation detection.

Authors: Giannis Iakovidis, Ankit Pensia

We study fast algorithms for sparse-PCA certification. Given a positive semidefinite matrix $M$, the problem asks either to rule out a large $k$-sparse quadratic form or to return a high-value (relaxed) witness. The standard semidefinite relaxation provides such certificates, but existing general-purpose solvers require $Ω(d^4)$ time. We give a bicriteria algorithm running in $O(d^2+d k^{O(\log k)})$ time: if some $k$-sparse unit vector has quadratic form greater than $2$, it returns either an $O(k^2)$-sparse unit vector or an SDP-feasible matrix of value at least $1$. For $k\leq\exp(O(\sqrt{\log d}))$, this running time is $O(d^2)$. We also go below the quadratic barrier in the sample-access model: Given $n=d^{o(1)}$ samples, our algorithm obtains a related one-sided certificate in $d^{2 - Ω(1)}$ time for $k=\mathrm{polylog}(d)$, without forming the empirical covariance matrix. As an application, these certificate routines yield the first quadratic and subquadratic-time algorithms for robust sparse estimation for broad families of distributions. Our sparse-PCA algorithm reduces a high-value sparse direction to a bounded-radius set in the graph of large correlations and searches the resulting candidate supports. The subquadratic implementation constructs this graph using fast correlation detection.

Scalable Composition of Byzantine Agreements under Reorder Attacks

from arXiv: Data Structures and Algorithms

Authors: Jing Chen, Jin Dong, Jichen Li, Xuanzhi Xia, Wentao Zhou

Byzantine agreement (BA) is a foundational building block in distributed systems, and the security analysis of BA protocols under multi-instance executions has attracted increasing attention. However, most existing adversary models focus solely on party corruption and neglect important threats posed by adversarial manipulations of communication channels in the network. Through channel attacks, messages can be reordered across multiple executions and lead to violations of the protocol's security guarantees, In this work, we present the first adversary model that combines party corruption and channel attacks. Based on this model, we establish new security thresholds for Byzantine agreement under parallel and concurrent compositions, supported by complementary impossibility and possibility results that match each other to form a tight bound. For the impossibility result, we show that even authenticated Byzantine agreement protocols cannot be secure under parallel composition when $n \leq 3t$ or $n \leq 2c + 2t + 1$, where $t$ and $c$ denote the number of corrupted parties and communication channels, respectively, and $n$ is the number of parties. For the possibility result, we prove the existence of secure protocols for unauthenticated Byzantine agreement under parallel and concurrent composition, when $n > \max\{3t, 2c+2t+1\}$. We first provide general black-box compilers that transform any single-instance secure BA protocol into one that is secure under parallel and concurrent executions without additional security assumptions. To optimize performance, we further design refined compilers using erasure-correcting codes. These refined versions significantly reduce communication overhead, particularly for long messages, where they achieve a constant multiplicative overhead compared with the original protocol, thus achieving the same asymptotic communication complexity.

Authors: Jing Chen, Jin Dong, Jichen Li, Xuanzhi Xia, Wentao Zhou

Byzantine agreement (BA) is a foundational building block in distributed systems, and the security analysis of BA protocols under multi-instance executions has attracted increasing attention. However, most existing adversary models focus solely on party corruption and neglect important threats posed by adversarial manipulations of communication channels in the network. Through channel attacks, messages can be reordered across multiple executions and lead to violations of the protocol's security guarantees, In this work, we present the first adversary model that combines party corruption and channel attacks. Based on this model, we establish new security thresholds for Byzantine agreement under parallel and concurrent compositions, supported by complementary impossibility and possibility results that match each other to form a tight bound. For the impossibility result, we show that even authenticated Byzantine agreement protocols cannot be secure under parallel composition when $n \leq 3t$ or $n \leq 2c + 2t + 1$, where $t$ and $c$ denote the number of corrupted parties and communication channels, respectively, and $n$ is the number of parties. For the possibility result, we prove the existence of secure protocols for unauthenticated Byzantine agreement under parallel and concurrent composition, when $n > \max\{3t, 2c+2t+1\}$. We first provide general black-box compilers that transform any single-instance secure BA protocol into one that is secure under parallel and concurrent executions without additional security assumptions. To optimize performance, we further design refined compilers using erasure-correcting codes. These refined versions significantly reduce communication overhead, particularly for long messages, where they achieve a constant multiplicative overhead compared with the original protocol, thus achieving the same asymptotic communication complexity.

Streaming Algorithms for Gaussian Kernel Density Statistics

from arXiv: Data Structures and Algorithms

Authors: Qin Zhang

Motivated by data produced by generative systems, \cite{LZ26b} formulates similarity-aware statistics via a weighted similarity graph, replacing equality with similarity in classical frequency-based statistics. Although this framework captures semantic relationships between nonidentical items, under general similarity functions even coarse one-pass approximation can require linear space. We therefore ask whether the geometric structure present in natural vector similarities can overcome this barrier. We answer this question affirmatively for the Gaussian kernel. For fixed-dimensional Euclidean vector streams, we study similarity-aware analogues of classical frequency statistics, including the number of distinct elements and frequency moments, through the diversity index and Gaussian density moments. We give one-pass sublinear-space approximation algorithms that exploit the geometric and analytic properties of the Gaussian kernel, and complement them with lower bounds. Our results show that geometric structure can fundamentally change the streaming complexity of similarity-aware statistical analysis.

Authors: Qin Zhang

Motivated by data produced by generative systems, \cite{LZ26b} formulates similarity-aware statistics via a weighted similarity graph, replacing equality with similarity in classical frequency-based statistics. Although this framework captures semantic relationships between nonidentical items, under general similarity functions even coarse one-pass approximation can require linear space. We therefore ask whether the geometric structure present in natural vector similarities can overcome this barrier. We answer this question affirmatively for the Gaussian kernel. For fixed-dimensional Euclidean vector streams, we study similarity-aware analogues of classical frequency statistics, including the number of distinct elements and frequency moments, through the diversity index and Gaussian density moments. We give one-pass sublinear-space approximation algorithms that exploit the geometric and analytic properties of the Gaussian kernel, and complement them with lower bounds. Our results show that geometric structure can fundamentally change the streaming complexity of similarity-aware statistical analysis.

Oracle Complexity of Stochastic Fixed-Point Equations with Nonexpansive Maps

from arXiv: Data Structures and Algorithms

Authors: Jelena Diakonikolas, Cristóbal Guzmán, David Martínez-Rubio

We study the oracle complexity of computing a point with small fixed-point residual $\|T(x)-x\| \leq ε$, for a general norm $\|\cdot\|$ and a self-map $T$ of a compact convex set. We study this problem in the setting where $T$ is nonexpansive with respect to the same norm $\|\cdot\|$ and accessed via an unbiased stochastic oracle with bounded variance $σ^2$. We provide an algorithm that solves such instances for any norm with a weak Rademacher type $q > 1$, with high probability. The algorithm is based on a recursive anchoring technique. For type-$2$ spaces, such as $\ell_p$-spaces for $p \in [2, \infty]$, our algorithm attains stochastic oracle complexity $\tilde O(σ^2 ε^{-3} + ε^{-1})$. We further prove a near-matching lower bound (i.e., matching up to poly-log factors) for such $\ell_{\infty}$-norm instances in high dimensions. Our lower bound holds against any randomized algorithm that succeeds with constant probability. It further extends to settings with ``sparse'' noise, where variance measured with respect to any $\ell_p$ norm is of the same order, ruling out the possibility of improving oracle complexity as a function of $\varepsilon$ by measuring variance in a non-matching $\ell_p$ norm.

Authors: Jelena Diakonikolas, Cristóbal Guzmán, David Martínez-Rubio

We study the oracle complexity of computing a point with small fixed-point residual $\|T(x)-x\| \leq ε$, for a general norm $\|\cdot\|$ and a self-map $T$ of a compact convex set. We study this problem in the setting where $T$ is nonexpansive with respect to the same norm $\|\cdot\|$ and accessed via an unbiased stochastic oracle with bounded variance $σ^2$. We provide an algorithm that solves such instances for any norm with a weak Rademacher type $q > 1$, with high probability. The algorithm is based on a recursive anchoring technique. For type-$2$ spaces, such as $\ell_p$-spaces for $p \in [2, \infty]$, our algorithm attains stochastic oracle complexity $\tilde O(σ^2 ε^{-3} + ε^{-1})$. We further prove a near-matching lower bound (i.e., matching up to poly-log factors) for such $\ell_{\infty}$-norm instances in high dimensions. Our lower bound holds against any randomized algorithm that succeeds with constant probability. It further extends to settings with ``sparse'' noise, where variance measured with respect to any $\ell_p$ norm is of the same order, ruling out the possibility of improving oracle complexity as a function of $\varepsilon$ by measuring variance in a non-matching $\ell_p$ norm.

Approximate Nearest Neighbor in Ultra-High Dimensional $\ell_\infty$

from arXiv: Data Structures and Algorithms

Authors: Nathan White, Tian Zhang

We study the approximate nearest neighbor problem under $\ell_\infty$ in the ultra-high dimensional setting where the dimension $d$ is significantly larger than the number of points $n$. Thus, we desire data structures with no dependence on $d$ in the query time. [Herold-Nanongkai-Spoerhase-Varma-Wu, SoCG 2025] introduce this problem and give data structures in $\ell_p$: for $p=1,2$, they give $(1+\varepsilon)$-approximation data structures with space $\tilde{O}(n\log d/\text{poly}(\varepsilon))$ and query time $\tilde{O}(n/\text{poly}(\varepsilon))$. Since any data structure must have query time $Ω(\min \{n,d\})$, this query time is nearly tight. However, their results are inefficient for $\ell_\infty$, with query time $Ω(nd)$. In order to handle the challenges of $\ell_\infty$, we introduce a notion of subset embeddings, which embed points by simply selecting a subset of dimensions. In particular, we show one may preserve all pairwise distances of an $n$ point dataset up to a factor of $O(c)$ by computing distances on only $n^{1+1/c}$ coordinates. We also show a matching lower bound: for any $c > 1$, there exists a set of $n$ points in $\mathbb{R}^{d}$ such that any subset embedding for the set with approximation $c$ must have at least $n^{1+Ω(1/c)}$ coordinates. Using our subset embeddings, we give data structures for approximate nearest neighbor in $\ell_\infty$ with space $O(n^2\log d)$, query time $\tilde{O}(n^{1+1/c})$, and approximation $O(c\log\log n)$ for any $c \geq 1$. Finally, we give another data structure for the approximate nearest neighbor under $\ell_\infty$ with the same space and query time as our subset embedding approach, but with approximation $O(c^{\log_2 3}) \approx O(c^{1.58})$. This allows us to achieve $O(1)$-approximation with query time e.g.~$n^{1.01}$

Authors: Nathan White, Tian Zhang

We study the approximate nearest neighbor problem under $\ell_\infty$ in the ultra-high dimensional setting where the dimension $d$ is significantly larger than the number of points $n$. Thus, we desire data structures with no dependence on $d$ in the query time. [Herold-Nanongkai-Spoerhase-Varma-Wu, SoCG 2025] introduce this problem and give data structures in $\ell_p$: for $p=1,2$, they give $(1+\varepsilon)$-approximation data structures with space $\tilde{O}(n\log d/\text{poly}(\varepsilon))$ and query time $\tilde{O}(n/\text{poly}(\varepsilon))$. Since any data structure must have query time $Ω(\min \{n,d\})$, this query time is nearly tight. However, their results are inefficient for $\ell_\infty$, with query time $Ω(nd)$. In order to handle the challenges of $\ell_\infty$, we introduce a notion of subset embeddings, which embed points by simply selecting a subset of dimensions. In particular, we show one may preserve all pairwise distances of an $n$ point dataset up to a factor of $O(c)$ by computing distances on only $n^{1+1/c}$ coordinates. We also show a matching lower bound: for any $c > 1$, there exists a set of $n$ points in $\mathbb{R}^{d}$ such that any subset embedding for the set with approximation $c$ must have at least $n^{1+Ω(1/c)}$ coordinates. Using our subset embeddings, we give data structures for approximate nearest neighbor in $\ell_\infty$ with space $O(n^2\log d)$, query time $\tilde{O}(n^{1+1/c})$, and approximation $O(c\log\log n)$ for any $c \geq 1$. Finally, we give another data structure for the approximate nearest neighbor under $\ell_\infty$ with the same space and query time as our subset embedding approach, but with approximation $O(c^{\log_2 3}) \approx O(c^{1.58})$. This allows us to achieve $O(1)$-approximation with query time e.g.~$n^{1.01}$

Subexponential Approximation of the Permanent in Deterministic Polynomial Time

from arXiv: Data Structures and Algorithms

Authors: Sergei Kudria, Jason Luo, Mahbod Majid

We give the first deterministic polynomial time algorithm that approximates the permanent of arbitrary nonnegative rational matrices within a subexponential factor. For a matrix of order $n$, the approximation factor is \[ \exp\!\left(O\!\left(\frac{n(\log\log n)^2}{\log n}\right)\right)=\exp(o(n)). \] All previously known deterministic polynomial time guarantees for unrestricted inputs had approximation factors $\exp(Ω(n))$. Our proof uses convex optimization to tighten an upper bound on the permanent. The bound is based on weighted sums over all matchings in a bipartite graph representing the matrix, and correlations between unmatched vertices control its error. We approximate these sums deterministically using correlation decay and a bound on the effect of vertex deletion.

Authors: Sergei Kudria, Jason Luo, Mahbod Majid

We give the first deterministic polynomial time algorithm that approximates the permanent of arbitrary nonnegative rational matrices within a subexponential factor. For a matrix of order $n$, the approximation factor is \[ \exp\!\left(O\!\left(\frac{n(\log\log n)^2}{\log n}\right)\right)=\exp(o(n)). \] All previously known deterministic polynomial time guarantees for unrestricted inputs had approximation factors $\exp(Ω(n))$. Our proof uses convex optimization to tighten an upper bound on the permanent. The bound is based on weighted sums over all matchings in a bipartite graph representing the matrix, and correlations between unmatched vertices control its error. We approximate these sums deterministically using correlation decay and a bound on the effect of vertex deletion.

Optimal Low-Rank Quantum State Tomography with Bounded-Sample Joint Measurements

from arXiv: Data Structures and Algorithms

Authors: Ashwin Nayak, Xingyu Zhou

We determine the optimal sample complexity of low-rank quantum state tomography when each measurement may act jointly on at most $t$ samples. For sufficiently small $\varepsilon$, estimating an unknown state on $\mathbb{C}^d$ of rank at most $r$ to trace norm error $\varepsilon$ with constant success probability requires, and is achievable with, $$ Θ\left( \frac{dr}{\varepsilon^2} \max\left\{1,\frac r{\sqrt t}\right\} \right)$$ samples. The lower bound allows the protocol to choose each joint measurement adaptively using all previous classical outcomes; the matching upper bound is nonadaptive. Thus joint measurements on at most $t$ samples improve the complexity of algorithms making single-sample measurements by at most a factor $\sqrt t$. Further, measuring order $r^2$ samples jointly is necessary and sufficient to attain the unrestricted collective rate. For the lower bound, we vary the support of a state with fixed uniform spectrum and bound the Fisher information trace of every joint measurement on $t$ samples. The adaptive Fisher chain rule and the van Trees inequality then give the trace norm lower bound. For the upper bound, we construct and analyze a nonadaptive tomography protocol based on a Gaussian joint measurement. An explicit second moment identity and a conditional Gaussian law outside the state's support give a rank-dependent error analysis, yielding the matching rate.

Authors: Ashwin Nayak, Xingyu Zhou

We determine the optimal sample complexity of low-rank quantum state tomography when each measurement may act jointly on at most $t$ samples. For sufficiently small $\varepsilon$, estimating an unknown state on $\mathbb{C}^d$ of rank at most $r$ to trace norm error $\varepsilon$ with constant success probability requires, and is achievable with, $$ Θ\left( \frac{dr}{\varepsilon^2} \max\left\{1,\frac r{\sqrt t}\right\} \right)$$ samples. The lower bound allows the protocol to choose each joint measurement adaptively using all previous classical outcomes; the matching upper bound is nonadaptive. Thus joint measurements on at most $t$ samples improve the complexity of algorithms making single-sample measurements by at most a factor $\sqrt t$. Further, measuring order $r^2$ samples jointly is necessary and sufficient to attain the unrestricted collective rate. For the lower bound, we vary the support of a state with fixed uniform spectrum and bound the Fisher information trace of every joint measurement on $t$ samples. The adaptive Fisher chain rule and the van Trees inequality then give the trace norm lower bound. For the upper bound, we construct and analyze a nonadaptive tomography protocol based on a Gaussian joint measurement. An explicit second moment identity and a conditional Gaussian law outside the state's support give a rank-dependent error analysis, yielding the matching rate.

Testing the Binary Rank with Polynomial Query Complexity

from arXiv: Data Structures and Algorithms

Authors: Michal Parnas

We provide an adaptive two-sided error testing algorithm for the binary rank of a $0,1$ matrix $M$ with query complexity $O(d^3\log(d+1)/ε^2)$, where $d$ is the tested binary rank bound and $ε$ is the distance parameter. This answers an open question posed by Parnas, Ron and Shraibman~\cite{parnas2021property}, who asked whether the binary rank can be tested with query complexity polynomial in $d$ and $1/ε$. Furthermore, our testing algorithm can be used to find an approximate binary decomposition of $M$ with an additional $d(n+m)$ queries. That is, under the promise that the binary rank of $M$ is at most $d$, we show how to find, with probability at least $5/6$, two $0,1$ matrices $A',B'$ such that $M' = A' \cdot B'$ is a $0,1$ matrix which differs from $M$ on at most an $O(ε)$ fraction of its entries.

Authors: Michal Parnas

We provide an adaptive two-sided error testing algorithm for the binary rank of a $0,1$ matrix $M$ with query complexity $O(d^3\log(d+1)/ε^2)$, where $d$ is the tested binary rank bound and $ε$ is the distance parameter. This answers an open question posed by Parnas, Ron and Shraibman~\cite{parnas2021property}, who asked whether the binary rank can be tested with query complexity polynomial in $d$ and $1/ε$. Furthermore, our testing algorithm can be used to find an approximate binary decomposition of $M$ with an additional $d(n+m)$ queries. That is, under the promise that the binary rank of $M$ is at most $d$, we show how to find, with probability at least $5/6$, two $0,1$ matrices $A',B'$ such that $M' = A' \cdot B'$ is a $0,1$ matrix which differs from $M$ on at most an $O(ε)$ fraction of its entries.

A Deadline-Driven Algorithm for Polyamorous Scheduling

from arXiv: Data Structures and Algorithms

Authors: Arjun Maneesh Agarwal

In Polyamorous Scheduling Problem, we are given an edge-weighted graph and must find a periodic schedule of matchings in this graph which minimizes the maximal weighted waiting time between consecutive occurrences of the same edge. This NP-hard problem generalises Bamboo Garden Trimming and is motivated by the need to find schedules of pairwise meetings in a complex social group. We present a $4 G^*$ algorithm for Polyamorous Scheduling, improving the previously known bound of $3 + \sqrt{5} \approx 5.236$. Our algorithm is inspired by the Deadline-Driven Heuristic which is optimal for Bamboo Garden Trimming (BGT).

Authors: Arjun Maneesh Agarwal

In Polyamorous Scheduling Problem, we are given an edge-weighted graph and must find a periodic schedule of matchings in this graph which minimizes the maximal weighted waiting time between consecutive occurrences of the same edge. This NP-hard problem generalises Bamboo Garden Trimming and is motivated by the need to find schedules of pairwise meetings in a complex social group. We present a $4 G^*$ algorithm for Polyamorous Scheduling, improving the previously known bound of $3 + \sqrt{5} \approx 5.236$. Our algorithm is inspired by the Deadline-Driven Heuristic which is optimal for Bamboo Garden Trimming (BGT).

Wednesday, September 09

On AI

from Emanuele Viola

Whoa! Overnight (i.e., during the course of a short summer) AI in theoretical research went from a toy to an indispensable tool, immeasurably speeding up research, including solving problems on its own. Now all the talk is the Math AI crisis, something which would have been unthinkable even last May. The people in the trades […]

Whoa! Overnight (i.e., during the course of a short summer) AI in theoretical research went from a toy to an indispensable tool, immeasurably speeding up research, including solving problems on its own. Now all the talk is the Math AI crisis, something which would have been unthinkable even last May. The people in the trades would be laughing their butts off, if they didn’t know better than following up what’s going on (they’re probably fishing, their Ford 250 truck with their company’s cute logo parked at the edge of the lake). What we thought quintessential human, creating art, math, thinking, it’s all done by machines which we fed with so many papers, movies, songs, conversations, that they got better than us. In the meantime, robotics is lagging behind and crimping a cable, scraping a siding, cutting pvc pipes still require the wear and tear of our bodies. Not even science fiction dystopia.

I consider what’s happening in AI the most exciting technology since the invention of computers. I don’t make this statement lightly. My two previous picks would be electricity and the telephone. I didn’t expect to see this in my lifetime, or even to ever happen. And I don’t know anyone who wasn’t shocked, except Ray Kurzweil: I met him last year and he told me: wait two months.

AI is disrupting academia. This is scary, but also the system wasn’t so good that a good shake isn’t necessarily for the best. Driven by cut-throat competition and the limitations of the human brain (as well as the general instability of the geopolitical landscape) academia became publish-or-perish, the unbearable pressure to push incremental papers, or to show off mathematical weight-lifting in the form of super-technical papers. Can we finally stop celebrating weight lifting? The goal has never been making things look complicated, but this was a target, and even specifically given as advice to young researchers (sad). There is now no point in this, given that AI can easily fill pages with integrals, and also to some extent simplify (though this is less clear, as it doesn’t seem AI has a good sense of what’s easy for us, understandably given it’s a machine and the awful training it was given). Your goal is to make things easy! But naturally things can also become worse. Before, you could write a paper from a remote region of the world and instantly get recognition. Today, this is to be evaluated under the lens of AI, which might actually end up making public relations even more critical for success in academia. Still, it may be that all “low-hanging AI fruits” are taken soon, and human contributions become more transparent. I tend to believe this will happen, basically for the reason that the math problems were outlined before AI and so it is natural that many fall under the new tool, while at the same time they cannot be produced by humans at a high rate. Hopefully we can also move away from incremental research and use AI to do something big, like progress in computational complexity theory, an area which still emerges relatively unscathed (in terms of big breakthroughs). Use all means at your disposal to solve P vs NP!

It’s palpable in the air the quest for guidance, principles. Many conferences have been set up in place before the revolution, and are now struggling to evaluate the onslaught of single-author AI-slob which would have appeared solid last May. This is no small problem, especially given the culture of conferences in computer science. You obviously don’t want to fill a conference with people who have no idea what they are talking about. Might be hi-time to rethink conferences… And there’s the fear of missing the next generation of scientists.

The excitement at this awesome new power at my disposal comes with a bitter, nostalgic feeling. I had arranged much of my life around working math in my head, since my teen age years when I would enjoy sitting in the sun, reading my calculus homework, then closing my eyes and solving it in the head. (I once got zero at an exam because I didn’t show the steps, my prof was however flexible enough to challenge me to repeat the feat in front of him, and then gave me full score.) This is not at all to say I’m particularly good at this sort of thing, but I enjoyed the feeling. In the last 30 years or so, I worked countless problems in my head, during walks, swims, bike rides. The problems never left me, at the doctor’s office, during parties, when everybody else was bored, when I was waiting in line, while traveling. At night I had to fight them off with utmost concentration so that I could sleep (the one big, big drawback). I was sometimes exhausted, sleepless, even nauseated with math, but I never felt bored. But there’s more, even when I wasn’t actively thinking, maybe resting or playing a videogame, I always had the intense background feeling that I was only doing that so that my brain would cool off and later be in a better position to produce math, a main metric I’ve been measuring my life by.

Now this seems all gone. Thinking without an AI companion appears pointless. I’ll be sitting in the woods and pull out my phone, or just talk, and understand one more line of the proof that AI generated. What will be missing forever is the feeling that all this had to happen entirely in my head, that there was no comparable tool, nothing else that could match what my own concentration could produce. This is what justified the endless videogaming, hikes, gazing through the window for hours, meditation, the endless fine tuning of my sleep, walks, food, so that at some point, even just for a brief but explosive moment, I could unleash my thought.

By Manu

Postdoc fellowship at Simons Institute for the Theory of Computing (apply by December 1, 2026)

from CCI: jobs

The Simons Institute for the Theory of Computing invites applications for research fellowships and Long-Term Visitors for Spring 2027. Fellowships are for exceptional young scientists, and long-term visitors are active researchers beyond their postdoctoral years. The Institute will host a program on “Symmetry in Efficient Computation with Local Constraints” in Spring 2027. Website: simons.berkeley.edu/research-fellowship-call-applications Email: […]

The Simons Institute for the Theory of Computing invites applications for research fellowships and Long-Term Visitors for Spring 2027. Fellowships are for exceptional young scientists, and long-term visitors are active researchers beyond their postdoctoral years. The Institute will host a program on “Symmetry in Efficient Computation with Local Constraints” in Spring 2027.

Website: https://simons.berkeley.edu/research-fellowship-call-applications
Email: simonsassociatedirector@berkeley.edu

By shacharlovett

Navier-Stokes and Lean

from Computational Complexity

I was working on this week's post on Lean after reading Kevin Hartnett's book The Proof in the Code: How a Truth Machine Is Transforming Math and AI. And then yesterday OpenAI announced a solution to Navier-Stokes, one of the Millennium problems. An incredible accomplishment to say the least. Hours earlier Tristan Buckmaster posted about his progress with Levent Alpöge based on a program started by Diego Córdoba and Luis Martínez-Zoroa, and his interactions with OpenAI. I'm still trying to understand what happened and will write more later but I recommend the Quanta article to get you up to speed.

Lean plays a major role for both projects. OpenAI fully formulated their results in Lean. Buckmaster said they have Lean-verified proofs for the three results they made public but held back on the "blowup for hypo-dissipative Navier Stokes" because the Lean verification has not finished. So it's worth taking a look back.

Leonardo de Moura developed the first version of Lean in 2013 as a Microsoft project for proving code correct. Hartnett tells the story of the people involved in the development of the various versions of Lean capturing the excitements and disagreements. What caught me was the lack of backward compatibility, the definitions and theorems formalized in one version of Lean might break in the next. There was a constant need to get the libraries back up to date until the development stabilized.

The best part of the book focused on some big projects in Lean.

  • Tom Hales wanting to convince the world that his proof of the Kepler Conjecture was correct.
  • Peter Scholze wanting to convince himself of the correctness of his liquid tensor experiment.
  • Kevin Buzzard, Johan Commelin and Patrick Massot formalizing Scholze's Perfectoid Spaces to show Lean can handle modern mathematical objects.
  • Terence Tao wanting to formalize in Lean the proof with Tim Gowers, Ben Green and Freddie Manners of the polynomial Freiman-Ruzsa conjecture, just to show it could be done.
None of these were individual efforts but required teams of volunteers to fill in details. Tao approached his proof as a polymath project and his superstar status in the math community really helped get volunteers and popularized Lean.
I did learn a new term from the book, the "de Bruijn factor", the ratio of the length of the formal computer proof to the length of the human proof. I suspect the factor is very large for proofs in theoretical computer science, particularly computational complexity which is why we haven't seen many computer science theorems formalized in Lean.
Hartnett's story ends at the January 2025 Joint Math Meetings in Seattle, a conference I attended. He talks about the initial connections between AI and Lean, but not the uncertain AI future that started to worry mathematicians. By the time the book was published in June 2026, we had seen tremendous progress in AI proving theorems. We've seen even more progress in the three months since, even in the last three days given the news above.
Lean plays a different role now. No longer do you need to write Lean code, any more than you need to write Python code, you can just use AI to generate it. Anthropic fully formalized Fermat's Last Theorem in Lean just last week and hardly caused a stir. 

And now Lean, particularly in the Navier-Stokes papers, is being used as a time-stamp, a way to claim your theorem before having to write it up properly in an explainable way. Buckmaster even held back a result because it wasn't yet Lean verified. The way we even publish results is a-changing.

By Lance Fortnow

I was working on this week's post on Lean after reading Kevin Hartnett's book The Proof in the Code: How a Truth Machine Is Transforming Math and AI. And then yesterday OpenAI announced a solution to Navier-Stokes, one of the Millennium problems. An incredible accomplishment to say the least. Hours earlier Tristan Buckmaster posted about his progress with Levent Alpöge based on a program started by Diego Córdoba and Luis Martínez-Zoroa, and his interactions with OpenAI. I'm still trying to understand what happened and will write more later but I recommend the Quanta article to get you up to speed.

Lean plays a major role for both projects. OpenAI fully formulated their results in Lean. Buckmaster said they have Lean-verified proofs for the three results they made public but held back on the "blowup for hypo-dissipative Navier Stokes" because the Lean verification has not finished. So it's worth taking a look back.

Leonardo de Moura developed the first version of Lean in 2013 as a Microsoft project for proving code correct. Hartnett tells the story of the people involved in the development of the various versions of Lean capturing the excitements and disagreements. What caught me was the lack of backward compatibility, the definitions and theorems formalized in one version of Lean might break in the next. There was a constant need to get the libraries back up to date until the development stabilized.

The best part of the book focused on some big projects in Lean.

  • Tom Hales wanting to convince the world that his proof of the Kepler Conjecture was correct.
  • Peter Scholze wanting to convince himself of the correctness of his liquid tensor experiment.
  • Kevin Buzzard, Johan Commelin and Patrick Massot formalizing Scholze's Perfectoid Spaces to show Lean can handle modern mathematical objects.
  • Terence Tao wanting to formalize in Lean the proof with Tim Gowers, Ben Green and Freddie Manners of the polynomial Freiman-Ruzsa conjecture, just to show it could be done.
None of these were individual efforts but required teams of volunteers to fill in details. Tao approached his proof as a polymath project and his superstar status in the math community really helped get volunteers and popularized Lean.

I did learn a new term from the book, the "de Bruijn factor", the ratio of the length of the formal computer proof to the length of the human proof. I suspect the factor is very large for proofs in theoretical computer science, particularly computational complexity which is why we haven't seen many computer science theorems formalized in Lean.

Hartnett's story ends at the January 2025 Joint Math Meetings in Seattle, a conference I attended. He talks about the initial connections between AI and Lean, but not the uncertain AI future that started to worry mathematicians. By the time the book was published in June 2026, we had seen tremendous progress in AI proving theorems. We've seen even more progress in the three months since, even in the last three days given the news above.

Lean plays a different role now. No longer do you need to write Lean code, any more than you need to write Python code, you can just use AI to generate it. Anthropic fully formalized Fermat's Last Theorem in Lean just last week and hardly caused a stir. 

And now Lean, particularly in the Navier-Stokes papers, is being used as a time-stamp, a way to claim your theorem before having to write it up properly in an explainable way. Buckmaster even held back a result because it wasn't yet Lean verified. The way we even publish results is a-changing.

By Lance Fortnow