Last Update

OPML feed of all feeds.

Subscribe to the Atom feed, RSS feed to stay up to date.

Thank you to arXiv for use of its open access interoperability.

Note: the date of arXiv entries announced right after publication holidays might incorrectly show up as the date of the publication holiday itself. This is due to our ad hoc method of inferring announcement dates, which are not returned by the arXiv API.

Powered by Pluto.

Source on GitHub.

Maintained by Nima Anari, Arnab Bhattacharyya, Gautam Kamath.

Theory of Computing Report

Friday, October 02

Beyond Light Cones: State Preparation Complexity in Quantum Spin Glasses

from arXiv: Computational Complexity

Authors: Omar Al-Ghattas, David Gamarnik, Bobak T Kiani

We introduce a method for studying state preparation complexity in dense quantum $p$-spin Hamiltonians on $n$ qubits, going beyond bounds based only on circuit lightcones. The key input is the class's effective profile complexity, which is derived from the metric entropy of its Pauli profiles. These profiles record expectations of all Pauli operators supported on exactly $p$ qubits. Classes with uniformly bounded quadratic effective profile complexity remain separated from the ground-state energy by a positive multiple of $\sqrt n$ for sufficiently large fixed $p$. At subquadratic effective profile complexity, the class cannot outperform a suitable benchmark class at leading order, with product states providing a universal benchmark. The proof combines an adaptation of a nonsymmetric quantum de Finetti theorem of Berta et al. (arXiv:1810.12197) with Gaussian process entropy bounds. Applying this framework, we show that attaining near-ground-state energy requires $Ω(n^2/\log n)$ one- and two-qubit gates, even with arbitrary discardable ancillas. We also obtain depth-width tradeoffs, entanglement-depth and matrix product state bond-dimension lower bounds, and obstructions for both orientations at every fixed level of Parham's magic hierarchy (arXiv:2504.19966), with total circuit width $O(n)$. In first-level reverse magic, a shallow circuit is followed by an unrestricted Clifford circuit. The latter can spread local observables across the system, preventing a direct application of small-lightcone bounds. For this first-level class, our bounds also allow arbitrarily many clean ancillas at fixed shallow-circuit depth. A sharper benchmark shows that Clifford+$T$ circuits with $o(n)$ $T$-gates have no leading-order energy advantage over product stabilizer states, even with unrestricted Clifford operations and arbitrary discardable ancillas.

Authors: Omar Al-Ghattas, David Gamarnik, Bobak T Kiani

We introduce a method for studying state preparation complexity in dense quantum $p$-spin Hamiltonians on $n$ qubits, going beyond bounds based only on circuit lightcones. The key input is the class's effective profile complexity, which is derived from the metric entropy of its Pauli profiles. These profiles record expectations of all Pauli operators supported on exactly $p$ qubits. Classes with uniformly bounded quadratic effective profile complexity remain separated from the ground-state energy by a positive multiple of $\sqrt n$ for sufficiently large fixed $p$. At subquadratic effective profile complexity, the class cannot outperform a suitable benchmark class at leading order, with product states providing a universal benchmark. The proof combines an adaptation of a nonsymmetric quantum de Finetti theorem of Berta et al. (arXiv:1810.12197) with Gaussian process entropy bounds. Applying this framework, we show that attaining near-ground-state energy requires $Ω(n^2/\log n)$ one- and two-qubit gates, even with arbitrary discardable ancillas. We also obtain depth-width tradeoffs, entanglement-depth and matrix product state bond-dimension lower bounds, and obstructions for both orientations at every fixed level of Parham's magic hierarchy (arXiv:2504.19966), with total circuit width $O(n)$. In first-level reverse magic, a shallow circuit is followed by an unrestricted Clifford circuit. The latter can spread local observables across the system, preventing a direct application of small-lightcone bounds. For this first-level class, our bounds also allow arbitrarily many clean ancillas at fixed shallow-circuit depth. A sharper benchmark shows that Clifford+$T$ circuits with $o(n)$ $T$-gates have no leading-order energy advantage over product stabilizer states, even with unrestricted Clifford operations and arbitrary discardable ancillas.

The Robustness of QAC0

from arXiv: Computational Complexity

Authors: Daniel Grier, Jackson Morris, Kewen Wu

In this work we study the robustness of $\mathsf{QAC}^0$ with respect to error tolerance and modifications to its gate-set. First, we investigate whether the non-zero error typically allowed for $\mathsf{QAC}^0$ circuits computing Boolean functions is truly necessary. We show that the error inherent in the parallel $W$-test of \cite{grier_morris_wu} can be eliminated entirely via a novel application of exact amplitude amplification in the many-copies context. Consequently, we find that $\mathsf{QAC}^0$ can \textit{exactly} simulate $\mathsf{TC}^0$ with polynomially many copies of the classical input and that for every fixed prime $p$ exact $\mathsf{QAC}^0$, $\mathsf{EQAC}^0$, can compute total Boolean functions outside of $\mathsf{AC}^0[p]$. Second, we ask to what extent the computational power of $\mathsf{QAC}^0$ follows from the fact that arbitrary single-qubit gates may be used at any point in the circuit. We find that $\mathsf{QAC}^0$ is in fact robust to restrictions on which single-qubit gates are permitted: every $\mathsf{QAC}^0$ circuit can be approximately implemented by a $\mathsf{QAC}^0$ circuit consisting of just generalized Toffoli, $S$, and Hadamard gates. Moreover, this approximating circuit can be constructed efficiently from a classical description of the original circuit.

Authors: Daniel Grier, Jackson Morris, Kewen Wu

In this work we study the robustness of $\mathsf{QAC}^0$ with respect to error tolerance and modifications to its gate-set. First, we investigate whether the non-zero error typically allowed for $\mathsf{QAC}^0$ circuits computing Boolean functions is truly necessary. We show that the error inherent in the parallel $W$-test of \cite{grier_morris_wu} can be eliminated entirely via a novel application of exact amplitude amplification in the many-copies context. Consequently, we find that $\mathsf{QAC}^0$ can \textit{exactly} simulate $\mathsf{TC}^0$ with polynomially many copies of the classical input and that for every fixed prime $p$ exact $\mathsf{QAC}^0$, $\mathsf{EQAC}^0$, can compute total Boolean functions outside of $\mathsf{AC}^0[p]$. Second, we ask to what extent the computational power of $\mathsf{QAC}^0$ follows from the fact that arbitrary single-qubit gates may be used at any point in the circuit. We find that $\mathsf{QAC}^0$ is in fact robust to restrictions on which single-qubit gates are permitted: every $\mathsf{QAC}^0$ circuit can be approximately implemented by a $\mathsf{QAC}^0$ circuit consisting of just generalized Toffoli, $S$, and Hadamard gates. Moreover, this approximating circuit can be constructed efficiently from a classical description of the original circuit.

Optimal transducers using symmetries

from arXiv: Computational Complexity

Authors: Benoît Dubus, Julien Ladeuze, Jérémie Roland

Transducers (Belovs, Jeffery and Yolcu, 2024) are a quantum computing framework describing a quantum algorithm as a unitary converting an input state into a target state using a catalyst, an auxiliary vector that is left unchanged. They are a powerful tool in quantum algorithm design, especially in the context of quantum query complexity: feasible points of the (dual) adversary semidefinite program directly translate into transducers and the optimal transduction complexity is equal to the adversary bound, i.e. the Las Vegas complexity, which is known to characterize bounded-error quantum query complexity. Moreover, contrary to bounded-error algorithms, transducers compose exactly, which limits overheads due to controlling errors in algorithms constructed by composition. Constructing efficient, let alone optimal, transducers in terms of quantum query complexity nevertheless remains a hard task since it still requires solving the adversary SDP and constructing the unitary to obtain an explicit algorithm. In this paper, we show how using the symmetry group of state-conversion problems simplifies both steps. First, using a symmetrization argument, we prove an optimal catalyst can always be chosen covariant under a representation of the symmetry group. Second, we prove that the transducer intertwines two different representations of the group and can thus be chosen block diagonal in the isotypic decomposition of the Hilbert space. Using those methods, we then derive optimal transducers, with optimal constants, for different widely used quantum algorithmic primitives, such as unstructured search, amplitude amplification and amplitude estimation. Our approach extends previous work on the use of representation theory to compute adversary lower bounds (Høyer, Lee, and {\v S}palek, 2007; Ambainis, Magnin, Roetteler and Roland, 2011) to the systematic construction of optimal algorithms.

Authors: Benoît Dubus, Julien Ladeuze, Jérémie Roland

Transducers (Belovs, Jeffery and Yolcu, 2024) are a quantum computing framework describing a quantum algorithm as a unitary converting an input state into a target state using a catalyst, an auxiliary vector that is left unchanged. They are a powerful tool in quantum algorithm design, especially in the context of quantum query complexity: feasible points of the (dual) adversary semidefinite program directly translate into transducers and the optimal transduction complexity is equal to the adversary bound, i.e. the Las Vegas complexity, which is known to characterize bounded-error quantum query complexity. Moreover, contrary to bounded-error algorithms, transducers compose exactly, which limits overheads due to controlling errors in algorithms constructed by composition. Constructing efficient, let alone optimal, transducers in terms of quantum query complexity nevertheless remains a hard task since it still requires solving the adversary SDP and constructing the unitary to obtain an explicit algorithm. In this paper, we show how using the symmetry group of state-conversion problems simplifies both steps. First, using a symmetrization argument, we prove an optimal catalyst can always be chosen covariant under a representation of the symmetry group. Second, we prove that the transducer intertwines two different representations of the group and can thus be chosen block diagonal in the isotypic decomposition of the Hilbert space. Using those methods, we then derive optimal transducers, with optimal constants, for different widely used quantum algorithmic primitives, such as unstructured search, amplitude amplification and amplitude estimation. Our approach extends previous work on the use of representation theory to compute adversary lower bounds (Høyer, Lee, and {\v S}palek, 2007; Ambainis, Magnin, Roetteler and Roland, 2011) to the systematic construction of optimal algorithms.

Time-space lower bounds for breaking quantum cryptography

from arXiv: Computational Complexity

Authors: Fangqi Dong, Alex Lombardi

We prove near-optimal time-space lower bounds for breaking quantum cryptography in the random oracle model. Specifically, we show that a $T$-query adversary with $S$ qubits of non-uniform advice can recover a random key $k$ from the $n$-qubit binary phase state $|ψ_k\rangle \propto \sum_{x} R(k,x) |x\rangle$ with probability at most $O(\frac{T^2 + \sqrt{ST}}{N})$ for $N=2^n$. In contrast, the best known bound for post-quantum one-way functions is $O(\frac{T^2 + ST}{N})$, with a trivial attack at $S = N$. This demonstrates a new advantage of quantum cryptography over classical cryptography: $n$ qubits of communication suffice for security against preprocessing attacks with space up to $N^2$ rather than $N$. Our methodology is simple: express the optimal preprocessing attack as the operator norm of a random matrix, and bound this value in expectation over the random oracle via the trace-moment method. These trace moments have a natural interpretation using compressed oracles [Zhandry, Crypto 2019], which we then analyze. This can be viewed as a simplification and generalization of the approach of Liu [Eurocrypt 2023] for proving time-space tradeoffs for breaking post-quantum cryptography. We also prove the following results: (1) We tighten Liu's analysis of post-quantum PRGs in QROM, achieving a distinguishing advantage bound of $O(\frac{T^2}N + \sqrt{\frac{ST}N})$. (2) For unitary synthesis, we extend the one-query lower bound of Lombardi-Ma-Wright [STOC 2024] to hold against adversaries that can make one arbitrary function query along with polynomially many (adaptive) queries to the random oracle, either before or after the function query. This also interprets the original LMW24 result in terms of compressed oracles. (3) Finally, we prove a tight $O(\frac{\sqrt{S}}N)$ bound for the pseudorandomness of random binary phase states against space $S$ distinguishers.

Authors: Fangqi Dong, Alex Lombardi

We prove near-optimal time-space lower bounds for breaking quantum cryptography in the random oracle model. Specifically, we show that a $T$-query adversary with $S$ qubits of non-uniform advice can recover a random key $k$ from the $n$-qubit binary phase state $|ψ_k\rangle \propto \sum_{x} R(k,x) |x\rangle$ with probability at most $O(\frac{T^2 + \sqrt{ST}}{N})$ for $N=2^n$. In contrast, the best known bound for post-quantum one-way functions is $O(\frac{T^2 + ST}{N})$, with a trivial attack at $S = N$. This demonstrates a new advantage of quantum cryptography over classical cryptography: $n$ qubits of communication suffice for security against preprocessing attacks with space up to $N^2$ rather than $N$. Our methodology is simple: express the optimal preprocessing attack as the operator norm of a random matrix, and bound this value in expectation over the random oracle via the trace-moment method. These trace moments have a natural interpretation using compressed oracles [Zhandry, Crypto 2019], which we then analyze. This can be viewed as a simplification and generalization of the approach of Liu [Eurocrypt 2023] for proving time-space tradeoffs for breaking post-quantum cryptography. We also prove the following results: (1) We tighten Liu's analysis of post-quantum PRGs in QROM, achieving a distinguishing advantage bound of $O(\frac{T^2}N + \sqrt{\frac{ST}N})$. (2) For unitary synthesis, we extend the one-query lower bound of Lombardi-Ma-Wright [STOC 2024] to hold against adversaries that can make one arbitrary function query along with polynomially many (adaptive) queries to the random oracle, either before or after the function query. This also interprets the original LMW24 result in terms of compressed oracles. (3) Finally, we prove a tight $O(\frac{\sqrt{S}}N)$ bound for the pseudorandomness of random binary phase states against space $S$ distinguishers.

Short Resolution Refutations for CNFs with Bounded Weighted Incidence Treewidth

from arXiv: Computational Complexity

Authors: Shaowei Cai, Ziqun Li

It is an open problem in proof complexity whether every unsatisfiable CNF formula has an FPT-sized resolution refutation parameterized by incidence treewidth. In this paper, we establish several upper bounds on resolution refutation length related to this problem. Consider an unsatisfiable CNF formula $F$ with $n$ variables, $m$ clauses, maximum clause width $k$, and incidence treewidth $\mathrm{tw}^*(F)$. In this paper, we introduce two variants of incidence treewidth. Their definitions can be stated informally as follows. The first is log-weighted incidence treewidth $\mathrm{tw}_{\log}^*(F)$, which is the treewidth of the weighted incidence graph, in which variables have weight one and each clause has weight equal to the logarithm of its width. The second is partially log-weighted incidence treewidth $\mathrm{tw}^*_{\mathrm{plog}}(F)$, which is a refinement of log-weighted incidence treewidth. In this variant, for a nice tree decomposition of the incidence graph, each clause has weight one along a path selected for that clause and elsewhere has weight equal to the logarithm of one plus the number of its literals whose variables do not appear in any bag on that path, and variables have weight one. For every unsatisfiable CNF formula $F$, we prove the existence of (i) an FPT-sized resolution refutation parameterized by log-weighted incidence treewidth, with width at most $\mathrm{tw}_{\log}^*(F)+k$; (ii) a resolution refutation of length $(n+m)k^{O(\mathrm{tw}^*(F))}$ and width at most $\mathrm{tw}^*(F)+k$; (iii) an FPT-sized resolution refutation parameterized by partially log-weighted incidence treewidth; and (iv) an FPT-sized regular resolution refutation parameterized by log-weighted incidence treewidth. Our main idea is to construct FPT-sized $k$-DNF resolution refutations parameterized by incidence treewidth, and then convert them into resolution refutations.

Authors: Shaowei Cai, Ziqun Li

It is an open problem in proof complexity whether every unsatisfiable CNF formula has an FPT-sized resolution refutation parameterized by incidence treewidth. In this paper, we establish several upper bounds on resolution refutation length related to this problem. Consider an unsatisfiable CNF formula $F$ with $n$ variables, $m$ clauses, maximum clause width $k$, and incidence treewidth $\mathrm{tw}^*(F)$. In this paper, we introduce two variants of incidence treewidth. Their definitions can be stated informally as follows. The first is log-weighted incidence treewidth $\mathrm{tw}_{\log}^*(F)$, which is the treewidth of the weighted incidence graph, in which variables have weight one and each clause has weight equal to the logarithm of its width. The second is partially log-weighted incidence treewidth $\mathrm{tw}^*_{\mathrm{plog}}(F)$, which is a refinement of log-weighted incidence treewidth. In this variant, for a nice tree decomposition of the incidence graph, each clause has weight one along a path selected for that clause and elsewhere has weight equal to the logarithm of one plus the number of its literals whose variables do not appear in any bag on that path, and variables have weight one. For every unsatisfiable CNF formula $F$, we prove the existence of (i) an FPT-sized resolution refutation parameterized by log-weighted incidence treewidth, with width at most $\mathrm{tw}_{\log}^*(F)+k$; (ii) a resolution refutation of length $(n+m)k^{O(\mathrm{tw}^*(F))}$ and width at most $\mathrm{tw}^*(F)+k$; (iii) an FPT-sized resolution refutation parameterized by partially log-weighted incidence treewidth; and (iv) an FPT-sized regular resolution refutation parameterized by log-weighted incidence treewidth. Our main idea is to construct FPT-sized $k$-DNF resolution refutations parameterized by incidence treewidth, and then convert them into resolution refutations.

Can AI Oversight Be Zero Knowledge?

from arXiv: Computational Complexity

Authors: Alessandro Chiesa, Ziyi Guan, Burcu Yildiz

AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.

Authors: Alessandro Chiesa, Ziyi Guan, Burcu Yildiz

AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.

Trapdoored Clifford Operators and Applications

from arXiv: Computational Complexity

Authors: Minki Hhan, Hojune Lee

Random Clifford operators have numerous applications in quantum computing, including randomized benchmarking, classical shadows, and quantum authentication. However, sampling and implementing uniformly random $n$-qubit Clifford incur near-quadratic complexity due to the size of Clifford group. We introduce a cryptographic way to overcome these barriers: trapdoored Clifford operator distributions whose samples are computationally indistinguishable from uniformly random Cliffords, yet implementing them can be much faster given the trapdoor. We construct a distribution of trapdoored Clifford operators whose elements can be sampled and implemented in near-linear time under a variant of the learning parity with noise assumption. Our constructions allow fast tableau action on Pauli labels for classical simulation, and also can be optimized to admit polylogarithmic-depth implementation. Along the way, we construct trapdoored matrices over finite fields that support efficient multiplication by both a matrix and its inverse, resolving an open question left by Vaikuntanathan and Zamir [SODA'26]. We use these constructions to obtain faster protocols based on random Cliffords. We also explore their applications to the worst-case to average-case reductions for matrix and Clifford problems including the iterated matrix multiplication and Clifford circuit synthesis. In particular, we show the hardness of batching Clifford circuits: synthesizing circuits that apply the same Clifford to multiple registers is at least as hard as worst-case matrix multiplication, even when synthesis succeeds on a small constant fraction of random Cliffords. This extends to approximate implementations by general quantum circuits.

Authors: Minki Hhan, Hojune Lee

Random Clifford operators have numerous applications in quantum computing, including randomized benchmarking, classical shadows, and quantum authentication. However, sampling and implementing uniformly random $n$-qubit Clifford incur near-quadratic complexity due to the size of Clifford group. We introduce a cryptographic way to overcome these barriers: trapdoored Clifford operator distributions whose samples are computationally indistinguishable from uniformly random Cliffords, yet implementing them can be much faster given the trapdoor. We construct a distribution of trapdoored Clifford operators whose elements can be sampled and implemented in near-linear time under a variant of the learning parity with noise assumption. Our constructions allow fast tableau action on Pauli labels for classical simulation, and also can be optimized to admit polylogarithmic-depth implementation. Along the way, we construct trapdoored matrices over finite fields that support efficient multiplication by both a matrix and its inverse, resolving an open question left by Vaikuntanathan and Zamir [SODA'26]. We use these constructions to obtain faster protocols based on random Cliffords. We also explore their applications to the worst-case to average-case reductions for matrix and Clifford problems including the iterated matrix multiplication and Clifford circuit synthesis. In particular, we show the hardness of batching Clifford circuits: synthesizing circuits that apply the same Clifford to multiple registers is at least as hard as worst-case matrix multiplication, even when synthesis succeeds on a small constant fraction of random Cliffords. This extends to approximate implementations by general quantum circuits.

Lower Bound of 22 for 3x3 Matrix Multiplication over the Integers

from arXiv: Computational Complexity

Authors: Isaac Rudich, Louis-Martin Rousseau

Strassen showed that two 2x2 matrices can be multiplied with 7 multiplications instead of 8. Applied recursively, his algorithm multiplies two nxn matrices with O(n^2.807) multiplications, beating the naive O(n^3). The best known 3x3 recursive matrix multiplication algorithm uses 23 multiplications O(n^2.854). The best published lower bound of 21 (on algorithms with integer constants) leaves room for an algorithm with O(n^2.771) multiplications, and thus does not rule out the possibility of an algorithm that would beat Strassen's. We prove a lower bound of 22 multiplications for any 3x3 recursive algorithm with integer constants, proving that no such algorithm can do better than O(n^2.814) multiplications, and eliminating the possibility of a 3x3 algorithm that beats Strassen's 2x2 method. The proof builds on a recent decomposition method from Wang, who approached the problem by turning it into 496 subproblems. We provide exact solutions for 359 of them. The proof is in Lean; verification requires auditing only a few short files. The Lean formalization directly encodes statements about the limitations of recursive algorithms for matrix multiplication, as opposed to just a statement about the rank of the problem.

Authors: Isaac Rudich, Louis-Martin Rousseau

Strassen showed that two 2x2 matrices can be multiplied with 7 multiplications instead of 8. Applied recursively, his algorithm multiplies two nxn matrices with O(n^2.807) multiplications, beating the naive O(n^3). The best known 3x3 recursive matrix multiplication algorithm uses 23 multiplications O(n^2.854). The best published lower bound of 21 (on algorithms with integer constants) leaves room for an algorithm with O(n^2.771) multiplications, and thus does not rule out the possibility of an algorithm that would beat Strassen's. We prove a lower bound of 22 multiplications for any 3x3 recursive algorithm with integer constants, proving that no such algorithm can do better than O(n^2.814) multiplications, and eliminating the possibility of a 3x3 algorithm that beats Strassen's 2x2 method. The proof builds on a recent decomposition method from Wang, who approached the problem by turning it into 496 subproblems. We provide exact solutions for 359 of them. The proof is in Lean; verification requires auditing only a few short files. The Lean formalization directly encodes statements about the limitations of recursive algorithms for matrix multiplication, as opposed to just a statement about the rank of the problem.

Integer reachability in VASS with transfers: a refined complexity analysis

from arXiv: Computational Complexity

Authors: Tymoteusz Kucharek, Piotr Hofman

Integer reachability is NP-complete for vector addition systems with states (VASS), but becomes PSPACE-complete in the presence of transfer operations. We refine this complexity gap for single-transfer VASS by identifying structural features of transfers responsible for the increase in complexity. Each system induces a transfer graph whose vertices are counters and whose edges represent possible transfers. We classify its vertices as good or bad, according to the branching and cyclic structure of their reachable subgraphs. Let $b$ be the number of bad vertices. We show that every positive instance admits a polynomially verifiable certificate of size $|I|^{O(b+1)}$, where $|I|$ is the input size. Consequently, integer reachability for single-transfer VASS can be decided in nondeterministic time $|I|^{O(b+1)}$; in particular, it belongs to NP for every class with a bounded number of bad counters. Conversely, we show that bad counters provide sufficient structural power to encode space-bounded computation. For every transfer graph with $b$ bad vertices, we construct a single-transfer VASS that encodes the acceptance of a Turing machine using $b^{O(1)}$ tape cells. This yields PSPACE-hardness for every polynomial-time constructible family of transfer graphs containing linearly many bad vertices. Our results isolate the transfer patterns responsible for the complexity of integer reachability.

Authors: Tymoteusz Kucharek, Piotr Hofman

Integer reachability is NP-complete for vector addition systems with states (VASS), but becomes PSPACE-complete in the presence of transfer operations. We refine this complexity gap for single-transfer VASS by identifying structural features of transfers responsible for the increase in complexity. Each system induces a transfer graph whose vertices are counters and whose edges represent possible transfers. We classify its vertices as good or bad, according to the branching and cyclic structure of their reachable subgraphs. Let $b$ be the number of bad vertices. We show that every positive instance admits a polynomially verifiable certificate of size $|I|^{O(b+1)}$, where $|I|$ is the input size. Consequently, integer reachability for single-transfer VASS can be decided in nondeterministic time $|I|^{O(b+1)}$; in particular, it belongs to NP for every class with a bounded number of bad counters. Conversely, we show that bad counters provide sufficient structural power to encode space-bounded computation. For every transfer graph with $b$ bad vertices, we construct a single-transfer VASS that encodes the acceptance of a Turing machine using $b^{O(1)}$ tape cells. This yields PSPACE-hardness for every polynomial-time constructible family of transfer graphs containing linearly many bad vertices. Our results isolate the transfer patterns responsible for the complexity of integer reachability.

Exact $T$-counts of Toffoli layers from an isotropy bound

from arXiv: Computational Complexity

Authors: Arul Rhik Mazumder

The $T$-count is the dominant cost of fault-tolerant Clifford+$T$ computation. We prove that a layer of $m$ disjoint CCZ gates, the diagonal core of a parallel Toffoli layer, needs exactly $6m+1$ $T$ gates in every Hadamard-free Clifford+$T$ circuit with clean ancillas. Campbell and Howard gave the matching construction. To our knowledge this is the first proof that it is optimal for general $m$ (for $m=1$ the value $7$ is classical, and Campbell and Howard state the value $13$ at $m=2$). On every controlled unitary the same floor comes within one of their exact count, and it recovers their $4m+3$ for a fan-out of $m$ Toffolis from one control. The proof rests on an isotropy constraint: for a pure-cubic phase, the vectors recording which $T$ gates touch each qubit span a totally isotropic subspace. In general the constraint gives the isotropy floor $δ\ge2(n-d^{\ast})-r$ for every diagonal level-three gate, computed from the phase polynomial in polynomial time. The floor is never below stabilizer nullity $ν$, which equals $n-d^{\ast}$ on this class, and a separate parity argument raises it to $2ν+1$ on non-Clifford pure-cubic gates. On the output of the TODD optimizer for the $24$ benchmark circuits it completes, the floor certifies $193$ of its $311$ merged phase-polynomial blocks optimal for their Hadamard layering ($186$ to $193$ across five optimizer seeds), against $113$ for nullity. The floor also holds, under stated conditions, for circuits whose only internal Hadamards form unitarily uncomputed temporary AND blocks, while under adaptive feedforward only $t\geν$ is proved.

Authors: Arul Rhik Mazumder

The $T$-count is the dominant cost of fault-tolerant Clifford+$T$ computation. We prove that a layer of $m$ disjoint CCZ gates, the diagonal core of a parallel Toffoli layer, needs exactly $6m+1$ $T$ gates in every Hadamard-free Clifford+$T$ circuit with clean ancillas. Campbell and Howard gave the matching construction. To our knowledge this is the first proof that it is optimal for general $m$ (for $m=1$ the value $7$ is classical, and Campbell and Howard state the value $13$ at $m=2$). On every controlled unitary the same floor comes within one of their exact count, and it recovers their $4m+3$ for a fan-out of $m$ Toffolis from one control. The proof rests on an isotropy constraint: for a pure-cubic phase, the vectors recording which $T$ gates touch each qubit span a totally isotropic subspace. In general the constraint gives the isotropy floor $δ\ge2(n-d^{\ast})-r$ for every diagonal level-three gate, computed from the phase polynomial in polynomial time. The floor is never below stabilizer nullity $ν$, which equals $n-d^{\ast}$ on this class, and a separate parity argument raises it to $2ν+1$ on non-Clifford pure-cubic gates. On the output of the TODD optimizer for the $24$ benchmark circuits it completes, the floor certifies $193$ of its $311$ merged phase-polynomial blocks optimal for their Hadamard layering ($186$ to $193$ across five optimizer seeds), against $113$ for nullity. The floor also holds, under stated conditions, for circuits whose only internal Hadamards form unitarily uncomputed temporary AND blocks, while under adaptive feedforward only $t\geν$ is proved.

Beyond IP = PSPACE and QIP = PSPACE: Interactive Proofs in Arbitrary Physical Theories

from arXiv: Computational Complexity

Authors: Kishor Bharti

The equalities IP = PSPACE and QIP = PSPACE, the latter achievable with three messages, raise a basic question: how much of an interactive proof's power comes from the underlying physical theory? We study interactive proofs in general probabilistic theories, which include classical and quantum theory. The answer depends on what the prover and verifier exchange and how the theory specifies efficient operations. When they exchange only classical messages, protocols in every theory satisfying our standard assumptions decide exactly PSPACE. For protocols with a quantum verifier and quantum messages, allowing a prover to use any theory containing quantum theory does not increase the maximum acceptance probability. Thus, the three-message PSPACE result remains valid against such provers. When messages may be arbitrary systems, the interactive-proof class can strictly exceed PSPACE.

Authors: Kishor Bharti

The equalities IP = PSPACE and QIP = PSPACE, the latter achievable with three messages, raise a basic question: how much of an interactive proof's power comes from the underlying physical theory? We study interactive proofs in general probabilistic theories, which include classical and quantum theory. The answer depends on what the prover and verifier exchange and how the theory specifies efficient operations. When they exchange only classical messages, protocols in every theory satisfying our standard assumptions decide exactly PSPACE. For protocols with a quantum verifier and quantum messages, allowing a prover to use any theory containing quantum theory does not increase the maximum acceptance probability. Thus, the three-message PSPACE result remains valid against such provers. When messages may be arbitrary systems, the interactive-proof class can strictly exceed PSPACE.

A Degree--Size Relation for Resolution over Polynomials

from arXiv: Computational Complexity

Authors: Shuo Pang

For every constant-width CNF, we show that linear degree in polynomial calculus (PC) implies exponential size in resolution over constant-degree polynomials, over the same prime field. Applications include exponential lower bounds for CNFs in $\operatorname{Res}(\operatorname{PC}_r/\mathbb{F}_p)$ and hence in $\operatorname{Res}(\oplus_p)$, separations between different moduli, improved lower bounds for $\operatorname{Res}(k)$ up to $k=\varepsilon\log n$, proof-search consequences, and an implication of super-polynomial $AC^0[p]$-Frege bounds from very strong PC degree lower bounds. The proof uses the common-multiplier idea isolated from Braun [arXiv:2609.23015] to construct a Razborov--Smolensky approximation that preserves inferences, without introducing extension variables. The approximation errors are measured by ranks of the multiplication maps induced by the error-witness polynomials, modulo bounded-degree PC consequences.

Authors: Shuo Pang

For every constant-width CNF, we show that linear degree in polynomial calculus (PC) implies exponential size in resolution over constant-degree polynomials, over the same prime field. Applications include exponential lower bounds for CNFs in $\operatorname{Res}(\operatorname{PC}_r/\mathbb{F}_p)$ and hence in $\operatorname{Res}(\oplus_p)$, separations between different moduli, improved lower bounds for $\operatorname{Res}(k)$ up to $k=\varepsilon\log n$, proof-search consequences, and an implication of super-polynomial $AC^0[p]$-Frege bounds from very strong PC degree lower bounds. The proof uses the common-multiplier idea isolated from Braun [arXiv:2609.23015] to construct a Razborov--Smolensky approximation that preserves inferences, without introducing extension variables. The approximation errors are measured by ranks of the multiplication maps induced by the error-witness polynomials, modulo bounded-degree PC consequences.

Entrywise Logarithmic Matrix Algebra and Dichotomy of Planar Graph Homomorphisms (Part I)

from arXiv: Computational Complexity

Authors: Jin-Yi Cai, Zhuxiao Tang

We prove a complexity classification of counting planar graph homomorphisms with non-negative weights. For a real symmetric matrix $M$ with non-negative entries, the problem $\PlGH(M)$ is either (1) P-time computable over all graphs, or (2) \#P-hard in general but P-time computable over planar graphs, or (3) \#P-hard over planar graphs. Furthermore, $\PlGH(M)$ in (2) consists of precisely those that involve the P-time FKT algorithm to count planar perfect matchings with a holographic transformation. The dichotomy is achieved by forming a (centered) logarithmic matrix algebra (a vector space with bilinear multiplication) by taking entrywise logarithms of all realizable matrices from $M$ using planar edge gadgets and polynomial interpolation. The current version is part I, which contains the proof for the dichotomy of entrywise positive and positive definite matrices, which is at the core of the dichotomy for non-negative matrices. Part II contains the extension from entrywise positive and positive definite matrices to non-negative matrices.

Authors: Jin-Yi Cai, Zhuxiao Tang

We prove a complexity classification of counting planar graph homomorphisms with non-negative weights. For a real symmetric matrix $M$ with non-negative entries, the problem $\PlGH(M)$ is either (1) P-time computable over all graphs, or (2) \#P-hard in general but P-time computable over planar graphs, or (3) \#P-hard over planar graphs. Furthermore, $\PlGH(M)$ in (2) consists of precisely those that involve the P-time FKT algorithm to count planar perfect matchings with a holographic transformation. The dichotomy is achieved by forming a (centered) logarithmic matrix algebra (a vector space with bilinear multiplication) by taking entrywise logarithms of all realizable matrices from $M$ using planar edge gadgets and polynomial interpolation. The current version is part I, which contains the proof for the dichotomy of entrywise positive and positive definite matrices, which is at the core of the dichotomy for non-negative matrices. Part II contains the extension from entrywise positive and positive definite matrices to non-negative matrices.

Approximate Polynomial Satisfiability is in the Counting Hierarchy

from arXiv: Computational Complexity

Authors: Nikhil Balaji, Mahsa Shirmohammadi, Sébastien Tavenas, James Worrell

The Approximate polynomial satisfiability problem (APS), introduced by Guo, Saxena, and Sinhababu (CCC 2018), asks whether the zero vector lies in the Zariski closure of the image of a given polynomial map. Specifically, for a field $k$ with algebraic closure~$K$, the problem asks whether $\boldsymbol 0 \in\overline{\boldsymbol f(K^n)}$ for a polynomial map $\boldsymbol f=(f_1,\ldots,f_m)$ with $f_i\in k[X_1,\ldots,X_n]$. APS is a natural topological analogue of Hilbert's Nullstellensatz, namely the question of whether a given system of polynomial equations has a common zero. APS captures several problems in algebraic complexity, including border rank, hitting sets for border classes, and null-cone membership; it is known to be NP-hard and in PSPACE. We show that APS lies in the Counting Hierarchy (CH) over both the rationals and finite fields, substantially improving the known PSPACE upper bound. Our proof builds on a recent breakthrough due to Andrews, Garg, and Schost (FOCS 2026) on deciding Hilbert's Nullstellensatz in CH. As a corollary, our result improves the complexity of certifying hitting sets for border classes from PSPACE to CH. We also give a polynomial-time reduction of Hilbert's Nullstellensatz to APS, valid in any characteristic. In characteristic zero, we give a reduction of APS to the decision problem for the existential theory of real closed fields. Overall, our results place approximate polynomial satisfiability closer in complexity to exact polynomial feasibility and as a byproduct give improved complexity bounds for several problems arising in approximative complexity.

Authors: Nikhil Balaji, Mahsa Shirmohammadi, Sébastien Tavenas, James Worrell

The Approximate polynomial satisfiability problem (APS), introduced by Guo, Saxena, and Sinhababu (CCC 2018), asks whether the zero vector lies in the Zariski closure of the image of a given polynomial map. Specifically, for a field $k$ with algebraic closure~$K$, the problem asks whether $\boldsymbol 0 \in\overline{\boldsymbol f(K^n)}$ for a polynomial map $\boldsymbol f=(f_1,\ldots,f_m)$ with $f_i\in k[X_1,\ldots,X_n]$. APS is a natural topological analogue of Hilbert's Nullstellensatz, namely the question of whether a given system of polynomial equations has a common zero. APS captures several problems in algebraic complexity, including border rank, hitting sets for border classes, and null-cone membership; it is known to be NP-hard and in PSPACE. We show that APS lies in the Counting Hierarchy (CH) over both the rationals and finite fields, substantially improving the known PSPACE upper bound. Our proof builds on a recent breakthrough due to Andrews, Garg, and Schost (FOCS 2026) on deciding Hilbert's Nullstellensatz in CH. As a corollary, our result improves the complexity of certifying hitting sets for border classes from PSPACE to CH. We also give a polynomial-time reduction of Hilbert's Nullstellensatz to APS, valid in any characteristic. In characteristic zero, we give a reduction of APS to the decision problem for the existential theory of real closed fields. Overall, our results place approximate polynomial satisfiability closer in complexity to exact polynomial feasibility and as a byproduct give improved complexity bounds for several problems arising in approximative complexity.

Polynomial-time local-unitary equivalence of graph states

from arXiv: Computational Complexity

Authors: Yuxuan Zhang

Local-unitary (LU) equivalence asks whether two quantum states differ only by independent changes of basis on their qubits. For graph states, whether this relation can be decided in polynomial time has remained open for over a decade. We give a deterministic algorithm that decides LU equivalence for graphs on $n$ labelled vertices in $\widetilde O(n^{6.38})$ bit operations and constructs exact single-qubit unitaries whenever the states are equivalent. Building on Claudet and Perdrix's quasipolynomial algorithm, we replace the enumeration of vertex subsets by a compact system of constraints generated from pairs and triples. The remaining graph transformation is found by solving linear equations over the binary field. These new steps cost $\widetilde O(n^5)$ bit operations; the inherited graph preprocessing sets the overall bound. We also count the local-Clifford (LC) classes of graph states within any LU class: their number is a power of two, computable within the same bound. For any given graph state, this decides whether single-qubit Clifford gates reach every graph state in its LU class, and supplies a counterexample when they do not. The method also decides LU equivalence of stabilizer codes encoding one logical qubit.

Authors: Yuxuan Zhang

Local-unitary (LU) equivalence asks whether two quantum states differ only by independent changes of basis on their qubits. For graph states, whether this relation can be decided in polynomial time has remained open for over a decade. We give a deterministic algorithm that decides LU equivalence for graphs on $n$ labelled vertices in $\widetilde O(n^{6.38})$ bit operations and constructs exact single-qubit unitaries whenever the states are equivalent. Building on Claudet and Perdrix's quasipolynomial algorithm, we replace the enumeration of vertex subsets by a compact system of constraints generated from pairs and triples. The remaining graph transformation is found by solving linear equations over the binary field. These new steps cost $\widetilde O(n^5)$ bit operations; the inherited graph preprocessing sets the overall bound. We also count the local-Clifford (LC) classes of graph states within any LU class: their number is a power of two, computable within the same bound. For any given graph state, this decides whether single-qubit Clifford gates reach every graph state in its LU class, and supplies a counterexample when they do not. The method also decides LU equivalence of stabilizer codes encoding one logical qubit.

Good Quantum Locally Testable Codes from Lossless Cubical Complexes

from arXiv: Computational Complexity

Authors: Itay Cohen, Itai Leigh, Assaf Reiner, Amnon Ta-Shma, Elad Tzalik

Sipser and Spielman constructed LDPC codes from either bipartite \emph{spectral} expanders or one-sided \emph{lossless} expanders. In higher dimensions, \emph{spectral} expansion similarly played a central role in the constructions of asymptotically good classical LTCs and qLDPC codes by Dinur, Evra, Livne, Lubotzky, and Mozes and by Panteleev and Kalachev. Alternatively, Lin and Hsieh constructed classical LTCs and qLDPC codes from two-dimensional \emph{lossless} cubical complexes. In this work we develop the higher-dimensional \emph{lossless} approach. We do not construct the required high-dimensional lossless cubical complexes; rather, we investigate what their existence would imply. We associate with a high-dimensional cubical complex a \emph{level chain complex}, whose chain groups are supported on the level sets of the Boolean cube rather than on its cells. Our main technical contribution is a clean local-to-global theorem for this structure: suitable one-dimensional lossless expansion in the directional graphs implies small-set coboundary expansion of the global level complex. As a consequence, sufficiently imbalanced, two-sided lossless four-dimensional cubical complexes give rise to asymptotically good quantum locally testable codes. We expect the local-to-global principle developed here to have further applications.

Authors: Itay Cohen, Itai Leigh, Assaf Reiner, Amnon Ta-Shma, Elad Tzalik

Sipser and Spielman constructed LDPC codes from either bipartite \emph{spectral} expanders or one-sided \emph{lossless} expanders. In higher dimensions, \emph{spectral} expansion similarly played a central role in the constructions of asymptotically good classical LTCs and qLDPC codes by Dinur, Evra, Livne, Lubotzky, and Mozes and by Panteleev and Kalachev. Alternatively, Lin and Hsieh constructed classical LTCs and qLDPC codes from two-dimensional \emph{lossless} cubical complexes. In this work we develop the higher-dimensional \emph{lossless} approach. We do not construct the required high-dimensional lossless cubical complexes; rather, we investigate what their existence would imply. We associate with a high-dimensional cubical complex a \emph{level chain complex}, whose chain groups are supported on the level sets of the Boolean cube rather than on its cells. Our main technical contribution is a clean local-to-global theorem for this structure: suitable one-dimensional lossless expansion in the directional graphs implies small-set coboundary expansion of the global level complex. As a consequence, sufficiently imbalanced, two-sided lossless four-dimensional cubical complexes give rise to asymptotically good quantum locally testable codes. We expect the local-to-global principle developed here to have further applications.

Noisy Quantum Query Complexity via Fractional Block Sensitivity

from arXiv: Computational Complexity

Authors: Mehil Agarwal, Shravas Rao, Fang Song

We study quantum query complexity under several models of imperfect oracle access, and develop lower bounds through a common framework based on fractional block sensitivity (\(\fbs\)). For a negligent oracle that applies the correct query with probability \(1-p\), we prove a lower bound in terms of \(\fbs(f,0^n)\). We also show, perhaps surprisingly, that negligence need not destroy quantum speedups: any function \(f\) can be transformed into a partial function \(f'\) whose negligent query complexity essentially preserves the quantum query complexity of \(f\). This gives partial functions with exponential quantum speedups even under negligent queries, and in particular rules out a general lower bound in terms of \(\fbs(f)\) for partial functions in this model. For two other models, we obtain general lower bounds in terms of \(\fbs(f)\) for all Boolean functions. For hybrid algorithms using \(Q\) coherent and \(C\) classical queries, we prove the tradeoff \(C+Q^2=Ω(\fbs(f))\). For an IID dephasing noisy model where each query dephases the query-index register at rate \(p\) independently, we prove \(Ω\!\left(p\,\fbs(f)\right)\) queries are necessary. Finally, we introduce a broader family of time-varying dephasing models and identify a variational resource that is always lower bounded by \(\fbs(f)\). Computing this resource reduces to a convex optimization problem, providing a simple way to derive lower bounds for new noise schedules. As applications, we recover the hybrid and IID dephasing bounds and determine the query complexity of unstructured search when the dephasing rate grows over time.

Authors: Mehil Agarwal, Shravas Rao, Fang Song

We study quantum query complexity under several models of imperfect oracle access, and develop lower bounds through a common framework based on fractional block sensitivity (\(\fbs\)). For a negligent oracle that applies the correct query with probability \(1-p\), we prove a lower bound in terms of \(\fbs(f,0^n)\). We also show, perhaps surprisingly, that negligence need not destroy quantum speedups: any function \(f\) can be transformed into a partial function \(f'\) whose negligent query complexity essentially preserves the quantum query complexity of \(f\). This gives partial functions with exponential quantum speedups even under negligent queries, and in particular rules out a general lower bound in terms of \(\fbs(f)\) for partial functions in this model. For two other models, we obtain general lower bounds in terms of \(\fbs(f)\) for all Boolean functions. For hybrid algorithms using \(Q\) coherent and \(C\) classical queries, we prove the tradeoff \(C+Q^2=Ω(\fbs(f))\). For an IID dephasing noisy model where each query dephases the query-index register at rate \(p\) independently, we prove \(Ω\!\left(p\,\fbs(f)\right)\) queries are necessary. Finally, we introduce a broader family of time-varying dephasing models and identify a variational resource that is always lower bounded by \(\fbs(f)\). Computing this resource reduces to a convex optimization problem, providing a simple way to derive lower bounds for new noise schedules. As applications, we recover the hybrid and IID dephasing bounds and determine the query complexity of unstructured search when the dephasing rate grows over time.

Hyperbolic Sphericity

from arXiv: Computational Geometry

Authors: Thomas Bläsius, Lennart Großkreutz, Jean-Pierre von der Heydt

The sphericity of a graph is the minimum dimension d such that the graph has an intersection representation of d-dimensional balls of equal radius. While sphericity has been studied in Euclidean space, we initiate the study of hyperbolic sphericity. The hyperbolic sphericity of a graph can be significantly smaller than its Euclidean counterpart, but, contrary to the Euclidean setting, depends strongly on the radius of the balls. We show that, if the radius of the balls can be chosen depending on the graph, the hyperbolic sphericity is upper bounded by the Euclidean sphericity. This extends a previous result for 2-dimensional hyperbolic space, i.e., uniform disk graphs, to arbitrary dimensions. Moreover, our proof is significantly simpler. If we fix the radius, i.e., do not make it dependent on the graph, we show that hyperbolic sphericity can be larger than Euclidean sphericity, but by at most 1. Additionally, we study how hyperbolic sphericity changes with the ball radius. We show that choosing a larger radius can substantially decrease the sphericity while increasing it by at most 1. We also provide a construction of a graph where the sphericity oscillates between different values as the radius increases. Besides being theoretically interesting, we note that these results are relevant for graph embeddings in machine learning, where one is interested in low-dimensional numeric representations of symbolic data like graphs.

Authors: Thomas Bläsius, Lennart Großkreutz, Jean-Pierre von der Heydt

The sphericity of a graph is the minimum dimension d such that the graph has an intersection representation of d-dimensional balls of equal radius. While sphericity has been studied in Euclidean space, we initiate the study of hyperbolic sphericity. The hyperbolic sphericity of a graph can be significantly smaller than its Euclidean counterpart, but, contrary to the Euclidean setting, depends strongly on the radius of the balls. We show that, if the radius of the balls can be chosen depending on the graph, the hyperbolic sphericity is upper bounded by the Euclidean sphericity. This extends a previous result for 2-dimensional hyperbolic space, i.e., uniform disk graphs, to arbitrary dimensions. Moreover, our proof is significantly simpler. If we fix the radius, i.e., do not make it dependent on the graph, we show that hyperbolic sphericity can be larger than Euclidean sphericity, but by at most 1. Additionally, we study how hyperbolic sphericity changes with the ball radius. We show that choosing a larger radius can substantially decrease the sphericity while increasing it by at most 1. We also provide a construction of a graph where the sphericity oscillates between different values as the radius increases. Besides being theoretically interesting, we note that these results are relevant for graph embeddings in machine learning, where one is interested in low-dimensional numeric representations of symbolic data like graphs.

Optimal Coresets for Hyperbolic Farthest-Point Queries via Ideal-Boundary Envelopes

from arXiv: Computational Geometry

Authors: Eunku Park

We study coresets for farthest-point queries in hyperbolic space. Given a nonempty finite set $P \subset \mathbb{H}^D$ and $0<\varepsilon \le 1$, we seek a coreset $P_{\varepsilon} \subseteq P$ whose farthest distance from every query point underestimates that of $P$ by at most an additive $\varepsilon$ and retains at least a $1-\varepsilon$ fraction of it. For every fixed $D \ge 2$, we prove that the optimal worst-case coreset size is $Θ\bigl(\varepsilon^{-(D-1)/2}\bigr)$. Our main geometric ingredient is an exact reduction from hyperbolic queries to an upper envelope on the ideal boundary. In the hyperboloid model, each input point induces a positive boundary-score function whose logarithm gives its asymptotic distance offset along geodesic rays. We define the \emph{ideal-boundary envelope} as the pointwise maximum of these functions and prove that the supremum additive loss over all queries equals the maximum logarithmic gap between the input and coreset envelopes. For the upper bound, we move the minimum-enclosing-ball center to the origin and normalize the spatial coordinates, obtaining a bounded Euclidean point set whose boundary envelope is bounded away from zero. A standard Euclidean kernel then approximates all directional score maxima simultaneously, and the structure theorem yields both guarantees. For the lower bound, a spherical packing on a fixed-radius hyperbolic sphere, together with antipodal queries and the hyperbolic cosine law, makes every input point indispensable, matching the upper bound even for either guarantee separately.

Authors: Eunku Park

We study coresets for farthest-point queries in hyperbolic space. Given a nonempty finite set $P \subset \mathbb{H}^D$ and $0<\varepsilon \le 1$, we seek a coreset $P_{\varepsilon} \subseteq P$ whose farthest distance from every query point underestimates that of $P$ by at most an additive $\varepsilon$ and retains at least a $1-\varepsilon$ fraction of it. For every fixed $D \ge 2$, we prove that the optimal worst-case coreset size is $Θ\bigl(\varepsilon^{-(D-1)/2}\bigr)$. Our main geometric ingredient is an exact reduction from hyperbolic queries to an upper envelope on the ideal boundary. In the hyperboloid model, each input point induces a positive boundary-score function whose logarithm gives its asymptotic distance offset along geodesic rays. We define the \emph{ideal-boundary envelope} as the pointwise maximum of these functions and prove that the supremum additive loss over all queries equals the maximum logarithmic gap between the input and coreset envelopes. For the upper bound, we move the minimum-enclosing-ball center to the origin and normalize the spatial coordinates, obtaining a bounded Euclidean point set whose boundary envelope is bounded away from zero. A standard Euclidean kernel then approximates all directional score maxima simultaneously, and the structure theorem yields both guarantees. For the lower bound, a spherical packing on a fixed-radius hyperbolic sphere, together with antipodal queries and the hyperbolic cosine law, makes every input point indispensable, matching the upper bound even for either guarantee separately.

Towards Strongly Aperiodic Monotiles in Higher Dimensions

from arXiv: Computational Geometry

Authors: Dmitry Kamenetsky

The discovery of Chair44 (Tsiokos, 2026) settled the three-dimensional einstein problem with a strongly aperiodic polyhedral monotile in $\mathbb{R}^3$. This note extends the underlying mechanism---the rep-$2^N$ chair $C_N = [0,2]^N \setminus (1,2]^N$ with corner/socket markings---to $\mathbb{R}^N$. Besides expository material (the rep-$2^N$ dissection and a conditional strong-aperiodicity theorem under lattice registration and hierarchical enforcement), the note makes a new computational contribution. We introduce a frame-marking formalism in which the marking of a tile is its full orientation frame and the matching rule is the contact language generated by the substitution itself; this makes the search for matching rules finite in every dimension. We give a finite certificate (coarsening closure, tightness, and a two-shell enclosure analysis) whose validity implies that every lattice-registered tiling by the marked tile is uniquely hierarchical, hence strongly aperiodic. For $N=3$ the certificate passes: it yields explicit facet matching rules on the 24 panels of $C_3$ (135 admissible facet-contact triples) and reproduces, from first principles and independently of published constructions, the Chair44 statistics 2388 $\to$ 44 admissible contacts (30 occurring), 33 one-shell clusters, 15 extendable, each forcing a unique supertile. Among the 2187 homochiral frame assignments of the 3D substitution with a translated central child, the certified one is unique up to conjugation. For $N=4$ the same pipeline is run on several structured families of frame assignments (canonical, $D_4$-, $Z_2\times Z_2$- and $Z_4$-symmetric, and a lift of the 3D solution); none is coarsening-closed, and we report the failure data. A self-similar marking of $C_4$ thus remains an explicitly finite, open computational problem, which we state precisely. Code: github.com/dimkadimon/Monotile-RN

Authors: Dmitry Kamenetsky

The discovery of Chair44 (Tsiokos, 2026) settled the three-dimensional einstein problem with a strongly aperiodic polyhedral monotile in $\mathbb{R}^3$. This note extends the underlying mechanism---the rep-$2^N$ chair $C_N = [0,2]^N \setminus (1,2]^N$ with corner/socket markings---to $\mathbb{R}^N$. Besides expository material (the rep-$2^N$ dissection and a conditional strong-aperiodicity theorem under lattice registration and hierarchical enforcement), the note makes a new computational contribution. We introduce a frame-marking formalism in which the marking of a tile is its full orientation frame and the matching rule is the contact language generated by the substitution itself; this makes the search for matching rules finite in every dimension. We give a finite certificate (coarsening closure, tightness, and a two-shell enclosure analysis) whose validity implies that every lattice-registered tiling by the marked tile is uniquely hierarchical, hence strongly aperiodic. For $N=3$ the certificate passes: it yields explicit facet matching rules on the 24 panels of $C_3$ (135 admissible facet-contact triples) and reproduces, from first principles and independently of published constructions, the Chair44 statistics 2388 $\to$ 44 admissible contacts (30 occurring), 33 one-shell clusters, 15 extendable, each forcing a unique supertile. Among the 2187 homochiral frame assignments of the 3D substitution with a translated central child, the certified one is unique up to conjugation. For $N=4$ the same pipeline is run on several structured families of frame assignments (canonical, $D_4$-, $Z_2\times Z_2$- and $Z_4$-symmetric, and a lift of the 3D solution); none is coarsening-closed, and we report the failure data. A self-similar marking of $C_4$ thus remains an explicitly finite, open computational problem, which we state precisely. Code: https://github.com/dimkadimon/Monotile-RN

Minimum Spanning Trees for Square Crop Plots

from arXiv: Computational Geometry

Authors: Mingyang Gong, Adiesha Liyanage, Braeden Sopp, Muzhou Chen, Binhai Zhu

Motivated by accessing crop plots in a field, where each crop plot can only be visited by a given pair of entry/exit points (we have three different types, each having four cases), we study the corresponding Minimum Spanning Tree (MST) problem of such a planar set of axis-aligned unit squares. It turns out that this crop plot distance does not satisfy triangle inequality, even though the whole setup of the problem is geometric. Hence additional care must be taken. The main results of this paper are as follows: (1) we prove that computing the MST of a set $P$ of unit squares under the crop plot distance is NP-hard, (2) the MST of $P$ can be approximated with a factor-4 approximation.

Authors: Mingyang Gong, Adiesha Liyanage, Braeden Sopp, Muzhou Chen, Binhai Zhu

Motivated by accessing crop plots in a field, where each crop plot can only be visited by a given pair of entry/exit points (we have three different types, each having four cases), we study the corresponding Minimum Spanning Tree (MST) problem of such a planar set of axis-aligned unit squares. It turns out that this crop plot distance does not satisfy triangle inequality, even though the whole setup of the problem is geometric. Hence additional care must be taken. The main results of this paper are as follows: (1) we prove that computing the MST of a set $P$ of unit squares under the crop plot distance is NP-hard, (2) the MST of $P$ can be approximated with a factor-4 approximation.

$k$-arrangements of pseudolines and pseudocircles

from arXiv: Computational Geometry

Authors: Jan Kynčl, Carolina Medina, Gelasio Salazar

A $k$-arrangement of pseudolines is a set of bi-infinite curves in the plane such that any two of them intersect each other in exactly $k$ points, at which they cross, and it is simple if no three curves meet at a common point. Cyclic arrangements are the only simple $1$-arrangements of pseudolines that are unavoidable, in the Ramsey spirit: for each fixed $m\ge 1$, every sufficiently large simple $1$-arrangement of pseudolines has a cyclic subarrangement of size $m$. We show that, for every $m\ge 3$, the number of unavoidable simple $k$-arrangements of pseudolines of size $m$ grows exponentially with $k$, independently of $m$. For even $k$, we prove an analogous result for $k$-arrangements of pseudocircles.

Authors: Jan Kynčl, Carolina Medina, Gelasio Salazar

A $k$-arrangement of pseudolines is a set of bi-infinite curves in the plane such that any two of them intersect each other in exactly $k$ points, at which they cross, and it is simple if no three curves meet at a common point. Cyclic arrangements are the only simple $1$-arrangements of pseudolines that are unavoidable, in the Ramsey spirit: for each fixed $m\ge 1$, every sufficiently large simple $1$-arrangement of pseudolines has a cyclic subarrangement of size $m$. We show that, for every $m\ge 3$, the number of unavoidable simple $k$-arrangements of pseudolines of size $m$ grows exponentially with $k$, independently of $m$. For even $k$, we prove an analogous result for $k$-arrangements of pseudocircles.

Polynomial-time classical and quantum simulation of quantum impurity models

from arXiv: Data Structures and Algorithms

Authors: Jiaqing Jiang, Nathan Ju, Ojas Parekh, Chaithanya Rayudu, Andrew Zhao

Quantum impurity models are paradigmatic models of interacting quantum matter, as well as key computational primitives for modern electronic-structure methods. They describe a small subsystem of interacting fermions coupled to a large, noninteracting bath. We perform a comprehensive study of the computational complexity of simulating impurity models, delineating the boundary between classical and quantum tractability for this class of problems. Our main finding is that static properties of quantum impurity models can be calculated efficiently on a classical computer. Specifically, we give classical algorithms that (1) estimate the ground-state energy to additive precision $δ$ in time $\mathrm{poly}(n,δ^{-1})$, and (2) estimate the partition function at inverse temperature $β$ to relative precision $δ$ in time $\mathrm{poly}(n,β,δ^{-1})$, where $n$ is the system size. These results improve the previous best-known complexity for ground-state energy estimation from quasipolynomial to polynomial time, while establishing for the first time rigorous polynomial-time guarantees for simulating impurity models in thermal equilibrium. On the other hand, we find that simulating dynamical properties of impurity models is hard for classical computers but easy on a quantum computer. As a canonical example, we show that computing their nonequilibrium Green's functions captures the full power of quantum computation, even at finite temperature. Taken together, our results rule out superpolynomial quantum speedups for computing static properties, but provide an avenue for quantum advantage in simulating impurity physics out of equilibrium.

Authors: Jiaqing Jiang, Nathan Ju, Ojas Parekh, Chaithanya Rayudu, Andrew Zhao

Quantum impurity models are paradigmatic models of interacting quantum matter, as well as key computational primitives for modern electronic-structure methods. They describe a small subsystem of interacting fermions coupled to a large, noninteracting bath. We perform a comprehensive study of the computational complexity of simulating impurity models, delineating the boundary between classical and quantum tractability for this class of problems. Our main finding is that static properties of quantum impurity models can be calculated efficiently on a classical computer. Specifically, we give classical algorithms that (1) estimate the ground-state energy to additive precision $δ$ in time $\mathrm{poly}(n,δ^{-1})$, and (2) estimate the partition function at inverse temperature $β$ to relative precision $δ$ in time $\mathrm{poly}(n,β,δ^{-1})$, where $n$ is the system size. These results improve the previous best-known complexity for ground-state energy estimation from quasipolynomial to polynomial time, while establishing for the first time rigorous polynomial-time guarantees for simulating impurity models in thermal equilibrium. On the other hand, we find that simulating dynamical properties of impurity models is hard for classical computers but easy on a quantum computer. As a canonical example, we show that computing their nonequilibrium Green's functions captures the full power of quantum computation, even at finite temperature. Taken together, our results rule out superpolynomial quantum speedups for computing static properties, but provide an avenue for quantum advantage in simulating impurity physics out of equilibrium.

Polynomial-time additive-error estimation of output probabilities for shallow quantum circuits

from arXiv: Data Structures and Algorithms

Authors: Matthew Coudron, Michael J. Gullans, Jon Nelson, Joel Rajakumar, Shi Jie Samuel Tan

We give a deterministic classical algorithm that estimates $|\langle x|U|0^n\rangle|^2$ to additive error $\varepsilon$ in $\mathrm{poly}(n, 1/\varepsilon)$ time, where $U$ is a constant-depth quantum circuit comprised of gates with bounded fan-in and arbitrary connectivity, and $x$ is an arbitrary $n$-bit output string. This improves over prior state-of-the-art algorithms that takes $n^{O(log(n))}$ time for the same task, $n^{O(log(log(n))}$ when $U$ is geometrically local, and $n^{O(1)}$ for 2D geometrically-local circuits.

Authors: Matthew Coudron, Michael J. Gullans, Jon Nelson, Joel Rajakumar, Shi Jie Samuel Tan

We give a deterministic classical algorithm that estimates $|\langle x|U|0^n\rangle|^2$ to additive error $\varepsilon$ in $\mathrm{poly}(n, 1/\varepsilon)$ time, where $U$ is a constant-depth quantum circuit comprised of gates with bounded fan-in and arbitrary connectivity, and $x$ is an arbitrary $n$-bit output string. This improves over prior state-of-the-art algorithms that takes $n^{O(log(n))}$ time for the same task, $n^{O(log(log(n))}$ when $U$ is geometrically local, and $n^{O(1)}$ for 2D geometrically-local circuits.

A provable quantum advantage for approximate optimization via decoded quantum interferometry

from arXiv: Data Structures and Algorithms

Authors: Maximilian J. Kramer, Elies Gil-Fuster, Benjamin D. M. Jones, Jens Eisert, Franz J. Schreiber

Decoded quantum interferometry (DQI) is a novel paradigm for tackling approximate optimization problems on quantum computers. This framework comes with strong performance guarantees and exploits a well-established duality between optimization and coding theory. A central question, however, is whether DQI can actually provably outperform all polynomial-time classical algorithms. In this work, we establish such an advantage in an oracle setting: we consider an optimization task called folded optimal polynomial intersection (folded OPI), where the acceptance sets are chosen randomly and accessed through membership oracles. We establish a strict gap between the approximation ratio achievable by any polynomial-time classical algorithm and the approximation ratio achieved by the DQI algorithm. Our proof builds on Jordan et al.'s DQI framework for approximate optimization and extends the classical lower-bound method underlying Yamakawa and Zhandry's exact-search oracle separation to approximation. Building on recent developments by Sun and Wootters, Horinaga and Yamakawa, and Jo, we further show that a modified version of the DQI algorithm achieves a strictly larger gap on the folded OPI problem, yielding an even stronger quantum separation. As a concrete example, for code rate $0.3$, DQI and the modified algorithm achieve expected scores of approximately $0.85$ and $0.95$, respectively. In contrast, exceeding the classical threshold of $0.65$ by any fixed amount with constant probability on sampled instances requires super-polynomially many classical membership queries.

Authors: Maximilian J. Kramer, Elies Gil-Fuster, Benjamin D. M. Jones, Jens Eisert, Franz J. Schreiber

Decoded quantum interferometry (DQI) is a novel paradigm for tackling approximate optimization problems on quantum computers. This framework comes with strong performance guarantees and exploits a well-established duality between optimization and coding theory. A central question, however, is whether DQI can actually provably outperform all polynomial-time classical algorithms. In this work, we establish such an advantage in an oracle setting: we consider an optimization task called folded optimal polynomial intersection (folded OPI), where the acceptance sets are chosen randomly and accessed through membership oracles. We establish a strict gap between the approximation ratio achievable by any polynomial-time classical algorithm and the approximation ratio achieved by the DQI algorithm. Our proof builds on Jordan et al.'s DQI framework for approximate optimization and extends the classical lower-bound method underlying Yamakawa and Zhandry's exact-search oracle separation to approximation. Building on recent developments by Sun and Wootters, Horinaga and Yamakawa, and Jo, we further show that a modified version of the DQI algorithm achieves a strictly larger gap on the folded OPI problem, yielding an even stronger quantum separation. As a concrete example, for code rate $0.3$, DQI and the modified algorithm achieve expected scores of approximately $0.85$ and $0.95$, respectively. In contrast, exceeding the classical threshold of $0.65$ by any fixed amount with constant probability on sampled instances requires super-polynomially many classical membership queries.

Exponential quantum advantages for decoded quantum interferometry in the streaming setting

from arXiv: Data Structures and Algorithms

Authors: Kewen Wu, Guangxu Yang

Decoded quantum interferometry (DQI) is a polynomial-time quantum algorithm introduced by Jordan et al. (Nature 2025). For a natural optimization problem, known as optimal polynomial intersection (OPI), it achieves approximation guarantees in regimes where all known classical algorithms require exponential time. Besides time, space is another central resource: storing and manipulating a massive input can be very challenging, especially when logical qubits carry substantial fault-tolerant implementation overhead. This motivates the following question: does DQI yield quantum advantages in memory, and can we prove it unconditionally? We give an affirmative answer to this question in the streaming setting. In particular, we consider a natural generalization of OPI using Hermite interpolation and Hasse derivatives, which asks for a low-degree polynomial satisfying as many constraints on its values and derivatives as possible. As a concrete example, we show [Quantum efficiency.] An adaptation of the DQI algorithm produces a polynomial satisfying $93\%$ of the constraints; moreover, it only reads the input stream in one pass, uses polylogarithmic space, and has polylogarithmic computation time per stream entry. [Classical hardness.] Any classical algorithm that produces an answer satisfying just $76\%$ of the constraints requires polynomial space, even if it can read the input stream with polynomially many passes and can use unlimited time. Our result provides a complete tradeoff curve for the tunable parameters, and implies that DQI has provable quantum advantages for the original OPI problem.

Authors: Kewen Wu, Guangxu Yang

Decoded quantum interferometry (DQI) is a polynomial-time quantum algorithm introduced by Jordan et al. (Nature 2025). For a natural optimization problem, known as optimal polynomial intersection (OPI), it achieves approximation guarantees in regimes where all known classical algorithms require exponential time. Besides time, space is another central resource: storing and manipulating a massive input can be very challenging, especially when logical qubits carry substantial fault-tolerant implementation overhead. This motivates the following question: does DQI yield quantum advantages in memory, and can we prove it unconditionally? We give an affirmative answer to this question in the streaming setting. In particular, we consider a natural generalization of OPI using Hermite interpolation and Hasse derivatives, which asks for a low-degree polynomial satisfying as many constraints on its values and derivatives as possible. As a concrete example, we show [Quantum efficiency.] An adaptation of the DQI algorithm produces a polynomial satisfying $93\%$ of the constraints; moreover, it only reads the input stream in one pass, uses polylogarithmic space, and has polylogarithmic computation time per stream entry. [Classical hardness.] Any classical algorithm that produces an answer satisfying just $76\%$ of the constraints requires polynomial space, even if it can read the input stream with polynomially many passes and can use unlimited time. Our result provides a complete tradeoff curve for the tunable parameters, and implies that DQI has provable quantum advantages for the original OPI problem.

Stable and Online Algorithms for Random Matrix Discrepancy

from arXiv: Data Structures and Algorithms

Authors: Eren C. Kızıldağ, Shuangping Li

We study the average-case matrix discrepancy problem: given independent normalized $d\times d$ Gaussian orthogonal ensemble matrices $A_1,\dots,A_N$ and a fixed margin $κ>0$, find signs $σ_1,\dots,σ_N\in\{-1,1\}$ such that the operator norm of $\sum_{i=1}^N σ_i A_i$ is at most $κ\sqrt{N}$. Focusing on the proportional regime $N/d^2\to τ\in(0,\infty)$ as $d\to\infty$ followed by the small-margin limit $κ\downarrow 0$, we characterize the density required by stable offline algorithms and by online algorithms. In the offline setting, we construct a polynomial-time \emph{recenter-and-round} algorithm that is noise-stable and succeeds whenever $τ=Ω(\frac{1}{κ^2\log(1/κ)})$, along with a matching lower bound for all stable algorithms. In the online setting where each sign must be chosen irrevocably upon observing the corresponding matrix, we determine the exact limiting performance of the \emph{Frobenius-greedy} algorithm, establishing that it succeeds when $τ>τ_{\rm FG}(κ)\sim \fracπ{4κ^2}$, as well as a matching lower bound for all online algorithms by conditioning on a revealed prefix. At the core of our algorithms lies rotational symmetry, which enables us to transfer Frobenius norm control into operator norm guarantees. Together, our results identify the algorithmic phase transition points for random matrix discrepancy: $Θ(\frac{1}{κ^2\log(1/κ)})$ for stable offline algorithms and $Θ(\frac{1}{κ^2})$ for online algorithms. Both thresholds lie far above the satisfiability scale $Θ(\log(1/κ))$, as shown by Maillard~\cite{maillard2025}.

Authors: Eren C. Kızıldağ, Shuangping Li

We study the average-case matrix discrepancy problem: given independent normalized $d\times d$ Gaussian orthogonal ensemble matrices $A_1,\dots,A_N$ and a fixed margin $κ>0$, find signs $σ_1,\dots,σ_N\in\{-1,1\}$ such that the operator norm of $\sum_{i=1}^N σ_i A_i$ is at most $κ\sqrt{N}$. Focusing on the proportional regime $N/d^2\to τ\in(0,\infty)$ as $d\to\infty$ followed by the small-margin limit $κ\downarrow 0$, we characterize the density required by stable offline algorithms and by online algorithms. In the offline setting, we construct a polynomial-time \emph{recenter-and-round} algorithm that is noise-stable and succeeds whenever $τ=Ω(\frac{1}{κ^2\log(1/κ)})$, along with a matching lower bound for all stable algorithms. In the online setting where each sign must be chosen irrevocably upon observing the corresponding matrix, we determine the exact limiting performance of the \emph{Frobenius-greedy} algorithm, establishing that it succeeds when $τ>τ_{\rm FG}(κ)\sim \fracπ{4κ^2}$, as well as a matching lower bound for all online algorithms by conditioning on a revealed prefix. At the core of our algorithms lies rotational symmetry, which enables us to transfer Frobenius norm control into operator norm guarantees. Together, our results identify the algorithmic phase transition points for random matrix discrepancy: $Θ(\frac{1}{κ^2\log(1/κ)})$ for stable offline algorithms and $Θ(\frac{1}{κ^2})$ for online algorithms. Both thresholds lie far above the satisfiability scale $Θ(\log(1/κ))$, as shown by Maillard~\cite{maillard2025}.

Coloring 3-colorable graphs with $O(n^{4/23})$ colors via a Gaussian-cover recursion

from arXiv: Data Structures and Algorithms

Authors: Emile Anand

We give a randomized polynomial-time algorithm that colors any promised $3$-colorable graph on $n$ vertices with $\smash{O(n^{4/23}) = O(n^{0.17391\ldots})}$ colors, improving on the recent bounds of $O(n^{0.19539})$ by Bansal, Huang, and Lee and Narang and Tang who obtained $O(n^{(13-\sqrt{97})/18+ε})=O(n^{0.17506\dots + ε})$ colors for every fixed $\smash{ε>0}$. To prove our result, we start from a fixed-level semidefinite relaxation, where we use a finite-depth recursion on Gaussian covers. Fixing a root vertex, we group vertices by correlation with the root vector. Here, each step extends a cover of directions by one edge and transfers it to a successor group. Our key analytic ingredient is a variance bound for Gaussian maxima: for a maximum of $m\geq 2$ centered linear forms with coefficient norms at most $r$, mean $μ$, and variance $v$, we prove $v\leq r^2-μ^2/(2\log m)$ using Chen's Gaussian convexity theorem. Together with a variance-scale lower-tail estimate, this controls the threshold loss at each extension, which shows that root-conditioned vector colorings can either extract a large independent set from a group or bound its size, forcing a contradiction after constantly many steps. The resulting sparse-case guarantee combines with the dense progress bound of Kawarabayashi, Thorup, and Yoneda, and the recursion's numerical inequalities are verified via rational interval arithmetic.

Authors: Emile Anand

We give a randomized polynomial-time algorithm that colors any promised $3$-colorable graph on $n$ vertices with $\smash{O(n^{4/23}) = O(n^{0.17391\ldots})}$ colors, improving on the recent bounds of $O(n^{0.19539})$ by Bansal, Huang, and Lee and Narang and Tang who obtained $O(n^{(13-\sqrt{97})/18+ε})=O(n^{0.17506\dots + ε})$ colors for every fixed $\smash{ε>0}$. To prove our result, we start from a fixed-level semidefinite relaxation, where we use a finite-depth recursion on Gaussian covers. Fixing a root vertex, we group vertices by correlation with the root vector. Here, each step extends a cover of directions by one edge and transfers it to a successor group. Our key analytic ingredient is a variance bound for Gaussian maxima: for a maximum of $m\geq 2$ centered linear forms with coefficient norms at most $r$, mean $μ$, and variance $v$, we prove $v\leq r^2-μ^2/(2\log m)$ using Chen's Gaussian convexity theorem. Together with a variance-scale lower-tail estimate, this controls the threshold loss at each extension, which shows that root-conditioned vector colorings can either extract a large independent set from a group or bound its size, forcing a contradiction after constantly many steps. The resulting sparse-case guarantee combines with the dense progress bound of Kawarabayashi, Thorup, and Yoneda, and the recursion's numerical inequalities are verified via rational interval arithmetic.

Linear Programming Representations and Strongly Polynomial Algorithms for Robust Markov Decision Processes

from arXiv: Data Structures and Algorithms

Authors: Han Zhong, Yinyu Ye

We study linear programming (LP) representations and strongly polynomial algorithms for robust Markov decision processes (RMDPs) with rational polyhedral state-action rectangular uncertainty in rewards and transitions. By encoding a finite sequence of robust policy-iteration steps, we construct a single LP whose optimal solutions recover the robust optimal value and all optimal stationary randomized policies. At fixed discount, the LP has polynomial dimension and encoding length and can be constructed in strongly polynomial time. We also develop a general complexity analysis of robust policy iteration that combines the cost of minimizing over uncertainty sets with the number of iterations needed to evaluate a policy. For a fixed discount factor, we use this analysis to improve the known complexity bounds for $\ell_1$ and $\ell_\infty$ RMDPs and establish new strongly polynomial bounds for general interval, weighted $\ell_1$, and Wasserstein RMDPs, as well as turn-based stochastic games with these uncertainty sets.

Authors: Han Zhong, Yinyu Ye

We study linear programming (LP) representations and strongly polynomial algorithms for robust Markov decision processes (RMDPs) with rational polyhedral state-action rectangular uncertainty in rewards and transitions. By encoding a finite sequence of robust policy-iteration steps, we construct a single LP whose optimal solutions recover the robust optimal value and all optimal stationary randomized policies. At fixed discount, the LP has polynomial dimension and encoding length and can be constructed in strongly polynomial time. We also develop a general complexity analysis of robust policy iteration that combines the cost of minimizing over uncertainty sets with the number of iterations needed to evaluate a policy. For a fixed discount factor, we use this analysis to improve the known complexity bounds for $\ell_1$ and $\ell_\infty$ RMDPs and establish new strongly polynomial bounds for general interval, weighted $\ell_1$, and Wasserstein RMDPs, as well as turn-based stochastic games with these uncertainty sets.

Randomized Matvec Lower Bounds for Simplex-Based Matrix Games

from arXiv: Data Structures and Algorithms

Authors: Wendao Wu, Cong Fang

We prove randomized matrix-vector query lower bounds for two normalized matrix-game geometries: a Euclidean unit ball against a probability simplex, with row norms at most one, and two probability simplices, with entries of absolute value at most one. Each query returns $(Ax,A^\top y)$ for arbitrary real vectors. The algorithm must return a feasible pair with full saddle-point gap at most $\varepsilon$, with probability at least $2/3$ on every admissible matrix. For sufficiently small $\varepsilon$, the worst-case query complexities are $Ω(\varepsilon^{-2/3}/(\log^2(1/\varepsilon)\log\log(1/\varepsilon)))$ for ball-simplex games and $Ω(\varepsilon^{-2/3}/(\log^{7/3}(1/\varepsilon)\log\log(1/\varepsilon)))$ for simplex-simplex games. The hard instances have dimensions of order $\varepsilon^{-2/3}$ and $\varepsilon^{-2/3}/\log^{1/3}(1/\varepsilon)$, respectively, and the bounds extend to larger dimensions. These lower bounds match the deterministic upper bounds of Karmarkar, O'Carroll, and Sidford up to logarithmic factors. The proof extracts a fresh Gaussian core after adaptive two-sided queries and uses uncertainty in its smallest singular value to establish linear-system solve hardness. Two reductions transfer this hardness to matrix games by converting a small full gap into a small residual, with an additional logarithmic normalization loss only for simplex-simplex games.

Authors: Wendao Wu, Cong Fang

We prove randomized matrix-vector query lower bounds for two normalized matrix-game geometries: a Euclidean unit ball against a probability simplex, with row norms at most one, and two probability simplices, with entries of absolute value at most one. Each query returns $(Ax,A^\top y)$ for arbitrary real vectors. The algorithm must return a feasible pair with full saddle-point gap at most $\varepsilon$, with probability at least $2/3$ on every admissible matrix. For sufficiently small $\varepsilon$, the worst-case query complexities are $Ω(\varepsilon^{-2/3}/(\log^2(1/\varepsilon)\log\log(1/\varepsilon)))$ for ball-simplex games and $Ω(\varepsilon^{-2/3}/(\log^{7/3}(1/\varepsilon)\log\log(1/\varepsilon)))$ for simplex-simplex games. The hard instances have dimensions of order $\varepsilon^{-2/3}$ and $\varepsilon^{-2/3}/\log^{1/3}(1/\varepsilon)$, respectively, and the bounds extend to larger dimensions. These lower bounds match the deterministic upper bounds of Karmarkar, O'Carroll, and Sidford up to logarithmic factors. The proof extracts a fresh Gaussian core after adaptive two-sided queries and uses uncertainty in its smallest singular value to establish linear-system solve hardness. Two reductions transfer this hardness to matrix games by converting a small full gap into a small residual, with an additional logarithmic normalization loss only for simplex-simplex games.

Quantum state preparation for weighted d-DNNF

from arXiv: Data Structures and Algorithms

Authors: Steef Hegeman, Joon Hyung Lee, Alfons Laarman

The quantum state preparation problem is to, given a description of a quantum state, efficiently generate a quantum circuit computing the state. We show that for quantum states described by weighted d-DNNF (deterministic, decomposable pseudo-Boolean circuits) a quantum circuit computing the state can be obtained in linear time up to complex arithmetic.

Authors: Steef Hegeman, Joon Hyung Lee, Alfons Laarman

The quantum state preparation problem is to, given a description of a quantum state, efficiently generate a quantum circuit computing the state. We show that for quantum states described by weighted d-DNNF (deterministic, decomposable pseudo-Boolean circuits) a quantum circuit computing the state can be obtained in linear time up to complex arithmetic.

A computational phase diagram for the transverse field Ising model

from arXiv: Data Structures and Algorithms

Authors: Thuy-Duong Vuong

We study the transverse field Ising model, defined by the Hamiltonian $H =\frac{1}{2}\sum_{i, j\in [n]} J_{ij} Z_i Z_j +\sum_{i=1}^n h_i^z Z_i + η\sum_{i} X_i$ where $J $ is the symmetric interaction matrix, and $η$ is the transverse field strength. Let $Δ(J)=λ_{\max}(J)-λ_{\min}(J)$ be the spectral width of $J.$ When the inverse temperature $β\geq0$ satisfies $Δ(J)\cdot\frac{\tanh(βη)}η\leq1$, we give a randomized classical algorithm that approximates the partition function $Z(β)=\operatorname{Tr}(e^{-βH})$ to a given relative error $ε\in(0,1)$ in time polynomial in $n$, $β$, the model parameters, and $ε^{-1}$. When $ Δ(J) \cdot \frac{\tanh(βη)}η > 1 ,$ we show that approximating $ Z(β)$ within an $\exp(o(n))$-multiplicative factor is $\textbf{NP}$-hard, and thus unlikely to admit an efficient classical or quantum algorithms under standard complexity theoretic assumptions. Furthermore, in the regime $Δ(J)\cdot \frac{\tanh(βη)}η\leq 1,$ we provide an efficient randomized classical algorithm that approximates Pauli string observables of the Gibbs state $ ρ_β= \frac{e^{-βH}}{\operatorname{Tr}(e^{-βH})}$ within an arbitrarily small additive error. In the special case when the observable is also diagonal in the $X$-basis, i.e. $P \in \{I, X\}^{\otimes n}$, the algorithm further achieves arbitrarily small relative error.

Authors: Thuy-Duong Vuong

We study the transverse field Ising model, defined by the Hamiltonian $H =\frac{1}{2}\sum_{i, j\in [n]} J_{ij} Z_i Z_j +\sum_{i=1}^n h_i^z Z_i + η\sum_{i} X_i$ where $J $ is the symmetric interaction matrix, and $η$ is the transverse field strength. Let $Δ(J)=λ_{\max}(J)-λ_{\min}(J)$ be the spectral width of $J.$ When the inverse temperature $β\geq0$ satisfies $Δ(J)\cdot\frac{\tanh(βη)}η\leq1$, we give a randomized classical algorithm that approximates the partition function $Z(β)=\operatorname{Tr}(e^{-βH})$ to a given relative error $ε\in(0,1)$ in time polynomial in $n$, $β$, the model parameters, and $ε^{-1}$. When $ Δ(J) \cdot \frac{\tanh(βη)}η > 1 ,$ we show that approximating $ Z(β)$ within an $\exp(o(n))$-multiplicative factor is $\textbf{NP}$-hard, and thus unlikely to admit an efficient classical or quantum algorithms under standard complexity theoretic assumptions. Furthermore, in the regime $Δ(J)\cdot \frac{\tanh(βη)}η\leq 1,$ we provide an efficient randomized classical algorithm that approximates Pauli string observables of the Gibbs state $ ρ_β= \frac{e^{-βH}}{\operatorname{Tr}(e^{-βH})}$ within an arbitrarily small additive error. In the special case when the observable is also diagonal in the $X$-basis, i.e. $P \in \{I, X\}^{\otimes n}$, the algorithm further achieves arbitrarily small relative error.

Vertex-Failure Distance Oracles and Labeling Schemes: Compact and Constant-Approximate

from arXiv: Data Structures and Algorithms

Authors: Yaowei Long

We present new algorithms for the vertex-failure distance oracles and labeling schemes problems in undirected weighted graphs. A vertex-failure distance oracle is a data structure that, given two vertices $x$ and $y$ and a failed vertex set $F$ of size at most $f$, returns an approximation to the distance between $x$ and $y$ in $G \setminus F$. In the labeling-scheme setting, the data structure needs to be stored distributively as labels on the vertices, and each query $(x,y,F)$ must be answered by accessing only the labels of the vertices in $F \cup \{x,y\}$. For any $f\geq 1$ and $k \ge 1$, we obtain a vertex-failure distance oracle with $O(k^{6})$ approximation, space $\tilde{O}(f^{2}n^{1+1/k})$, query time $\tilde{O}(f^{5}n^{1/k})$, and polynomial preprocessing time. In particular, this is the first time-efficient oracle for multiple vertex failures with space close to linear, as well as the first constant-approximation oracle with polynomial space when tolerating $Ω(\log n)$ vertex failures. The previous results, due to [Duan-Gu-Ren, SODA'21], gave two alternatives: for any constant $c \ge 1$ and $ε>0$, one oracle has $\mathrm{poly}(\log n,f)$ approximation, space $n^{2+1/c}\mathrm{poly}(\log n,f)$, and query time $\mathrm{poly}(\log n,f^{c})$, while the other has $(1+ε)$ approximation, space $n^{2+1/c}(\log n/ε)^{O(f)}$, and query time $\mathrm{poly}(\log n,f^{c},1/ε)$. We also obtain a vertex-failure distance labeling scheme with $O(k^{6})$ approximation and label size $f^{3}n^{1/k}\log^{O(k)} n$. This is the first nontrivial distance labeling scheme for vertex failures. Our techniques build on recent tools related to length-constrained vertex expanders and also introduce a new expander-based shortcut sparsification. The latter also leads to a deterministic vertex-failure connectivity labeling scheme of size $\tilde{O}(f^{2})$.

Authors: Yaowei Long

We present new algorithms for the vertex-failure distance oracles and labeling schemes problems in undirected weighted graphs. A vertex-failure distance oracle is a data structure that, given two vertices $x$ and $y$ and a failed vertex set $F$ of size at most $f$, returns an approximation to the distance between $x$ and $y$ in $G \setminus F$. In the labeling-scheme setting, the data structure needs to be stored distributively as labels on the vertices, and each query $(x,y,F)$ must be answered by accessing only the labels of the vertices in $F \cup \{x,y\}$. For any $f\geq 1$ and $k \ge 1$, we obtain a vertex-failure distance oracle with $O(k^{6})$ approximation, space $\tilde{O}(f^{2}n^{1+1/k})$, query time $\tilde{O}(f^{5}n^{1/k})$, and polynomial preprocessing time. In particular, this is the first time-efficient oracle for multiple vertex failures with space close to linear, as well as the first constant-approximation oracle with polynomial space when tolerating $Ω(\log n)$ vertex failures. The previous results, due to [Duan-Gu-Ren, SODA'21], gave two alternatives: for any constant $c \ge 1$ and $ε>0$, one oracle has $\mathrm{poly}(\log n,f)$ approximation, space $n^{2+1/c}\mathrm{poly}(\log n,f)$, and query time $\mathrm{poly}(\log n,f^{c})$, while the other has $(1+ε)$ approximation, space $n^{2+1/c}(\log n/ε)^{O(f)}$, and query time $\mathrm{poly}(\log n,f^{c},1/ε)$. We also obtain a vertex-failure distance labeling scheme with $O(k^{6})$ approximation and label size $f^{3}n^{1/k}\log^{O(k)} n$. This is the first nontrivial distance labeling scheme for vertex failures. Our techniques build on recent tools related to length-constrained vertex expanders and also introduce a new expander-based shortcut sparsification. The latter also leads to a deterministic vertex-failure connectivity labeling scheme of size $\tilde{O}(f^{2})$.

Convergence of Kikuchi matrices to $Γ$-independent and $q$-Gaussian limits

from arXiv: Data Structures and Algorithms

Authors: Afonso S. Bandeira, Dmitriy Kunisky, Petar Nizić-Nikolac, Lucas Pesenti, Robert Wang

Kikuchi matrices are a family of structured matrices that were introduced to study problems involving tensors and hypergraphs. We show that, as the ambient dimension grows, dense random Kikuchi matrices have a limit described by a system of $Γ$-independent semicircular elements. This characterizes their limiting spectral distribution and yields improved bounds on their spectral norm, a key quantity in the analysis of algorithms for Tensor PCA. Finally, we show that, in an appropriate double limit, independent Kikuchi matrices converge to the $q$-Gaussian system, another central object in noncommutative probability.

Authors: Afonso S. Bandeira, Dmitriy Kunisky, Petar Nizić-Nikolac, Lucas Pesenti, Robert Wang

Kikuchi matrices are a family of structured matrices that were introduced to study problems involving tensors and hypergraphs. We show that, as the ambient dimension grows, dense random Kikuchi matrices have a limit described by a system of $Γ$-independent semicircular elements. This characterizes their limiting spectral distribution and yields improved bounds on their spectral norm, a key quantity in the analysis of algorithms for Tensor PCA. Finally, we show that, in an appropriate double limit, independent Kikuchi matrices converge to the $q$-Gaussian system, another central object in noncommutative probability.

Beating One Half for Online Bipartite Matching with Reusable Resources

from arXiv: Data Structures and Algorithms

Authors: Xiaohui Bei, Zhihao Gavin Tang, Wenhao Wu

We study online bipartite matching with unit-inventory reusable resources, where requests arrive in an adversarially fixed order, and each use of a resource makes it unavailable for an independent duration drawn from a resource-dependent distribution. The benchmark knows all requests in advance but cannot observe a duration before choosing the corresponding use. The classical Ranking algorithm of Karp, Vazirani, and Vazirani (STOC 1990) fixes a uniformly random priority order of the resources and matches each arriving request to its highest-priority available neighbor. It achieves the optimal competitive ratio $1-1/e$ for unweighted nonreusable resources, but whether it beats $1/2$ for reusable resources has remained open. We prove that, for unweighted resources with resource-dependent stochastic durations, Ranking achieves a competitive ratio of $(5-2\sqrt3)/3\approx0.511966$. We also give a black-box reduction from unweighted Ranking to resource-weighted matching: any unweighted competitive ratio $α>1/2$ yields a weighted ratio strictly above $1/2$. With independent sampling access to the duration distributions, the reduction gives a weighted ratio of $0.500034$. These results resolve two questions left open by Delong et al. (MOR 2024): whether Ranking beats $1/2$, and whether one can beat $1/2$ under stochastic durations. We analyze Ranking resource by resource, rather than request by request. For deterministic durations, this gives a reduction to random-order greedy for a coverage function. We then extend the analysis to stochastic durations by comparing the residual schedules of Ranking and a greedy algorithm, and apply a finer analysis of the random ranks to obtain the stated $0.511$ bound. For the weighted reduction, we apply Ranking within groups of similar weights and uses weighted greedy to control the loss between groups.

Authors: Xiaohui Bei, Zhihao Gavin Tang, Wenhao Wu

We study online bipartite matching with unit-inventory reusable resources, where requests arrive in an adversarially fixed order, and each use of a resource makes it unavailable for an independent duration drawn from a resource-dependent distribution. The benchmark knows all requests in advance but cannot observe a duration before choosing the corresponding use. The classical Ranking algorithm of Karp, Vazirani, and Vazirani (STOC 1990) fixes a uniformly random priority order of the resources and matches each arriving request to its highest-priority available neighbor. It achieves the optimal competitive ratio $1-1/e$ for unweighted nonreusable resources, but whether it beats $1/2$ for reusable resources has remained open. We prove that, for unweighted resources with resource-dependent stochastic durations, Ranking achieves a competitive ratio of $(5-2\sqrt3)/3\approx0.511966$. We also give a black-box reduction from unweighted Ranking to resource-weighted matching: any unweighted competitive ratio $α>1/2$ yields a weighted ratio strictly above $1/2$. With independent sampling access to the duration distributions, the reduction gives a weighted ratio of $0.500034$. These results resolve two questions left open by Delong et al. (MOR 2024): whether Ranking beats $1/2$, and whether one can beat $1/2$ under stochastic durations. We analyze Ranking resource by resource, rather than request by request. For deterministic durations, this gives a reduction to random-order greedy for a coverage function. We then extend the analysis to stochastic durations by comparing the residual schedules of Ranking and a greedy algorithm, and apply a finer analysis of the random ranks to obtain the stated $0.511$ bound. For the weighted reduction, we apply Ranking within groups of similar weights and uses weighted greedy to control the loss between groups.

Achieving Optimal Redundancy for Small Dynamic Rank/Select Dictionaries

from arXiv: Data Structures and Algorithms

Authors: Gabriel Marques Domingues

In this paper, we study the number of bits required to construct a dynamic dictionary with optimal time for $\texttt{rank}/\texttt{select}$ operations. Using the standard (multiplication) Word-RAM model with $w$-bit words, we construct a data-structure for a dynamic $\texttt{rank}/\texttt{select}$ dictionary for a set $S\subseteq\{0,1,\cdots,u-1\}$ of $n$ elements that, given a parameter $1\leq k\leq \log^*w$, uses $$\operatorname{lg}\binom{u}{n}+\mathcal{O}(n\log^{(k)}w)\text{ bits}$$ taking optimal $\mathcal{O}(k+\log_w n)$ time (worst-case) for all operations. We show optimality for $n=w^{\mathcal{O}(1)}$ by extending the lower bound of Li, Liang, Yu, and Zhou [FOCS 2023] to super-polynomial universes: any dynamic dictionary for $n\leq \sqrt{u}$ elements that uses $\operatorname{lg}\binom{u}{n}+\mathcal{O}(n\log^{(k)}n)$ bits requires $Ω(k)$ time for operations. Lastly, we extend the data-structure to a dynamic fully indexable dictionary (that also supports $\texttt{rank}/\texttt{select}$ on the complement of $S$).

Authors: Gabriel Marques Domingues

In this paper, we study the number of bits required to construct a dynamic dictionary with optimal time for $\texttt{rank}/\texttt{select}$ operations. Using the standard (multiplication) Word-RAM model with $w$-bit words, we construct a data-structure for a dynamic $\texttt{rank}/\texttt{select}$ dictionary for a set $S\subseteq\{0,1,\cdots,u-1\}$ of $n$ elements that, given a parameter $1\leq k\leq \log^*w$, uses $$\operatorname{lg}\binom{u}{n}+\mathcal{O}(n\log^{(k)}w)\text{ bits}$$ taking optimal $\mathcal{O}(k+\log_w n)$ time (worst-case) for all operations. We show optimality for $n=w^{\mathcal{O}(1)}$ by extending the lower bound of Li, Liang, Yu, and Zhou [FOCS 2023] to super-polynomial universes: any dynamic dictionary for $n\leq \sqrt{u}$ elements that uses $\operatorname{lg}\binom{u}{n}+\mathcal{O}(n\log^{(k)}n)$ bits requires $Ω(k)$ time for operations. Lastly, we extend the data-structure to a dynamic fully indexable dictionary (that also supports $\texttt{rank}/\texttt{select}$ on the complement of $S$).

Near-optimal quantum query lower bounds on bipartiteness and expansion testing in the bounded-degree graph model

from arXiv: Data Structures and Algorithms

Authors: Chandrima Kayal, Sayantan Sen, Dániel Szabó

In this work, we study bipartiteness and expansion testing, two canonical problems in graph property testing in the bounded-degree model through the lens of quantum query complexity. In the classical setting, it is known that $\widetildeΘ(\sqrt{N})$ queries are necessary and sufficient for both these testing problems (Goldreich and Ron, 1999, 2000 & 2002), where $N$ denotes the number of vertices of the input graph. Due to their significance, (Ambainis, Childs, and Liu, 2011) initiated the study of these problems in the quantum setting and designed quantum algorithms for bipartiteness and expansion testing that perform $\widetilde{O}(N^{1/3})$ queries, showing a polynomial speedup. They also proved that $\widetildeΩ(N^{1/4})$ queries are necessary for expansion testing, but the possibility of an exponential quantum advantage for bipartiteness testing remained open. Despite significant effort, there has been no improvement in these results in the last decade and a half. In this work, we prove essentially tight $\widetildeΩ(N^{1/3})$ quantum query lower bounds for both bipartiteness and expansion testing, thereby completely characterizing the quantum query complexity of these problems up to polylogarithmic factors. While our proofs use the polynomial method similarly to Ambainis, Childs, and Liu, we use intermediate problems that we relate to the main problems via reductions, and perform a more precise analysis of the resulting polynomials, leading to the near-optimal lower bounds.

Authors: Chandrima Kayal, Sayantan Sen, Dániel Szabó

In this work, we study bipartiteness and expansion testing, two canonical problems in graph property testing in the bounded-degree model through the lens of quantum query complexity. In the classical setting, it is known that $\widetildeΘ(\sqrt{N})$ queries are necessary and sufficient for both these testing problems (Goldreich and Ron, 1999, 2000 & 2002), where $N$ denotes the number of vertices of the input graph. Due to their significance, (Ambainis, Childs, and Liu, 2011) initiated the study of these problems in the quantum setting and designed quantum algorithms for bipartiteness and expansion testing that perform $\widetilde{O}(N^{1/3})$ queries, showing a polynomial speedup. They also proved that $\widetildeΩ(N^{1/4})$ queries are necessary for expansion testing, but the possibility of an exponential quantum advantage for bipartiteness testing remained open. Despite significant effort, there has been no improvement in these results in the last decade and a half. In this work, we prove essentially tight $\widetildeΩ(N^{1/3})$ quantum query lower bounds for both bipartiteness and expansion testing, thereby completely characterizing the quantum query complexity of these problems up to polylogarithmic factors. While our proofs use the polynomial method similarly to Ambainis, Childs, and Liu, we use intermediate problems that we relate to the main problems via reductions, and perform a more precise analysis of the resulting polynomials, leading to the near-optimal lower bounds.

Safe Hypergraph Contraction via Capacity-Aware Repair Certificates

from arXiv: Data Structures and Algorithms

Authors: Yu Deng, Xinyi Yang, Keren Zhu

Multilevel partitioners shrink circuit hypergraphs through vertex contractions, yet a contraction that satisfies block capacity can still eliminate every optimal balanced bipartition. We develop certified safe coarsening (CSC) to identify contractions that preserve an optimum without computing that optimum. CSC certifies a repair for any feasible partition that splits a candidate group: the repair must respect the fixed block capacities and must not increase the cut-net objective. Its bounds exclude hyperedges that capacity constraints force to be cut. A pair certificate checks individual merges, while a directed minimum-cut test certifies groups whose savings emerge only when vertices move together. We prove that certified disjoint batches and successive rounds with recertification retain at least one globally optimal feasible partition for hypergraphs with positive integer vertex and net weights. Experiments on exactly solvable instances confirm optimum preservation for every tested configuration; integration with KaHyPar lowers the sum of per-instance best cuts on circuit benchmarks, with additional runtime.

Authors: Yu Deng, Xinyi Yang, Keren Zhu

Multilevel partitioners shrink circuit hypergraphs through vertex contractions, yet a contraction that satisfies block capacity can still eliminate every optimal balanced bipartition. We develop certified safe coarsening (CSC) to identify contractions that preserve an optimum without computing that optimum. CSC certifies a repair for any feasible partition that splits a candidate group: the repair must respect the fixed block capacities and must not increase the cut-net objective. Its bounds exclude hyperedges that capacity constraints force to be cut. A pair certificate checks individual merges, while a directed minimum-cut test certifies groups whose savings emerge only when vertices move together. We prove that certified disjoint batches and successive rounds with recertification retain at least one globally optimal feasible partition for hypergraphs with positive integer vertex and net weights. Experiments on exactly solvable instances confirm optimum preservation for every tested configuration; integration with KaHyPar lowers the sum of per-instance best cuts on circuit benchmarks, with additional runtime.

Exact Locality Gaps for Matchable Semi-Matchings

from arXiv: Data Structures and Algorithms

Authors: Marek Gałązka, Hanna Wdowicka

An assignment of tasks to servers can resist every small improvement and still make tasks wait longer than necessary. We determine exactly how inefficient such an assignment can be when each task requires one unit of service and the eligibility constraints permit all tasks to use distinct servers. For every move size $r$ and maximum current server load $K$, we give a closed formula for the worst ratio between locally optimal and globally optimal total completion time. Local optimality here allows every feasible reassignment changing at most $r$ tasks. Every finite-cap bound is attained on a tree where each task has at most two eligible servers. Thus the worst behavior already occurs under simple eligibility constraints. At load cap two, the exact ratio is $1+1/(r+2)$, attained on a path with $r+2$ tasks. Without a load cap, the worst-case supremum is $3/2$ for single-task moves and approximately $1.294503159$ for two-task moves; its excess above one is $1/(r+2)+O(2^{-r}/r)$ as $r$ grows. The proof uses an explicit rational potential on a comparison graph and matching extremal constructions. These results give sharp guarantees for bounded-size local search on matchable semi-matchings, including exact guarantees under degree bounds.

Authors: Marek Gałązka, Hanna Wdowicka

An assignment of tasks to servers can resist every small improvement and still make tasks wait longer than necessary. We determine exactly how inefficient such an assignment can be when each task requires one unit of service and the eligibility constraints permit all tasks to use distinct servers. For every move size $r$ and maximum current server load $K$, we give a closed formula for the worst ratio between locally optimal and globally optimal total completion time. Local optimality here allows every feasible reassignment changing at most $r$ tasks. Every finite-cap bound is attained on a tree where each task has at most two eligible servers. Thus the worst behavior already occurs under simple eligibility constraints. At load cap two, the exact ratio is $1+1/(r+2)$, attained on a path with $r+2$ tasks. Without a load cap, the worst-case supremum is $3/2$ for single-task moves and approximately $1.294503159$ for two-task moves; its excess above one is $1/(r+2)+O(2^{-r}/r)$ as $r$ grows. The proof uses an explicit rational potential on a comparison graph and matching extremal constructions. These results give sharp guarantees for bounded-size local search on matchable semi-matchings, including exact guarantees under degree bounds.

Robust Non-Clairvoyant Scheduling with Classification Models

from arXiv: Data Structures and Algorithms

Authors: Anthony Dugois, Vincent Fagnon, Giorgio Lucarelli

We study the classical single-machine scheduling problem of minimizing the sum of completion times of jobs in a non-clairvoyant setting, where the processing time of each job remains unknown until its completion. This is a hard problem for which no constant competitive algorithm is possible. Inspired by robust optimization and learning-augmented algorithms, we introduce a novel robustness framework that leverages structural information provided by a classification model to overcome this limitation. Specifically, we assume that jobs are partitioned into classes and we have access to the confusion matrix of the classifier, whose entry $(k,\ell)$ indicates the number of jobs predicted to belong to class~$k$ but that actually belong to class~$\ell$. In this manner, we are able to characterize uncertainty as a set of permutations within each predicted class, rather than as a collection of discrete numerical scenarios, avoiding the computational difficulty of classical robust metrics, such as Min-Max and Min-Max Regret. In addition to these worst-case metrics, we also consider the expected objective over all scenarios. We first propose an optimal non-adaptive strategy that is oblivious with respect to all three robust criteria. We then investigate adaptive and randomized algorithms, showing that they can outperform the optimal non-adaptive strategy when the matrix exhibits particular structural properties.

Authors: Anthony Dugois, Vincent Fagnon, Giorgio Lucarelli

We study the classical single-machine scheduling problem of minimizing the sum of completion times of jobs in a non-clairvoyant setting, where the processing time of each job remains unknown until its completion. This is a hard problem for which no constant competitive algorithm is possible. Inspired by robust optimization and learning-augmented algorithms, we introduce a novel robustness framework that leverages structural information provided by a classification model to overcome this limitation. Specifically, we assume that jobs are partitioned into classes and we have access to the confusion matrix of the classifier, whose entry $(k,\ell)$ indicates the number of jobs predicted to belong to class~$k$ but that actually belong to class~$\ell$. In this manner, we are able to characterize uncertainty as a set of permutations within each predicted class, rather than as a collection of discrete numerical scenarios, avoiding the computational difficulty of classical robust metrics, such as Min-Max and Min-Max Regret. In addition to these worst-case metrics, we also consider the expected objective over all scenarios. We first propose an optimal non-adaptive strategy that is oblivious with respect to all three robust criteria. We then investigate adaptive and randomized algorithms, showing that they can outperform the optimal non-adaptive strategy when the matrix exhibits particular structural properties.

Factor Three Approximation for Edit Distance

from arXiv: Data Structures and Algorithms

Authors: Egor Gorbachev

We give randomized algorithms for $3$-approximate edit distance in $\widetilde{\mathcal{O}}(N^{11/6})$ time for unweighted edit distance and in $\widetilde{\mathcal{O}}(N^{40/21})$ time for arbitrary metric edit weights, where $N$ is the total input length. For non-metric costs, we prove an unconditional $Ω(N^2)$ oracle-query lower bound for every approximation factor depending only on $N$, even for symmetric weights or weights satisfying the triangle inequality (but not both). Under the Orthogonal Vectors Hypothesis, we show a similar result for constant-size alphabets. This holds even for symmetric weights over a size-$3$ alphabet or triangle-inequality weights over a size-$2$ alphabet. In contrast, for symmetric weights over a binary alphabet we show an $\widetilde{\mathcal{O}}(N^{40/21})$-time $3$-approximation algorithm.

Authors: Egor Gorbachev

We give randomized algorithms for $3$-approximate edit distance in $\widetilde{\mathcal{O}}(N^{11/6})$ time for unweighted edit distance and in $\widetilde{\mathcal{O}}(N^{40/21})$ time for arbitrary metric edit weights, where $N$ is the total input length. For non-metric costs, we prove an unconditional $Ω(N^2)$ oracle-query lower bound for every approximation factor depending only on $N$, even for symmetric weights or weights satisfying the triangle inequality (but not both). Under the Orthogonal Vectors Hypothesis, we show a similar result for constant-size alphabets. This holds even for symmetric weights over a size-$3$ alphabet or triangle-inequality weights over a size-$2$ alphabet. In contrast, for symmetric weights over a binary alphabet we show an $\widetilde{\mathcal{O}}(N^{40/21})$-time $3$-approximation algorithm.

When Is Deletion Ordering Tractable? From Update Dynamics to Permutation Structure

from arXiv: Data Structures and Algorithms

Authors: Xinyu Wang, Ziyu Zhao, Yixuan He, Xiaowen Chang Alex Smola

Given a fixed set of pending deletion requests, retraining from scratch after each request is prohibitive, so a prescribed request-wise policy processes them sequentially. The resulting terminal model can depend on their order. Rather than prescribing an ordering rule, we study the permutation objective induced by the fixed policy and ask when it admits simpler structure. We identify two independent reductions: position additivity represents the objective by request--position costs, reducing optimization to assignment and, with a shared positional profile, sorting; suffix localization removes dependence on the distant prefix while retaining interactions among the surviving requests. Under shared affine updates, we characterize the quadratic interactions that obstruct additivity, prove the reductions' independence, and show that suffix-conditioned assignment improves the approximation rate from O(p^L) toO(p^(2L)). Experiments recover both structures in executed objectives. A controlled damped-Newton sweep shows that stronger contraction shifts the objective toward shorter, more suffix-specific dependence, while two full-network policies exhibit distinct positional and within-suffix structure. Structures identified from compact execution sets also predict unseen orders. These results frame deletion ordering as identifying the computational structure induced by the executed updates.

Authors: Xinyu Wang, Ziyu Zhao, Yixuan He, Xiaowen Chang Alex Smola

Given a fixed set of pending deletion requests, retraining from scratch after each request is prohibitive, so a prescribed request-wise policy processes them sequentially. The resulting terminal model can depend on their order. Rather than prescribing an ordering rule, we study the permutation objective induced by the fixed policy and ask when it admits simpler structure. We identify two independent reductions: position additivity represents the objective by request--position costs, reducing optimization to assignment and, with a shared positional profile, sorting; suffix localization removes dependence on the distant prefix while retaining interactions among the surviving requests. Under shared affine updates, we characterize the quadratic interactions that obstruct additivity, prove the reductions' independence, and show that suffix-conditioned assignment improves the approximation rate from O(p^L) toO(p^(2L)). Experiments recover both structures in executed objectives. A controlled damped-Newton sweep shows that stronger contraction shifts the objective toward shorter, more suffix-specific dependence, while two full-network policies exhibit distinct positional and within-suffix structure. Structures identified from compact execution sets also predict unseen orders. These results frame deletion ordering as identifying the computational structure induced by the executed updates.

Settling the Pass Complexity of Streaming Set Cover

from arXiv: Data Structures and Algorithms

Authors: Sepehr Assadi, Janani Sundaresan

In the streaming set cover problem, $m$ sets from a universe of size $n$ are arriving one by one in a stream, and the algorithm is allowed to process the stream using one or a few passes and a space of $o(mn)$, which is sublinear in the input size. The goal is to determine the minimal (or approximately minimal) number of sets that cover the universe at the end of the last pass. This problem has been studied extensively over the years with rapid progress that led to several $O(\log{n})$-approximation algorithms in $\tilde{O}(mn^{1/p})$ space and $O(p)$ passes. However, progress on this front has largely stagnated over the past decade, despite the absence of any lower bounds that rule out even an $O(\log{n})$-approximation in $O(m)$ space and just two passes. We provide a simple explanation for this lack of progress by establishing an optimal three-way space-pass-approximation tradeoff for this problem: any $α$-approximation algorithm for streaming set cover requires $$ \widetildeΩ\Big(\frac{m}α \cdot \big(\frac{n}α\big)^{1/p}\Big) $$ space in $p$ passes whenever $α\ll n^{1/(p+1)}$. In light of prior work, this result is optimal up to constant factors in $p$ and logarithmic factors in $n,m$ for any $α\geq p$. Our bound is optimal with respect to the range of $α$ also, and fully settles the complexity of this fundamental problem in the streaming model. The proof of this result is (surprisingly) simple and non-technical and relies on a randomized reduction from a variant of the standard pointer chasing problem in communication complexity, using elementary properties of random sets.

Authors: Sepehr Assadi, Janani Sundaresan

In the streaming set cover problem, $m$ sets from a universe of size $n$ are arriving one by one in a stream, and the algorithm is allowed to process the stream using one or a few passes and a space of $o(mn)$, which is sublinear in the input size. The goal is to determine the minimal (or approximately minimal) number of sets that cover the universe at the end of the last pass. This problem has been studied extensively over the years with rapid progress that led to several $O(\log{n})$-approximation algorithms in $\tilde{O}(mn^{1/p})$ space and $O(p)$ passes. However, progress on this front has largely stagnated over the past decade, despite the absence of any lower bounds that rule out even an $O(\log{n})$-approximation in $O(m)$ space and just two passes. We provide a simple explanation for this lack of progress by establishing an optimal three-way space-pass-approximation tradeoff for this problem: any $α$-approximation algorithm for streaming set cover requires $$ \widetildeΩ\Big(\frac{m}α \cdot \big(\frac{n}α\big)^{1/p}\Big) $$ space in $p$ passes whenever $α\ll n^{1/(p+1)}$. In light of prior work, this result is optimal up to constant factors in $p$ and logarithmic factors in $n,m$ for any $α\geq p$. Our bound is optimal with respect to the range of $α$ also, and fully settles the complexity of this fundamental problem in the streaming model. The proof of this result is (surprisingly) simple and non-technical and relies on a randomized reduction from a variant of the standard pointer chasing problem in communication complexity, using elementary properties of random sets.

Best of Two Worlds: Combining High and Low Resolution to Compute Viewsheds on terrains

from arXiv: Data Structures and Algorithms

Authors: Laura Toma

The viewshed of a point $v$ on a grid terrain $T$, viewshed$_T(v)$, is defined as the set of grid points in $T$ that are visible from $v$. We describe a novel algorithm for computing viewshed$_T(v)$ using a multi-resolution approach: Given a parameter $k >1$ that represents the block size, we create a grid $T'$ which is a lower-resolution version of $T$, such that each point in $T'$ corresponds to a block of $\lceil \sqrt k \rceil $-by-$\lceil \sqrt k \rceil$ points in $T$. The key of our approach is using $T'$ to speed up the computation of viewshed$_T(v)$ while not introducing approximation. We compute viewshed$_T(v)$ in two steps: First we compute the viewshed of $v$ on $T'$, while maintaining the invariant that any block in $T'$ that is labeled as invisible may not contain any visible points. Thus, the first step's role is to use $T'$ to filter out blocks in $T$ that are guaranteed to be invisible. The second step considers the blocks that were labeled as visible in $T'$ and computes the visibility of their points with full accuracy using the data in $T$. Overall the algorithm runs in $O(n + \frac nk \lg \frac nk + k \lg k + l \cdot \lg n)$, where $l$ is the total size of visible blocks in $T'$. When $k = Ω(1)$ and $l = o(n) $, the running time of our algorithm improves on the previous best bound of $O(n \lg n)$. Our experimental results show the performance of the new algorithm in practice and a speedup of more than an order of magnitude compared to previous algorithms.

Authors: Laura Toma

The viewshed of a point $v$ on a grid terrain $T$, viewshed$_T(v)$, is defined as the set of grid points in $T$ that are visible from $v$. We describe a novel algorithm for computing viewshed$_T(v)$ using a multi-resolution approach: Given a parameter $k >1$ that represents the block size, we create a grid $T'$ which is a lower-resolution version of $T$, such that each point in $T'$ corresponds to a block of $\lceil \sqrt k \rceil $-by-$\lceil \sqrt k \rceil$ points in $T$. The key of our approach is using $T'$ to speed up the computation of viewshed$_T(v)$ while not introducing approximation. We compute viewshed$_T(v)$ in two steps: First we compute the viewshed of $v$ on $T'$, while maintaining the invariant that any block in $T'$ that is labeled as invisible may not contain any visible points. Thus, the first step's role is to use $T'$ to filter out blocks in $T$ that are guaranteed to be invisible. The second step considers the blocks that were labeled as visible in $T'$ and computes the visibility of their points with full accuracy using the data in $T$. Overall the algorithm runs in $O(n + \frac nk \lg \frac nk + k \lg k + l \cdot \lg n)$, where $l$ is the total size of visible blocks in $T'$. When $k = Ω(1)$ and $l = o(n) $, the running time of our algorithm improves on the previous best bound of $O(n \lg n)$. Our experimental results show the performance of the new algorithm in practice and a speedup of more than an order of magnitude compared to previous algorithms.

A Faster Auction Algorithm for Weighted Matroid Intersection

from arXiv: Data Structures and Algorithms

Authors: Tatsuya Terao

We consider the weighted matroid intersection problem in the independence-oracle model. A sequence of works by Huang--Kakimura--Kamiyama [SODA'16 \& Math. Program'19], Chekuri--Quanrud [SODA'16], Quanrud [ICALP'24], and Dudeja--Grilnberger [IPCO'26] has developed efficient $(1-\varepsilon)$-approximation algorithms for this problem. We present a simple deterministic auction algorithm that, given two matroids on a common ground set of size $n$, computes a $(1-\varepsilon)$-approximate maximum-weight common independent set using $O(n \varepsilon^{-2} \log^2(n))$ independence-oracle queries. This is the first deterministic $(1-\varepsilon)$-approximation algorithm for the weighted matroid intersection problem whose query complexity is nearly linear in $n$ and polynomial in $1/\varepsilon$. Our algorithm builds on the auction algorithm for unweighted matroid intersection by Huang--Kobayashi ['26], together with the analysis of the auction algorithm for weighted bipartite matching by Liu--Ke--Khuller [APPROX'23].

Authors: Tatsuya Terao

We consider the weighted matroid intersection problem in the independence-oracle model. A sequence of works by Huang--Kakimura--Kamiyama [SODA'16 \& Math. Program'19], Chekuri--Quanrud [SODA'16], Quanrud [ICALP'24], and Dudeja--Grilnberger [IPCO'26] has developed efficient $(1-\varepsilon)$-approximation algorithms for this problem. We present a simple deterministic auction algorithm that, given two matroids on a common ground set of size $n$, computes a $(1-\varepsilon)$-approximate maximum-weight common independent set using $O(n \varepsilon^{-2} \log^2(n))$ independence-oracle queries. This is the first deterministic $(1-\varepsilon)$-approximation algorithm for the weighted matroid intersection problem whose query complexity is nearly linear in $n$ and polynomial in $1/\varepsilon$. Our algorithm builds on the auction algorithm for unweighted matroid intersection by Huang--Kobayashi ['26], together with the analysis of the auction algorithm for weighted bipartite matching by Liu--Ke--Khuller [APPROX'23].

Beyond odd characteristic: Faster isomorphism testing of 2-groups of Frattini class 2

from arXiv: Data Structures and Algorithms

Authors: Joshua A. Grochow, Gábor Ivanyos, Youming Qiao, Xiaorui Sun

The finite group isomorphism problem asks whether two finite groups of order $N$ are isomorphic. The first algorithm, attributed to Tarjan (see Miller, STOC '78), runs in time $N^{\log N + O(1)}$. Despite intensive study, the current best known algorithm has a running time of $N^{(1 / 4 + o(1))\log N}$ (Rosenbaum, '13). $p$-groups of class $2$ have been recognized as the major bottleneck for faster group isomorphism. Recent progress has led to $N^{o(\log N)}$-time algorithms for $p$-groups of class $2$ where $p$ is odd (Sun, STOC '23; Ivanyos--Mendoza--Qiao--Sun--Zhang, FOCS '24; Grochow--Qiao--Stange--Sun, STOC '25). However, the case of $p=2$, which represents the majority of $p$-groups of class 2 assuming a well-known conjecture in group enumeration, remained elusive, with essentially no progress until now. In this paper, we present an algorithm for testing the isomorphism of two 2-groups of Frattini class 2 of order $N$ in time $N^{O((\log N)^{1/2})}$. To our knowledge, this is the first $N^{o(\log N)}$-time isomorphism algorithm for a class of $2$-groups that constitutes logarithmically almost all $2$-groups, in the sense that $\lim_{N \to \infty} \frac{\log(\text{\# 2-groups of Frattini class 2 and order } \leq N)}{\log(\text{\# 2-groups of order} \leq N)} = 1$. As our main tool, we present the first non-trivial algorithms for the quadratic form space/tuple isometry problems over $\mathbb{F}_2$. These algorithms rely on combinations of combinatorial and algebraic ideas, including finite matrix group algorithms developed by Luks (FOCS '92). As far as we know, this is the first time that matrix group algorithms are used to make progress on the worst-case complexity of $p$-group isomorphism.

Authors: Joshua A. Grochow, Gábor Ivanyos, Youming Qiao, Xiaorui Sun

The finite group isomorphism problem asks whether two finite groups of order $N$ are isomorphic. The first algorithm, attributed to Tarjan (see Miller, STOC '78), runs in time $N^{\log N + O(1)}$. Despite intensive study, the current best known algorithm has a running time of $N^{(1 / 4 + o(1))\log N}$ (Rosenbaum, '13). $p$-groups of class $2$ have been recognized as the major bottleneck for faster group isomorphism. Recent progress has led to $N^{o(\log N)}$-time algorithms for $p$-groups of class $2$ where $p$ is odd (Sun, STOC '23; Ivanyos--Mendoza--Qiao--Sun--Zhang, FOCS '24; Grochow--Qiao--Stange--Sun, STOC '25). However, the case of $p=2$, which represents the majority of $p$-groups of class 2 assuming a well-known conjecture in group enumeration, remained elusive, with essentially no progress until now. In this paper, we present an algorithm for testing the isomorphism of two 2-groups of Frattini class 2 of order $N$ in time $N^{O((\log N)^{1/2})}$. To our knowledge, this is the first $N^{o(\log N)}$-time isomorphism algorithm for a class of $2$-groups that constitutes logarithmically almost all $2$-groups, in the sense that $\lim_{N \to \infty} \frac{\log(\text{\# 2-groups of Frattini class 2 and order } \leq N)}{\log(\text{\# 2-groups of order} \leq N)} = 1$. As our main tool, we present the first non-trivial algorithms for the quadratic form space/tuple isometry problems over $\mathbb{F}_2$. These algorithms rely on combinations of combinatorial and algebraic ideas, including finite matrix group algorithms developed by Luks (FOCS '92). As far as we know, this is the first time that matrix group algorithms are used to make progress on the worst-case complexity of $p$-group isomorphism.

Sparsification Framework for Directed Densest Subgraph

from arXiv: Data Structures and Algorithms

Authors: Slobodan Mitrović, Theodore Pan

We develop a new approach for computing approximate directed densest subgraphs (DDS). Our main result is a sparsification procedure that reduces a directed graph $G$ on $n$ vertices to a graph with $n \cdot \text{poly} \log n$ edges while preserving enough structure to recover an approximate DDS of $G$. Instantiating this framework in several memory-constrained settings, we obtain the following improvements over the state of the art: In semi-streaming, we obtain a single-pass algorithm that computes a $(1-\varepsilon)$-approximate DDS. Previously, the only semi-streaming algorithm that computed a constant approximation of DDS was by Bahmani, Kumar, and Vassilvitskii (2012), providing a $0.5-\varepsilon$ approximation in $O(\log n)$ passes. Hence, our work completely closes the approximation gap between undirected and directed DS in the semi-streaming setting, matching the $(1-\varepsilon)$-approximate undirected DS algorithm by Esfandiari, Hajiaghayi, and Woodruff (2016). In the near-linear-memory MPC regime, we obtain an $O(1)$-round algorithm for $(1-\varepsilon)$-approximate DDS, improving over the $O(\sqrt{\log n})$-round $(0.5-\varepsilon)$-approximation algorithm of Mitrović and Pan (2024). In the sublinear-time setting, we obtain an algorithm using $\tilde{O}(n)$ time, space, and oracle queries to compute a $(1-\varepsilon)$-approximate DDS, improving over the $\tilde{O}(n^{1.5})$ time, space, and query algorithm of Esfandiari, Hajiaghayi, and Woodruff (2016).

Authors: Slobodan Mitrović, Theodore Pan

We develop a new approach for computing approximate directed densest subgraphs (DDS). Our main result is a sparsification procedure that reduces a directed graph $G$ on $n$ vertices to a graph with $n \cdot \text{poly} \log n$ edges while preserving enough structure to recover an approximate DDS of $G$. Instantiating this framework in several memory-constrained settings, we obtain the following improvements over the state of the art: In semi-streaming, we obtain a single-pass algorithm that computes a $(1-\varepsilon)$-approximate DDS. Previously, the only semi-streaming algorithm that computed a constant approximation of DDS was by Bahmani, Kumar, and Vassilvitskii (2012), providing a $0.5-\varepsilon$ approximation in $O(\log n)$ passes. Hence, our work completely closes the approximation gap between undirected and directed DS in the semi-streaming setting, matching the $(1-\varepsilon)$-approximate undirected DS algorithm by Esfandiari, Hajiaghayi, and Woodruff (2016). In the near-linear-memory MPC regime, we obtain an $O(1)$-round algorithm for $(1-\varepsilon)$-approximate DDS, improving over the $O(\sqrt{\log n})$-round $(0.5-\varepsilon)$-approximation algorithm of Mitrović and Pan (2024). In the sublinear-time setting, we obtain an algorithm using $\tilde{O}(n)$ time, space, and oracle queries to compute a $(1-\varepsilon)$-approximate DDS, improving over the $\tilde{O}(n^{1.5})$ time, space, and query algorithm of Esfandiari, Hajiaghayi, and Woodruff (2016).

The Power of Two-Choice Linear Probing

from arXiv: Data Structures and Algorithms

Authors: Amir Azarmehr, Michael A. Bender, William Kuszmaul, Rose Silver

This paper considers the following basic question: If an (ordered) linear-probing hash table is allowed \emph{two} hash functions, instead of one, how does this change the expected insertion and query time, as a function of the load factor $1 - ε$? We prove that the \emph{greedy two-choice insertion strategy} achieves polynomially better bounds than the single choice algorithm, but that one can even do \emph{much better} by using more sophisticated non-greedy strategies. Specifically, we show that there is an insertion strategy that does not evict elements (once an element is inserted, its hash choice is fixed) and that achieves expected query time $O(\log ε^{-1})$ with expected insertion time $O(ε^{-1})$. We then further show that, if one is allowed to evict elements (i.e., to change over time which hash function a given element uses), then it is possible to achieve expected query time $O(1)$ with expected insertion time $O(ε^{-1/2})$. This final result achieves an expected query time of $O(1)$ even when the hash table is filled to $100\%$ full. Combined, the results reveal that there is a surprisingly strong ``power of two choices'' phenomenon for linear-probing hash tables, allowing for a two-choice hash table to achieve significantly better bounds than what might at first seem to be possible.

Authors: Amir Azarmehr, Michael A. Bender, William Kuszmaul, Rose Silver

This paper considers the following basic question: If an (ordered) linear-probing hash table is allowed \emph{two} hash functions, instead of one, how does this change the expected insertion and query time, as a function of the load factor $1 - ε$? We prove that the \emph{greedy two-choice insertion strategy} achieves polynomially better bounds than the single choice algorithm, but that one can even do \emph{much better} by using more sophisticated non-greedy strategies. Specifically, we show that there is an insertion strategy that does not evict elements (once an element is inserted, its hash choice is fixed) and that achieves expected query time $O(\log ε^{-1})$ with expected insertion time $O(ε^{-1})$. We then further show that, if one is allowed to evict elements (i.e., to change over time which hash function a given element uses), then it is possible to achieve expected query time $O(1)$ with expected insertion time $O(ε^{-1/2})$. This final result achieves an expected query time of $O(1)$ even when the hash table is filled to $100\%$ full. Combined, the results reveal that there is a surprisingly strong ``power of two choices'' phenomenon for linear-probing hash tables, allowing for a two-choice hash table to achieve significantly better bounds than what might at first seem to be possible.

Query-efficient winner prediction in district-based elections

from arXiv: Data Structures and Algorithms

Authors: Koustav De, Debajyoti Kar, Swagato Sanyal

In a district-based election, N voters are partitioned into k districts, and each voter votes for one of m candidates. Each district elects a winner using the plurality rule (i.e. the candidate getting the largest number of votes is declared the winner, breaking ties as per some fixed rule), and the overall winner is determined by applying plurality to the district winners; we assume that there is a unique winner amongst the district winners. The margin of victory of such an election is the minimum number of votes that must be altered so that the current winner ceases to be the unique district winner. We study the problem of predicting the winner of a district-based election in the query complexity model, where one has query access to individual votes. The objective is to minimise the number of queries. This setting captures exit polling, where queries correspond to interviewing voters, and is closely related to problems in query complexity and property testing. Assuming that the margin of victory of the election is at least eps N, Dey, Kar and Sanyal (AAMAS 2023) gave algorithms for the case of two candidates with error probability del and query complexity tilde{O}(1/eps^6 log^2 1/del), which improves to tilde{O}(1/eps^4 log^2 1/del) under the additional assumption that district populations are balanced. Our main result is an adaptive randomised algorithm that, for an arbitrary district-based election and any error parameter del, with probability at least 1-del, predicts the winner correctly using tilde{O}(1/eps^2 log m/del log 1/del) queries. In particular, we improve the bounds of Dey et al. for arbitrary district populations and extend their results to any number of candidates. Furthermore, for constantly many candidates, our algorithm nearly matches a lower bound of Omega(1/eps^2 log 1/del) on the query complexity that holds even for two candidates and a single district.

Authors: Koustav De, Debajyoti Kar, Swagato Sanyal

In a district-based election, N voters are partitioned into k districts, and each voter votes for one of m candidates. Each district elects a winner using the plurality rule (i.e. the candidate getting the largest number of votes is declared the winner, breaking ties as per some fixed rule), and the overall winner is determined by applying plurality to the district winners; we assume that there is a unique winner amongst the district winners. The margin of victory of such an election is the minimum number of votes that must be altered so that the current winner ceases to be the unique district winner. We study the problem of predicting the winner of a district-based election in the query complexity model, where one has query access to individual votes. The objective is to minimise the number of queries. This setting captures exit polling, where queries correspond to interviewing voters, and is closely related to problems in query complexity and property testing. Assuming that the margin of victory of the election is at least eps N, Dey, Kar and Sanyal (AAMAS 2023) gave algorithms for the case of two candidates with error probability del and query complexity tilde{O}(1/eps^6 log^2 1/del), which improves to tilde{O}(1/eps^4 log^2 1/del) under the additional assumption that district populations are balanced. Our main result is an adaptive randomised algorithm that, for an arbitrary district-based election and any error parameter del, with probability at least 1-del, predicts the winner correctly using tilde{O}(1/eps^2 log m/del log 1/del) queries. In particular, we improve the bounds of Dey et al. for arbitrary district populations and extend their results to any number of candidates. Furthermore, for constantly many candidates, our algorithm nearly matches a lower bound of Omega(1/eps^2 log 1/del) on the query complexity that holds even for two candidates and a single district.

Faster Algorithms for Finding Small Induced Patterns in Sparse Host Graphs

from arXiv: Data Structures and Algorithms

Authors: Priyanshi Agrawal, Balagopal Komarath

We study algorithms for detecting induced subgraphs corresponding to fixed pattern graphs in host graphs. We show that at least five of the 21 connected graphs on five vertices can be detected in time roughly the product of the number of vertices and the number of edges, and that at least 65 of the 112 connected graphs on six vertices can be detected in time nearly quadratic in the number of edges. We also give algorithms for detecting induced paths and cycles on seven vertices, running in time roughly the number of vertices times the square of the number of edges. Our main technical tool is a generalized notion of tree decomposition width, called (p, q)-width. It yields algorithms whose running times depend on both the number of vertices and the number of edges, and are never worse than existing bounds. Whenever the host graph has fewer than roughly quadratically many edges in its number of vertices, our bounds are strictly faster. For some patterns, including the seven-vertex cycle, our algorithms are optimal under standard complexity-theoretic assumptions. We further develop this approach using pattern-based polynomials that exploit the structure of tree decompositions, not just their width. This gives algorithms for detecting induced paths and cycles on an even number of vertices in bipartite graphs, running in time roughly the (k-1)-th power of the number of edges for paths on 2k vertices, and that same bound times the number of vertices for cycles on 2k vertices. These are faster than the best known algorithms for general graphs.

Authors: Priyanshi Agrawal, Balagopal Komarath

We study algorithms for detecting induced subgraphs corresponding to fixed pattern graphs in host graphs. We show that at least five of the 21 connected graphs on five vertices can be detected in time roughly the product of the number of vertices and the number of edges, and that at least 65 of the 112 connected graphs on six vertices can be detected in time nearly quadratic in the number of edges. We also give algorithms for detecting induced paths and cycles on seven vertices, running in time roughly the number of vertices times the square of the number of edges. Our main technical tool is a generalized notion of tree decomposition width, called (p, q)-width. It yields algorithms whose running times depend on both the number of vertices and the number of edges, and are never worse than existing bounds. Whenever the host graph has fewer than roughly quadratically many edges in its number of vertices, our bounds are strictly faster. For some patterns, including the seven-vertex cycle, our algorithms are optimal under standard complexity-theoretic assumptions. We further develop this approach using pattern-based polynomials that exploit the structure of tree decompositions, not just their width. This gives algorithms for detecting induced paths and cycles on an even number of vertices in bipartite graphs, running in time roughly the (k-1)-th power of the number of edges for paths on 2k vertices, and that same bound times the number of vertices for cycles on 2k vertices. These are faster than the best known algorithms for general graphs.

Unifying and Extending Strong Simulation of Quantum Circuits

from arXiv: Data Structures and Algorithms

Authors: Floris Geerts, Rihan Hai, Matthias Lanzinger, Reinhard Pichler, Emanuel Sallinger, Daniel Unterberger

We establish functional aggregate queries (FAQs) as a unifying language for exact classical simulation of quantum circuits. A circuit becomes a sum-product query: factors encode gates, internal wire variables are aggregated, and free boundary variables index transition amplitudes. The central insight is that distinct sources of simulation tractability can be exploited within the same InsideOut evaluation scheme. The query specifies what is computed; the evaluation plan, semiring, and representation of intermediate factors determine the cost. This view unifies structural and algebraic simulation guarantees. With explicit factor representations, FAQ evaluation recovers the treewidth bound for tensor-network contraction and yields finer sparsity-sensitive bounds via fractional covers. Over a formal phase semiring, compressed intermediate factors recover rank-width-based simulation for compatible quadratic phase representations. For Clifford circuits, affine-quadratic factors are closed under multiplication and marginalization and remain polynomial in size, yielding polynomial-time exact amplitude computation without any bounded-width assumption. Beyond these recoveries, the framework yields a new tractability criterion: tensor layout symmetry width. This parameter combines local cut-rank with separator symmetry through exact tree-tensor representations. We give a constructive evaluation bound and exhibit a circuit family with bounded tensor layout symmetry width but unbounded phase-graph rank-width and circuit line-graph treewidth. These results establish representation-aware FAQ evaluation as a common algorithmic foundation for classical simulation and a systematic route to new tractable regimes.

Authors: Floris Geerts, Rihan Hai, Matthias Lanzinger, Reinhard Pichler, Emanuel Sallinger, Daniel Unterberger

We establish functional aggregate queries (FAQs) as a unifying language for exact classical simulation of quantum circuits. A circuit becomes a sum-product query: factors encode gates, internal wire variables are aggregated, and free boundary variables index transition amplitudes. The central insight is that distinct sources of simulation tractability can be exploited within the same InsideOut evaluation scheme. The query specifies what is computed; the evaluation plan, semiring, and representation of intermediate factors determine the cost. This view unifies structural and algebraic simulation guarantees. With explicit factor representations, FAQ evaluation recovers the treewidth bound for tensor-network contraction and yields finer sparsity-sensitive bounds via fractional covers. Over a formal phase semiring, compressed intermediate factors recover rank-width-based simulation for compatible quadratic phase representations. For Clifford circuits, affine-quadratic factors are closed under multiplication and marginalization and remain polynomial in size, yielding polynomial-time exact amplitude computation without any bounded-width assumption. Beyond these recoveries, the framework yields a new tractability criterion: tensor layout symmetry width. This parameter combines local cut-rank with separator symmetry through exact tree-tensor representations. We give a constructive evaluation bound and exhibit a circuit family with bounded tensor layout symmetry width but unbounded phase-graph rank-width and circuit line-graph treewidth. These results establish representation-aware FAQ evaluation as a common algorithmic foundation for classical simulation and a systematic route to new tractable regimes.

Dynamic Connectivity, Minimum Spanning Tree, and 2-Edge Connectivity with Polylogarithmic Worst-Case Update Time

from arXiv: Data Structures and Algorithms

Authors: Simon Meierhans, Maximilian Probst Gutenberg, Yu-Cheng Yeh

We give fully dynamic algorithms for maintaining connectivity, minimum spanning tree, and $2$-edge connectivity of a graph with worst-case polylogarithmic update time. Our algorithms are randomized and succeed with high probability against an adaptive adversary. For the minimum spanning tree and $2$-edge connectivity problems, this improves over the subpolynomial update time bounds obtained by Nanongkai, Saranurak, and Wulff-Nilsen [FOCS'17], Jin and Sun [FOCS'21], and Jin, Sun, and Thorup [SODA'24], respectively. The only randomized component of our algorithms is the computation of static expander decompositions, and a deterministic algorithm for said problem would directly imply deterministic algorithms for all three problems. This reduction is novel even for the connectivity problem.

Authors: Simon Meierhans, Maximilian Probst Gutenberg, Yu-Cheng Yeh

We give fully dynamic algorithms for maintaining connectivity, minimum spanning tree, and $2$-edge connectivity of a graph with worst-case polylogarithmic update time. Our algorithms are randomized and succeed with high probability against an adaptive adversary. For the minimum spanning tree and $2$-edge connectivity problems, this improves over the subpolynomial update time bounds obtained by Nanongkai, Saranurak, and Wulff-Nilsen [FOCS'17], Jin and Sun [FOCS'21], and Jin, Sun, and Thorup [SODA'24], respectively. The only randomized component of our algorithms is the computation of static expander decompositions, and a deterministic algorithm for said problem would directly imply deterministic algorithms for all three problems. This reduction is novel even for the connectivity problem.

Thursday, October 01

The Keynesian Subtweet

from Ben Recht

Utility maximization is unescapable. We should learn when it’s misapplied.

Hi there, argmin readers! Today’s post is a live blog of Class 9 of my graduate seminar “Forecasting: A Critical Retrospective.” The syllabus and list of past posts are here.

Every class I teach seems to require a lecture or two on expected utility maximization. It’s the basis of classification rules, so it’s integral to machine learning classes. It’s the simplest stochastic optimization problem, so it appears in my optimization classes. It’s a great motivator for probabilistic thinking, so it appears in probability classes. Perhaps a bit less well appreciated, it’s an indirect proper scoring rule for forecasts, so we have to talk about it in this class. And, I suppose, it’s a mathematical tool that scarily dominates the philosophy of many powerful people in government and industry alike. It’s worth understanding the nuts and bolts!

Part of what makes utility maximization so appealing is how simple the core idea is. We want to decide whether to act or not. We build a probabilistic model of the world under the action and compute the expected value of the benefit. We also build a model for what would happen if we don’t act and compute the expected value of the benefit of inaction. We choose to act if our calculations imply that action has more benefit than inaction.

What could be simpler? You compute probabilities. You compute costs. You multiply them together and add them up. Everything is beautifully quantified and calculable.

The only problem is these numbers are all made up. They are forecasts, and they are rarely justifiable. Outside of the casino, we rarely know how to calculate precise odds of outcomes. Worse, forecasting the prices out in the future is also inherently uncertain, often more uncertain than the odds calculations. Maximizing the utility of a rational actor or a general population is an appealing philosophical goal, but the precise rational calculations are always made about fantasy stories. Economists know this, and proceed with caution. In the comments of Tuesday’s post, Jordan Ellenberg flagged this quote from John Maynard Keynes.

“By ‘uncertain’ knowledge, let me explain, I do not mean merely to distinguish what is known for certain from what is only probable. The game of roulette is not subject, in this sense, to uncertainty… Even the weather is only moderately uncertain. The sense in which I am using the term is that in which the prospect of a European war is uncertain, or the price of copper and the rate of interest twenty years hence, or the obsolescence of a new invention, or the position of private wealthowners in the social system in 1970. About these matters there is no scientific basis on which to form any calculable probability whatever. We simply do not know. Nevertheless, the necessity for action and for decision compels us as practical men to do our best to overlook this awkward fact and to behave exactly as we should if we had behind us a good Benthamite calculation of a series of prospective advantages and disadvantages, each multiplied by its appropriate probability, waiting to be summed.”

Keynes here seems to be reluctantly defending utility maximization as the best of many bad options for decision making in the face of uncertainty. But I don’t think that’s his point at all! Keynes is defending his “General Theory of Employment” and arguing against this sort of utilitarian calculation.

As it is colloquially known, “Keynesian economics” — a term that sadly oversimplifies Keynes’ brilliant collected works — tells governments to stimulate economies during recessions. Keynes argues that we can’t accurately predict the future, and hence people tend to lean on status-quo bias. The safest bet is to assume nothing changes. But when the status quo is bad, people become risk-averse and hold their money. This hoarding prolongs recessions. Because the future is uncertain, people don’t act “rationally” when they fear they might not have money to spend later. Keynes thus argues that governments should step in to nudge citizens out of their undue pessimism by giving them extra money to spend.

Written in the pits of the Great Depression, Keynes’ argument in his 1937 article in the Quarterly Journal of Economics articulates how utility maximization goes wrong. The future is unknowable. Uncertainty makes people afraid. Fear makes hoarding more attractive. Mass hoarding perpetuates a vicious cycle of societal hardship. At this point, someone needs to step in to get the ball rolling, eating up some of the risk to inspire more confidence in people to spend their money.

In this class, we’re not going to debate the merits and challenges of recessionary stimulus spending. But Keynes’ article (which I have added to the reading list) articulates the nuance of forecasting under great uncertainty. Human psychology plays a key role. Today we’ll go through how this appears in the mathematics, with different utility functions expressing different models of risk aversion, and different algorithms for action. We’ll see that not only are the costs and benefits uncertain, but the algorithm itself can be shaped by different models of psychology. Expected utility maximization might be a reasonable way to engineer machines, but it’s not something that reasonable people do. Understanding the many hyperparameters of optimal decision making can help us better understand the psychology of people obsessed with forecasting.

Subscribe now

By Ben Recht

Computational Work Extraction: The Complexity of Catalysts

from arXiv: Computational Complexity

Authors: Atul Singh Arora, Shantanav Chakraborty, Alexandru Cojocaru, Sreyas Saminathan, Uttam Singh

We prove maximal separations: $n$-qubit systems can have $Θ(n)$ ergotropy, while every efficient process extracts negligible work, even for Hamiltonians consisting of single-qubit terms. We establish an unconditional existential separation and give an explicit construction in the random oracle model. Assuming the existence of quantum-secure pseudorandom functions, this separation extends to the plain model. This work uncovers an important connection between ergotropy and the complexity of catalytic computation---computation where auxiliary qubits must be finally restored to their initial state. Relative to a random oracle, we establish relational and decision problems that: (i) can be solved efficiently with $λ$ catalysts; but (ii) cannot be solved by any algorithm with $cλ$ catalysts, for any $c<1$. We show this by proving query lower bounds for quantum-space bounded algorithms. As a consequence, for computational ergotropy, catalysts prove to be surprisingly powerful---there is a family of Hamiltonians and states for which catalysts enable efficient extraction of the full $Θ(n)$ ergotropy, while every efficient non-catalytic process extracts negligible work. Furthermore, catalysts also allow us to introduce and instantiate the notion of pseudoergotropy---analogous to pseudorandomness. On the other hand, we show catalysts do not change (information-theoretic) ergotropy. Finally, our work also sheds light on the classical aspect of the problem. First, most of our constructions rely on classical states and Hamiltonians and therefore imply analogous results for classical ergotropy. Second, we show that certain proof of quantumness protocols can be used to generically separate classical and quantum catalytic ergotropy.

Authors: Atul Singh Arora, Shantanav Chakraborty, Alexandru Cojocaru, Sreyas Saminathan, Uttam Singh

We prove maximal separations: $n$-qubit systems can have $Θ(n)$ ergotropy, while every efficient process extracts negligible work, even for Hamiltonians consisting of single-qubit terms. We establish an unconditional existential separation and give an explicit construction in the random oracle model. Assuming the existence of quantum-secure pseudorandom functions, this separation extends to the plain model. This work uncovers an important connection between ergotropy and the complexity of catalytic computation---computation where auxiliary qubits must be finally restored to their initial state. Relative to a random oracle, we establish relational and decision problems that: (i) can be solved efficiently with $λ$ catalysts; but (ii) cannot be solved by any algorithm with $cλ$ catalysts, for any $c<1$. We show this by proving query lower bounds for quantum-space bounded algorithms. As a consequence, for computational ergotropy, catalysts prove to be surprisingly powerful---there is a family of Hamiltonians and states for which catalysts enable efficient extraction of the full $Θ(n)$ ergotropy, while every efficient non-catalytic process extracts negligible work. Furthermore, catalysts also allow us to introduce and instantiate the notion of pseudoergotropy---analogous to pseudorandomness. On the other hand, we show catalysts do not change (information-theoretic) ergotropy. Finally, our work also sheds light on the classical aspect of the problem. First, most of our constructions rely on classical states and Hamiltonians and therefore imply analogous results for classical ergotropy. Second, we show that certain proof of quantumness protocols can be used to generically separate classical and quantum catalytic ergotropy.

Computational Bounds for $f$-Routing

from arXiv: Computational Complexity

Authors: Oren Renard, Nicholas Spooner

The $f$-routing protocol is a leading candidate for quantum position verification (Kent, Munro, and Spiller, 2011), but security guarantees for explicit functions remain limited. We prove unconditional resource lower bounds for uniform attackers; our new techniques bypass communication-complexity bounds central to previous works, which are inherently at most linear in the input length. We show that, for input length $n$ and sufficiently small constant $ε>0$, a uniformly generated strategy using $q$ qubits and having description length $\mathrm{poly}(q)$, with success probability at least $1-ε$ on every input, implies the following computational bounds on $f$: 1. If the strategies are arbitrary quantum channels, then $f\in\mathrm{QSZK}(\mathrm{poly}(nq))$, where $\mathrm{QSZK}(T)$ is the class of languages having quantum statistical zero knowledge proofs in which the verifier runs in time $T$ (and the simulator in time $\mathrm{poly}(T)$). 2. If the strategies are explicit Pauli-sparse unitaries on $q$ qubits that have at most $s$ nonzero Pauli coefficients, then $f\in\mathrm{DTIME}(\mathrm{poly}(nqs))$. 3. If the strategies are Clifford+T circuits using at most $t$ magic gates, then $f\in\mathrm{DTIME}(\mathrm{poly}(nq2^t))$. Time and space hierarchies then yield explicit functions secure against polynomial and even quasipolynomial qubits $q$ under our computational restrictions. These bounds exceed the $q\le\log n$ bound of Bluhm, Christandl, and Speelman (2022) for inner product function $f=\mathrm{IP}$, at the cost of restricting adversarial computation and increasing honest evaluation complexity.

Authors: Oren Renard, Nicholas Spooner

The $f$-routing protocol is a leading candidate for quantum position verification (Kent, Munro, and Spiller, 2011), but security guarantees for explicit functions remain limited. We prove unconditional resource lower bounds for uniform attackers; our new techniques bypass communication-complexity bounds central to previous works, which are inherently at most linear in the input length. We show that, for input length $n$ and sufficiently small constant $ε>0$, a uniformly generated strategy using $q$ qubits and having description length $\mathrm{poly}(q)$, with success probability at least $1-ε$ on every input, implies the following computational bounds on $f$: 1. If the strategies are arbitrary quantum channels, then $f\in\mathrm{QSZK}(\mathrm{poly}(nq))$, where $\mathrm{QSZK}(T)$ is the class of languages having quantum statistical zero knowledge proofs in which the verifier runs in time $T$ (and the simulator in time $\mathrm{poly}(T)$). 2. If the strategies are explicit Pauli-sparse unitaries on $q$ qubits that have at most $s$ nonzero Pauli coefficients, then $f\in\mathrm{DTIME}(\mathrm{poly}(nqs))$. 3. If the strategies are Clifford+T circuits using at most $t$ magic gates, then $f\in\mathrm{DTIME}(\mathrm{poly}(nq2^t))$. Time and space hierarchies then yield explicit functions secure against polynomial and even quasipolynomial qubits $q$ under our computational restrictions. These bounds exceed the $q\le\log n$ bound of Bluhm, Christandl, and Speelman (2022) for inner product function $f=\mathrm{IP}$, at the cost of restricting adversarial computation and increasing honest evaluation complexity.

Exponential Quantum Advantage in Numbers-on-Forehead Communication

from arXiv: Computational Complexity

Authors: Haoyu Wang, Pei Wu, Guangxu Yang

We give the first exponential quantum advantage in the general interactive three-party Numbers-on-Forehead (NOF) model for a decision problem. Previous separations hold only for restricted protocols like one-way communication for a relation. We construct an explicit partial Boolean function, the Interleaved Unitary Product problem, that requires only $O(\log n)$ NOF quantum communication but $\widetildeΩ(n^{1/32})$ randomized communication. This function builds on the two-party unitary product problem of Arunachalam, Girish, and Lifshitz (TQC 2024). The main technical obstacle is that discrepancy, the standard lower-bound method for NOF, also lower-bounds quantum communication. We instead develop a regularity-based argument for randomized NOF lower bounds, building on the approach of Kelley, Lovett, and Meka and adapting the regularity decomposition of Abboud, Fischer, Kelley, Lovett, and Meka (STOC 2024) to cylinder intersections. Combined with matrix-product estimates of Arunachalam, Girish, and Lifshitz, this yields our randomized lower bound.

Authors: Haoyu Wang, Pei Wu, Guangxu Yang

We give the first exponential quantum advantage in the general interactive three-party Numbers-on-Forehead (NOF) model for a decision problem. Previous separations hold only for restricted protocols like one-way communication for a relation. We construct an explicit partial Boolean function, the Interleaved Unitary Product problem, that requires only $O(\log n)$ NOF quantum communication but $\widetildeΩ(n^{1/32})$ randomized communication. This function builds on the two-party unitary product problem of Arunachalam, Girish, and Lifshitz (TQC 2024). The main technical obstacle is that discrepancy, the standard lower-bound method for NOF, also lower-bounds quantum communication. We instead develop a regularity-based argument for randomized NOF lower bounds, building on the approach of Kelley, Lovett, and Meka and adapting the regularity decomposition of Abboud, Fischer, Kelley, Lovett, and Meka (STOC 2024) to cylinder intersections. Combined with matrix-product estimates of Arunachalam, Girish, and Lifshitz, this yields our randomized lower bound.

Quantum Algorithms for Minimum Generating Set

from arXiv: Computational Complexity

Authors: Bireswar Das, Udit Kumar, Kavita Samant, Dhara Thakkar

In this paper, we present a polynomial-time quantum algorithm for computing a minimum-sized generating set of solvable black-box groups. Next, we consider the class $Γ_d$ of black-box groups, where every non-abelian composition factor is isomorphic to a subgroup of the symmetric group $S_d$ for a fixed $d$. We design polynomial-time quantum algorithms to compute the direct product decomposition of abelian factor groups and solve the constructive membership problem for factor groups of groups from $Γ_d$. With the help of these algorithms, we design a quantum algorithm for computing a chief series of black-box groups from $Γ_d$. Using the chief series, we construct a polynomial-time quantum algorithm for computing minimum generating sets of black-box groups from $Γ_d$. Finally, we show that the minimum generating set problem for general black-box groups is in $\textrm{NP} \cap \textrm{coAM}$.

Authors: Bireswar Das, Udit Kumar, Kavita Samant, Dhara Thakkar

In this paper, we present a polynomial-time quantum algorithm for computing a minimum-sized generating set of solvable black-box groups. Next, we consider the class $Γ_d$ of black-box groups, where every non-abelian composition factor is isomorphic to a subgroup of the symmetric group $S_d$ for a fixed $d$. We design polynomial-time quantum algorithms to compute the direct product decomposition of abelian factor groups and solve the constructive membership problem for factor groups of groups from $Γ_d$. With the help of these algorithms, we design a quantum algorithm for computing a chief series of black-box groups from $Γ_d$. Using the chief series, we construct a polynomial-time quantum algorithm for computing minimum generating sets of black-box groups from $Γ_d$. Finally, we show that the minimum generating set problem for general black-box groups is in $\textrm{NP} \cap \textrm{coAM}$.

A physical and universal model of bosonic computations with Solovay-Kitaev theorem

from arXiv: Computational Complexity

Authors: Dorian Rudolph, Arsalan Motamedi, Dhruva Sambrani, Hamid Reza Naeij, Ulysse Chabaud, Sevag Gharibian, Saeed Mehraban

Bosonic quantum systems are among the leading architectures for quantum information processing, offering continuous-variable degrees of freedom with strong error-correction capabilities. However, standard bosonic quantum computation models such as the Lloyd-Braunstein [Lloyd and Braunstein, 1999] and hybrid oscillator-qubit models [Brenner, Dias, and Koenig, 2025; Liu et al., 2026] permit dramatic energy growth, leading to unphysical computational power and the breakdown of fundamental algorithmic tools such as universal and efficient compilation [Brenner et al., 2026; Rudolph et al., 2025]. To address this, we introduce a new model of bosonic quantum computation, Bosonic Energy-Preserving Quantum Computation (BEQC), in which energy is treated as a computational resource. Namely, energy is supplied solely through input coherent states, and all gates are generated by energy-preserving Hamiltonians. Thus, by construction, dramatic energy growth is impossible, making the model physically grounded. We next show that BEQC is a computationally robust and universal model in many respects, including: (1. Computational power) BEQC efficiently simulates all polynomial-energy computations in existing models, and exactly recovers BQP in the polynomial energy setting. It further admits several complexity-theoretic upper bounds when varying the energy, precision, and space parameters of the model. (2. Universal gate sets and state synthesis) BEQC has natural universal gate sets based on linear optics and Kerr interactions. In particular, we obtain a Solovay-Kitaev theorem which circumvents previous no-go results. We give various applications, including (a) a protocol for engineering GKP states with rigorous preparation guarantees, (b) Fock state preparation to exponential precision, and (c) native Fock space simulation of any qubit-based unitary.

Authors: Dorian Rudolph, Arsalan Motamedi, Dhruva Sambrani, Hamid Reza Naeij, Ulysse Chabaud, Sevag Gharibian, Saeed Mehraban

Bosonic quantum systems are among the leading architectures for quantum information processing, offering continuous-variable degrees of freedom with strong error-correction capabilities. However, standard bosonic quantum computation models such as the Lloyd-Braunstein [Lloyd and Braunstein, 1999] and hybrid oscillator-qubit models [Brenner, Dias, and Koenig, 2025; Liu et al., 2026] permit dramatic energy growth, leading to unphysical computational power and the breakdown of fundamental algorithmic tools such as universal and efficient compilation [Brenner et al., 2026; Rudolph et al., 2025]. To address this, we introduce a new model of bosonic quantum computation, Bosonic Energy-Preserving Quantum Computation (BEQC), in which energy is treated as a computational resource. Namely, energy is supplied solely through input coherent states, and all gates are generated by energy-preserving Hamiltonians. Thus, by construction, dramatic energy growth is impossible, making the model physically grounded. We next show that BEQC is a computationally robust and universal model in many respects, including: (1. Computational power) BEQC efficiently simulates all polynomial-energy computations in existing models, and exactly recovers BQP in the polynomial energy setting. It further admits several complexity-theoretic upper bounds when varying the energy, precision, and space parameters of the model. (2. Universal gate sets and state synthesis) BEQC has natural universal gate sets based on linear optics and Kerr interactions. In particular, we obtain a Solovay-Kitaev theorem which circumvents previous no-go results. We give various applications, including (a) a protocol for engineering GKP states with rigorous preparation guarantees, (b) Fock state preparation to exponential precision, and (c) native Fock space simulation of any qubit-based unitary.

A mixing time method for estimating the sample complexity of quantum state discrimination

from arXiv: Computational Complexity

Authors: Juntai Zhou, Felix Leditzky

We develop a mixing time method for estimating the sample complexity of quantum state discrimination. We start with considering the minimum-error discrimination of geometrically uniform pure state ensembles, and prove that its sample complexity has a tight estimate given by a quantum homogeneous mixing time [George et al., 2026] and a quantum version of the generalized Dobrushin coefficient [Wolfer, 2020]. This quantum mixing time further reduces to a classical one when the generating group $G$ forms a Gelfand pair with the stabilizer subgroup $H$ of the generator state. In this case the generalized Dobrushin coefficient can be fully expressed by representation-theoretic quantities of the commutative Hecke algebra $\operatorname{End}_G(\mathbb C[G/H])$. In particular, this method reduces the sample complexity estimation of learning quantum coupon collector states [Arunachalam et al., 2020] and learning phase states to classical mixing time problems. We apply this framework to answer the open problems of learning degree-$d$ phase states over $\mathbb F_q$ in [Alrabiah et al., 2026] and generalized Boolean phase states over $\mathbb Z_q$ [Arunachalam et al., 2023]. The framework also applies to hypergraph state ensembles, giving estimates expressed fully in terms of hypergraph data and recovering estimates for graph state ensembles in [Montanaro and Shao, 2022]. Finally, we extend the discussion to arbitrary mixed state ensembles with uniform priors, prove a sandwiched bound for minimum-error discrimination sample complexity by a quantum weakly mixing time, and provide a tight estimate for the minimax discrimination sample complexity from [D'Ariano et al., 2005] by a Dobrushin-type coefficient. We also discuss the method of strengthened data processing inequality [Gao and Rouz{é}, 2022] and give an upper bound in terms of a strengthened data processing inequality constant.

Authors: Juntai Zhou, Felix Leditzky

We develop a mixing time method for estimating the sample complexity of quantum state discrimination. We start with considering the minimum-error discrimination of geometrically uniform pure state ensembles, and prove that its sample complexity has a tight estimate given by a quantum homogeneous mixing time [George et al., 2026] and a quantum version of the generalized Dobrushin coefficient [Wolfer, 2020]. This quantum mixing time further reduces to a classical one when the generating group $G$ forms a Gelfand pair with the stabilizer subgroup $H$ of the generator state. In this case the generalized Dobrushin coefficient can be fully expressed by representation-theoretic quantities of the commutative Hecke algebra $\operatorname{End}_G(\mathbb C[G/H])$. In particular, this method reduces the sample complexity estimation of learning quantum coupon collector states [Arunachalam et al., 2020] and learning phase states to classical mixing time problems. We apply this framework to answer the open problems of learning degree-$d$ phase states over $\mathbb F_q$ in [Alrabiah et al., 2026] and generalized Boolean phase states over $\mathbb Z_q$ [Arunachalam et al., 2023]. The framework also applies to hypergraph state ensembles, giving estimates expressed fully in terms of hypergraph data and recovering estimates for graph state ensembles in [Montanaro and Shao, 2022]. Finally, we extend the discussion to arbitrary mixed state ensembles with uniform priors, prove a sandwiched bound for minimum-error discrimination sample complexity by a quantum weakly mixing time, and provide a tight estimate for the minimax discrimination sample complexity from [D'Ariano et al., 2005] by a Dobrushin-type coefficient. We also discuss the method of strengthened data processing inequality [Gao and Rouz{é}, 2022] and give an upper bound in terms of a strengthened data processing inequality constant.

Improved Quantum Random Self-Reduction for Linear Problems

from arXiv: Computational Complexity

Authors: Vahid R. Asadi, Shuichi Hirahara, Nobutaka Shimizu

We study quantum random self-reductions for linear problems over finite fields. Let $M\in\mathbb{F}^{n\times n}$ be an arbitrary matrix, and let $\mathcal{O}$ be an oracle that agrees with the linear map $x\mapsto Mx$ on an $\varepsilon$-fraction of inputs $x\sim\mathbb{F}^n$. Given coherent access to $\mathcal{O}$ and coherent entry access to $M$, we give a uniform quantum reduction that computes $Mx$ on any prescribed input $x$ with probability at least $2/3$ in time $\widetilde{O}(nT^{1/3})$, for $n\le T\le n^{3/2}$ and constant field size and $\varepsilon$, where $T$ is the cost of one coherent query to $\mathcal{O}$. In particular, when $T=\widetilde{O}(n)$, the reduction runs in time $\widetilde{O}(n^{4/3})$, improving the $\widetilde{O}(n^{3/2}+T)$ reduction of Asadi, Golovnev, Gur, Shinkar, and Subramanian (SODA 2024). Our reduction uses the Bogolyubov--Ruzsa subspace guaranteed by additive combinatorics, but it avoids learning this subspace explicitly, which was computationally expensive for the previous reduction; in particular, it does not recover a basis for its orthogonal complement. The main technical step is to decompose the inputs into sparse pieces and find a vector that lies outside the Bogolyubov--Ruzsa subspace via a quantum search based on amplitude amplification. This yields a tunable tradeoff between the cost of querying the average-case oracle and the cost of verifying matrix-vector products.

Authors: Vahid R. Asadi, Shuichi Hirahara, Nobutaka Shimizu

We study quantum random self-reductions for linear problems over finite fields. Let $M\in\mathbb{F}^{n\times n}$ be an arbitrary matrix, and let $\mathcal{O}$ be an oracle that agrees with the linear map $x\mapsto Mx$ on an $\varepsilon$-fraction of inputs $x\sim\mathbb{F}^n$. Given coherent access to $\mathcal{O}$ and coherent entry access to $M$, we give a uniform quantum reduction that computes $Mx$ on any prescribed input $x$ with probability at least $2/3$ in time $\widetilde{O}(nT^{1/3})$, for $n\le T\le n^{3/2}$ and constant field size and $\varepsilon$, where $T$ is the cost of one coherent query to $\mathcal{O}$. In particular, when $T=\widetilde{O}(n)$, the reduction runs in time $\widetilde{O}(n^{4/3})$, improving the $\widetilde{O}(n^{3/2}+T)$ reduction of Asadi, Golovnev, Gur, Shinkar, and Subramanian (SODA 2024). Our reduction uses the Bogolyubov--Ruzsa subspace guaranteed by additive combinatorics, but it avoids learning this subspace explicitly, which was computationally expensive for the previous reduction; in particular, it does not recover a basis for its orthogonal complement. The main technical step is to decompose the inputs into sparse pieces and find a vector that lies outside the Bogolyubov--Ruzsa subspace via a quantum search based on amplitude amplification. This yields a tunable tradeoff between the cost of querying the average-case oracle and the cost of verifying matrix-vector products.

Security Properties of Neural Networks as Decision Problems

from arXiv: Computational Complexity

Authors: Adrian Wurm

Certifying a deployed neural network raises decision problems that the verification literature has not classified: whether the model carries a backdoor planted in its training data, whether a fault in its stored parameters can drive it into an unsafe state, whether its output leaks a private part of its input. We formalise eight such problems and classify what we can. The organising observation is a logical one. The function computed by a piecewise linear network, together with all its node values, is definable by a quantifier-free formula of real addition of size linear in the network, so a property of the network is a quantifier-alternation sentence, which Sontag's 1985 theorem places in the polynomial hierarchy at the level of its prefix. Membership results are thus corollaries, and the argument makes plain what they need: that the quantified objects are inputs rather than the network's own parameters. Non-interference, monotonicity and counterfactual fairness have exactly the complexity of network equivalence and of interval verification, all co-NP- complete over ReLU. Detection of backdoor triggers from a quantised alphabet is Sigma_2^P-complete, one level above robustness certification, so it does not reduce to polynomially many robustness queries unless the hierarchy collapses. Inversion resistance is co-NP-complete for every l_p metric, p a fixed positive integer. Quantifying over parameters instead of inputs - the fault model of bit-flip attacks, radiation upsets and analog accelerators - makes verification exists-R-complete already for networks of identity nodes, for which every previously studied problem is in P, and it stays so when each parameter is confined to a box of inverse-polynomial width; the corresponding safety question is forall-R-complete for ReLU.

Authors: Adrian Wurm

Certifying a deployed neural network raises decision problems that the verification literature has not classified: whether the model carries a backdoor planted in its training data, whether a fault in its stored parameters can drive it into an unsafe state, whether its output leaks a private part of its input. We formalise eight such problems and classify what we can. The organising observation is a logical one. The function computed by a piecewise linear network, together with all its node values, is definable by a quantifier-free formula of real addition of size linear in the network, so a property of the network is a quantifier-alternation sentence, which Sontag's 1985 theorem places in the polynomial hierarchy at the level of its prefix. Membership results are thus corollaries, and the argument makes plain what they need: that the quantified objects are inputs rather than the network's own parameters. Non-interference, monotonicity and counterfactual fairness have exactly the complexity of network equivalence and of interval verification, all co-NP- complete over ReLU. Detection of backdoor triggers from a quantised alphabet is Sigma_2^P-complete, one level above robustness certification, so it does not reduce to polynomially many robustness queries unless the hierarchy collapses. Inversion resistance is co-NP-complete for every l_p metric, p a fixed positive integer. Quantifying over parameters instead of inputs - the fault model of bit-flip attacks, radiation upsets and analog accelerators - makes verification exists-R-complete already for networks of identity nodes, for which every previously studied problem is in P, and it stays so when each parameter is confined to a box of inverse-polynomial width; the corresponding safety question is forall-R-complete for ReLU.

QMA(2) with Limited Shared Entanglement

from arXiv: Computational Complexity

Authors: Alex Della Schiava, Ranitha Mataraarachchi

A $\mathsf{QMA}(2)$ protocol involves two provers submitting unentangled witnesses to a polynomial-time quantum verifier. In The Power of Unentanglement (ToC, 2009), Aaronson et al. proposed $\mathsf{QMA}(2;h)$, a variant of $\mathsf{QMA}(2)$ in which the two provers may share $h$ EPR pairs. Our main result shows that the power of $\mathsf{QMA}(2)$ remains unchanged for up to logarithmically many shared EPR pairs: $\mathsf{QMA}(2;h)=\mathsf{QMA}(2)$ for $h=O(\log n)$, where $n$ is the input length. The result follows from a simulation using four unentangled witnesses, combined with the Harrow-Montanaro equality $\mathsf{QMA}(4)=\mathsf{QMA}(2)$ (FOCS, 2010). We also prove monotonicity in the EPR budget: $\mathsf{QMA}(2;h)\subseteq\mathsf{QMA}(2;H)$ for $h\le H$, preserving completeness and soundness. Combined with input padding, this shows that establishing $\mathsf{QMA}(2;n^\varepsilon)=\mathsf{QMA}(2)$ for any fixed $\varepsilon>0$ would imply equality for every polynomially bounded budget, resolving the open problem raised by Aaronson et al. Finally, we extend these results to a variant of the model in which the provers may use local operations and classical communication (LOCC) during witness preparation. For logarithmic-size witnesses and inverse-polynomial gaps, both models remain equivalent to their unentangled counterpart when $h=O(\log n)$. We show that extending this equivalence to any superlogarithmic EPR budget in the LOCC model would imply $\mathsf{NP}\subseteq\mathsf{BQP}$.

Authors: Alex Della Schiava, Ranitha Mataraarachchi

A $\mathsf{QMA}(2)$ protocol involves two provers submitting unentangled witnesses to a polynomial-time quantum verifier. In The Power of Unentanglement (ToC, 2009), Aaronson et al. proposed $\mathsf{QMA}(2;h)$, a variant of $\mathsf{QMA}(2)$ in which the two provers may share $h$ EPR pairs. Our main result shows that the power of $\mathsf{QMA}(2)$ remains unchanged for up to logarithmically many shared EPR pairs: $\mathsf{QMA}(2;h)=\mathsf{QMA}(2)$ for $h=O(\log n)$, where $n$ is the input length. The result follows from a simulation using four unentangled witnesses, combined with the Harrow-Montanaro equality $\mathsf{QMA}(4)=\mathsf{QMA}(2)$ (FOCS, 2010). We also prove monotonicity in the EPR budget: $\mathsf{QMA}(2;h)\subseteq\mathsf{QMA}(2;H)$ for $h\le H$, preserving completeness and soundness. Combined with input padding, this shows that establishing $\mathsf{QMA}(2;n^\varepsilon)=\mathsf{QMA}(2)$ for any fixed $\varepsilon>0$ would imply equality for every polynomially bounded budget, resolving the open problem raised by Aaronson et al. Finally, we extend these results to a variant of the model in which the provers may use local operations and classical communication (LOCC) during witness preparation. For logarithmic-size witnesses and inverse-polynomial gaps, both models remain equivalent to their unentangled counterpart when $h=O(\log n)$. We show that extending this equivalence to any superlogarithmic EPR budget in the LOCC model would imply $\mathsf{NP}\subseteq\mathsf{BQP}$.

On quantum interactive proofs with a laconic prover

from arXiv: Computational Complexity

Authors: Zihan Hu, Yupan Liu

Interactive proof systems with a laconic prover, studied by Goldreich, Vadhan, and Wigderson (CC, 2002), capture problems verifiable with logarithmic prover communication in the classical setting. For two-message quantum analogs, even a single-bit prover response contains quantum statistical zero-knowledge ($\sf QSZK$), introduced by Watrous (FOCS 2002). However, restricting the verifier's question to classical public coins collapses the corresponding class to $\sf BQP$, as shown by Beigi, Shor, and Watrous (ToC, 2011). We further study two-message quantum interactive proof systems with a laconic prover. To this end, we introduce the class ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$, where $\ell$ is the length of the prover's response, and establish: 1. A natural complete characterization of ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$ by Multi-State Distinguishability. In particular, Quantum State Distinguishability (QSD) is ${\sf QIP}_{\rm bit}$-complete. Since QSD is $\sf QSZK$-hard, our result places ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$, for $\ell\geq 2$, in a landscape "just above" $\sf QSZK$. 2. Easy regimes for ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$ collapsing to $\sf QSZK$. We prove that QSD$[a,b]$ (and thus ${\sf QIP}_{\rm bit}[a,b]$) is in $\sf QSZK$ when $a(n)-b(n)\geq 1/O(\log n)$, and combine this with an answer compression from ${\sf QIP}_{\ell\text{-}{\rm bit}}[2,c,s]$ to ${\sf QIP}_{\rm bit}$ to obtain another easy regime when $2c>(1+2^{\ell/2})s$. Remarkably, our improved polarization applies to SD and $\sf SZK$, resolving an open problem in Sahai and Vadhan (JACM, 2003). 3. Quantum public coins also make the interaction useless: ${\sf qc}\text{-}{\sf QAM}[O(\sqrt{\log{n}})]$ with constant gap is in $\sf BQP$, where ${\sf qc}\text{-}{\sf QAM}[\ell]$ is a subclass of ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$ in which the verifier's question is exactly halves of EPR pairs.

Authors: Zihan Hu, Yupan Liu

Interactive proof systems with a laconic prover, studied by Goldreich, Vadhan, and Wigderson (CC, 2002), capture problems verifiable with logarithmic prover communication in the classical setting. For two-message quantum analogs, even a single-bit prover response contains quantum statistical zero-knowledge ($\sf QSZK$), introduced by Watrous (FOCS 2002). However, restricting the verifier's question to classical public coins collapses the corresponding class to $\sf BQP$, as shown by Beigi, Shor, and Watrous (ToC, 2011). We further study two-message quantum interactive proof systems with a laconic prover. To this end, we introduce the class ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$, where $\ell$ is the length of the prover's response, and establish: 1. A natural complete characterization of ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$ by Multi-State Distinguishability. In particular, Quantum State Distinguishability (QSD) is ${\sf QIP}_{\rm bit}$-complete. Since QSD is $\sf QSZK$-hard, our result places ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$, for $\ell\geq 2$, in a landscape "just above" $\sf QSZK$. 2. Easy regimes for ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$ collapsing to $\sf QSZK$. We prove that QSD$[a,b]$ (and thus ${\sf QIP}_{\rm bit}[a,b]$) is in $\sf QSZK$ when $a(n)-b(n)\geq 1/O(\log n)$, and combine this with an answer compression from ${\sf QIP}_{\ell\text{-}{\rm bit}}[2,c,s]$ to ${\sf QIP}_{\rm bit}$ to obtain another easy regime when $2c>(1+2^{\ell/2})s$. Remarkably, our improved polarization applies to SD and $\sf SZK$, resolving an open problem in Sahai and Vadhan (JACM, 2003). 3. Quantum public coins also make the interaction useless: ${\sf qc}\text{-}{\sf QAM}[O(\sqrt{\log{n}})]$ with constant gap is in $\sf BQP$, where ${\sf qc}\text{-}{\sf QAM}[\ell]$ is a subclass of ${\sf QIP}_{\ell\text{-}{\rm bit}}(2)$ in which the verifier's question is exactly halves of EPR pairs.

Information Potential: A Variational Approach to Quantum Information Complexity

from arXiv: Computational Complexity

Authors: Penghui Yao, Yifan Zhou

Quantum information complexity (QIC), introduced by Touchette [Touchette, STOC 2015], has been shown to be one of the most powerful methods for proving quantum communication complexity and has also been shown to be equal to amortized quantum communication complexity. Unfortunately, QIC is generally hard to analyze because it is a sum of quantum conditional mutual information terms, which are difficult to estimate. In this work, we introduce a new variational approach, the information potential, for analyzing QIC and proving quantum communication lower bounds inspired by the resolvent representation for quantum relative entropy. This approach connects the QIC of individual messages to a quadratic form, making it more amenable and thus enables us to bound the cumulative positive increments of the potential throughout an interactive quantum protocol by its QIC. Lower bounds on the growth of the potential therefore translate into lower bounds on both QIC and quantum communication complexity. As an application, we give an optimal $Ω(1/r)$ lower bound on the QIC of the two-bit $\mathsf{AND}$ function as well as an optimal $Ω(n/r)$ lower bound on the quantum communication complexity of $r$-round Set-Disjointness, answering an open problem in~[Braverman, Garg, Ko, Mao, Touchette FOCS 2015]. Moreover, we further prove a direct-sum theorem for bounded-round quantum communication complexity of Set Disjointness. With the tight bound on the QIC of $\mathsf{AND}$ function, we further establish a nearly tight tradeoff for the asymmetric quantum communication complexity of $\mathsf{Set} \mathsf{Disjointness}$: $(q_A+1)(q_B+1)=Ω(n)$, where $q_A$ and $q_B$ denote the total numbers of qubits sent by Alice and Bob, respectively.

Authors: Penghui Yao, Yifan Zhou

Quantum information complexity (QIC), introduced by Touchette [Touchette, STOC 2015], has been shown to be one of the most powerful methods for proving quantum communication complexity and has also been shown to be equal to amortized quantum communication complexity. Unfortunately, QIC is generally hard to analyze because it is a sum of quantum conditional mutual information terms, which are difficult to estimate. In this work, we introduce a new variational approach, the information potential, for analyzing QIC and proving quantum communication lower bounds inspired by the resolvent representation for quantum relative entropy. This approach connects the QIC of individual messages to a quadratic form, making it more amenable and thus enables us to bound the cumulative positive increments of the potential throughout an interactive quantum protocol by its QIC. Lower bounds on the growth of the potential therefore translate into lower bounds on both QIC and quantum communication complexity. As an application, we give an optimal $Ω(1/r)$ lower bound on the QIC of the two-bit $\mathsf{AND}$ function as well as an optimal $Ω(n/r)$ lower bound on the quantum communication complexity of $r$-round Set-Disjointness, answering an open problem in~[Braverman, Garg, Ko, Mao, Touchette FOCS 2015]. Moreover, we further prove a direct-sum theorem for bounded-round quantum communication complexity of Set Disjointness. With the tight bound on the QIC of $\mathsf{AND}$ function, we further establish a nearly tight tradeoff for the asymmetric quantum communication complexity of $\mathsf{Set} \mathsf{Disjointness}$: $(q_A+1)(q_B+1)=Ω(n)$, where $q_A$ and $q_B$ denote the total numbers of qubits sent by Alice and Bob, respectively.

The Commuting Local Hamiltonian Problem: Relativized Evidence Against BQP-Hardness

from arXiv: Computational Complexity

Authors: Itay Shalit, Mark Zhandry

The commuting local-Hamiltonian (CLH) problem is a restriction of the local-Hamiltonian problem, in which the terms of the Hamiltonian are required to pairwise commute. A long line of work has shown that the problem lies in $\mathsf{NP}$ for certain families of commuting local Hamiltonians. Nevertheless, there has been no formal evidence against the possibility that the general CLH problem is $\mathsf{QMA}$-complete. CLH is complete for the complexity class $\mathsf{QIMA}$, defined through quantum verifiers whose local gates commute. Therefore, CLH is $\mathsf{QMA}$-complete if and only if $\mathsf{QIMA}=\mathsf{QMA}$. In this work, we introduce a classical-oracle analogue $\mathsf{QIMA}^{\mathcal{O}}$ and construct a classical oracle $\mathcal{O}$ such that $\mathsf{BQP}^{\mathcal{O}}\not\subseteq\mathsf{QIMA}^{\mathcal{O}}$. Since $\mathsf{BQP}^{\mathcal{O}}\subseteq\mathsf{QMA}^{\mathcal{O}}$ for any classical oracle $\mathcal{O}$, this implies $\mathsf{QIMA}^{\mathcal{O}}\neq\mathsf{QMA}^{\mathcal{O}}$ for our constructed oracle. Thus, our result provides relativized evidence against the possibility that the general commuting local-Hamiltonian problem is $\mathsf{BQP}$-hard, and hence also against the possibility that it is $\mathsf{QMA}$-complete.

Authors: Itay Shalit, Mark Zhandry

The commuting local-Hamiltonian (CLH) problem is a restriction of the local-Hamiltonian problem, in which the terms of the Hamiltonian are required to pairwise commute. A long line of work has shown that the problem lies in $\mathsf{NP}$ for certain families of commuting local Hamiltonians. Nevertheless, there has been no formal evidence against the possibility that the general CLH problem is $\mathsf{QMA}$-complete. CLH is complete for the complexity class $\mathsf{QIMA}$, defined through quantum verifiers whose local gates commute. Therefore, CLH is $\mathsf{QMA}$-complete if and only if $\mathsf{QIMA}=\mathsf{QMA}$. In this work, we introduce a classical-oracle analogue $\mathsf{QIMA}^{\mathcal{O}}$ and construct a classical oracle $\mathcal{O}$ such that $\mathsf{BQP}^{\mathcal{O}}\not\subseteq\mathsf{QIMA}^{\mathcal{O}}$. Since $\mathsf{BQP}^{\mathcal{O}}\subseteq\mathsf{QMA}^{\mathcal{O}}$ for any classical oracle $\mathcal{O}$, this implies $\mathsf{QIMA}^{\mathcal{O}}\neq\mathsf{QMA}^{\mathcal{O}}$ for our constructed oracle. Thus, our result provides relativized evidence against the possibility that the general commuting local-Hamiltonian problem is $\mathsf{BQP}$-hard, and hence also against the possibility that it is $\mathsf{QMA}$-complete.

Frustration Free Stoquastic Local Hamiltonian with Sub-Constant Gap is in NP

from arXiv: Computational Complexity

Authors: Eshan Chattopadhyay, Oren Renard, Nicholas Spooner

We continue the study of the Stoquastic Local Hamiltonian problem, a physically motivated restriction of the QMA-complete Local Hamiltonian problem (Kitaev, Shen, and Vyalyi, 2002). For the $β$-gapped, frustration-free case, Bravyi, Bessen, and Terhal (2006) showed that the problem is MA-complete when $β= 1/\mathrm{poly}(n)$. Aharonov and Grilo (2019) derandomized this algorithm and proved membership in NP for constant gap $β= Ω(1)$. We present an improved algorithm and analysis, establishing membership in NP even when $β= Ω(1/(\log\log n))$. We complement our result with an explicit example demonstrating why the analysis does not extend directly to $β= o(1/\log\log n)$.

Authors: Eshan Chattopadhyay, Oren Renard, Nicholas Spooner

We continue the study of the Stoquastic Local Hamiltonian problem, a physically motivated restriction of the QMA-complete Local Hamiltonian problem (Kitaev, Shen, and Vyalyi, 2002). For the $β$-gapped, frustration-free case, Bravyi, Bessen, and Terhal (2006) showed that the problem is MA-complete when $β= 1/\mathrm{poly}(n)$. Aharonov and Grilo (2019) derandomized this algorithm and proved membership in NP for constant gap $β= Ω(1)$. We present an improved algorithm and analysis, establishing membership in NP even when $β= Ω(1/(\log\log n))$. We complement our result with an explicit example demonstrating why the analysis does not extend directly to $β= o(1/\log\log n)$.

Experimentally Testable Quantum Advantage in Shallow Circuits

from arXiv: Computational Complexity

Authors: Kishor Bharti, Adán Cabello

Experimental tests of shallow-circuit quantum advantage require explicit classical bounds at finite circuit sizes. We refine the finite-size classical soundness bound of Aasnaess's graph-distributed construction to depend linearly on the number of players. Combined with standard disjoint-player repetition, this gives a two-round test on a single processor for any finite nonlocal game with a finite-dimensional perfect quantum strategy and classical winning probability $γ<1$, with arbitrarily small classical success. The method is based on playing $m$ copies with disjoint players and teleporting each player's register to a uniformly random one of $N$ sites before the questions are revealed. Each answer bit of a depth-$D$, fan-in-$K$ classical response with fixed wiring depends on at most $K^D$ question wires, making cross-player dependencies unlikely when $N$ is large. The quantum implementation has constant depth per round and wins with certainty. A classical device wins with probability at most $γ^m+O(mK^D/N)$, which vanishes as $O(\log N/N)$ for $m=\lceil\log_{1/γ}N\rceil$ at fixed $D$ and $K$. We present an explicit proposal for an experimentally testable quantum advantage with 99 qubits.

Authors: Kishor Bharti, Adán Cabello

Experimental tests of shallow-circuit quantum advantage require explicit classical bounds at finite circuit sizes. We refine the finite-size classical soundness bound of Aasnaess's graph-distributed construction to depend linearly on the number of players. Combined with standard disjoint-player repetition, this gives a two-round test on a single processor for any finite nonlocal game with a finite-dimensional perfect quantum strategy and classical winning probability $γ<1$, with arbitrarily small classical success. The method is based on playing $m$ copies with disjoint players and teleporting each player's register to a uniformly random one of $N$ sites before the questions are revealed. Each answer bit of a depth-$D$, fan-in-$K$ classical response with fixed wiring depends on at most $K^D$ question wires, making cross-player dependencies unlikely when $N$ is large. The quantum implementation has constant depth per round and wins with certainty. A classical device wins with probability at most $γ^m+O(mK^D/N)$, which vanishes as $O(\log N/N)$ for $m=\lceil\log_{1/γ}N\rceil$ at fixed $D$ and $K$. We present an explicit proposal for an experimentally testable quantum advantage with 99 qubits.

A Separation Between Types of Quantum Oracle Separations

from arXiv: Computational Complexity

Authors: Scott Aaronson, Adam Bouland, Jordan Docter, Barak Nehoran

Recent works have demonstrated that quantum oracles have subtle behavior, as access to inverse, conjugate or controlled queries can exponentially change the query complexity of certain tasks. Inspired by these works, we introduce the notion of meta-complexity of quantum relativization. We ask: for any two quantum complexity classes, under which "types" of quantum oracles are they equal or separated? Different pantheons of oracles (or quantum oracle types, e.g. unitary vs state, poly- vs superpoly-dimensional, closed under inverse or not) form a partially ordered set based on their power in separating complexity classes. Moreover, two oracle pantheons A and B are separated if there exists a pair of complexity classes that are separated under an oracle from pantheon A but yet the complexity classes are equivalent under all oracles from pantheon B. We show that this meta-complexity can be nontrivial by giving a complete classification, within the family of oracle pantheons defined in this paper, of which models can separate the complexity classes $\mathsf{PostBQP}$ and $\mathsf{PreciseBQP}$, the exponentially precise variant of $\mathsf{BQP}$. Both classes equal $\mathsf{PP}$ in the unrelativized setting. Within our taxonomy, they remain equal relative to real or polynomial-dimensional unitary oracles and whenever inverse or conjugate access is supplied. In contrast, we give a separation relative to forward-only complex diagonal unitaries of superpolynomial dimension, as well as a separation relative to single-qubit state-preparation oracles. We view this as a test case for the meta-complexity of oracles which underscores the subtlety inherent to the relativization of quantum complexity classes.

Authors: Scott Aaronson, Adam Bouland, Jordan Docter, Barak Nehoran

Recent works have demonstrated that quantum oracles have subtle behavior, as access to inverse, conjugate or controlled queries can exponentially change the query complexity of certain tasks. Inspired by these works, we introduce the notion of meta-complexity of quantum relativization. We ask: for any two quantum complexity classes, under which "types" of quantum oracles are they equal or separated? Different pantheons of oracles (or quantum oracle types, e.g. unitary vs state, poly- vs superpoly-dimensional, closed under inverse or not) form a partially ordered set based on their power in separating complexity classes. Moreover, two oracle pantheons A and B are separated if there exists a pair of complexity classes that are separated under an oracle from pantheon A but yet the complexity classes are equivalent under all oracles from pantheon B. We show that this meta-complexity can be nontrivial by giving a complete classification, within the family of oracle pantheons defined in this paper, of which models can separate the complexity classes $\mathsf{PostBQP}$ and $\mathsf{PreciseBQP}$, the exponentially precise variant of $\mathsf{BQP}$. Both classes equal $\mathsf{PP}$ in the unrelativized setting. Within our taxonomy, they remain equal relative to real or polynomial-dimensional unitary oracles and whenever inverse or conjugate access is supplied. In contrast, we give a separation relative to forward-only complex diagonal unitaries of superpolynomial dimension, as well as a separation relative to single-qubit state-preparation oracles. We view this as a test case for the meta-complexity of oracles which underscores the subtlety inherent to the relativization of quantum complexity classes.

Top-Down Lower Bounds for All Depths

from arXiv: Computational Complexity

Authors: Oliver Korten

We prove that Parity requires $2^{n^{Ω(1)}}$ size De Morgan circuits of constant depth using a new method which is completely "top-down" in the sense of [HJP95]. The proof relies crucially on the core ideas developed in a line of work [HJP95, PPZ99, MW19, GRSS24] which previously established top-down lower bounds for circuits of depth 3 and 4. We first present a proof of a lower bound $\exp(n^{3^{-d}})$. In this case, nearly all of the relevant combinatorial ideas necessary for the proof are already present in some form in [GRSS24]. We then present two extensions of this argument, the first achieving a lower bound $\exp(ε_d n^{1/(2d-2)})$ for some $ε_d>0$ depending only on $d$, and the second achieving the essentially tight lower bound $\exp(ε_d n^{1/(d-1)})$. These improved results each hinge on establishing a key lemma which quantifies the extent to which a high entropy random variable in $\{0,1\}^n$ will look close to uniform after projecting it onto a random small set of coordinates $R \subseteq [n]$.

Authors: Oliver Korten

We prove that Parity requires $2^{n^{Ω(1)}}$ size De Morgan circuits of constant depth using a new method which is completely "top-down" in the sense of [HJP95]. The proof relies crucially on the core ideas developed in a line of work [HJP95, PPZ99, MW19, GRSS24] which previously established top-down lower bounds for circuits of depth 3 and 4. We first present a proof of a lower bound $\exp(n^{3^{-d}})$. In this case, nearly all of the relevant combinatorial ideas necessary for the proof are already present in some form in [GRSS24]. We then present two extensions of this argument, the first achieving a lower bound $\exp(ε_d n^{1/(2d-2)})$ for some $ε_d>0$ depending only on $d$, and the second achieving the essentially tight lower bound $\exp(ε_d n^{1/(d-1)})$. These improved results each hinge on establishing a key lemma which quantifies the extent to which a high entropy random variable in $\{0,1\}^n$ will look close to uniform after projecting it onto a random small set of coordinates $R \subseteq [n]$.

Karp's NP-complete problems over first-order definable structures

from arXiv: Computational Complexity

Authors: Aidan Healy, Bartek Klin

We determine the decidability of Karp's NP-complete problems on structures which are first-order definable over the theory of equality, also known as orbit-finite sets with atoms or nominal sets.

Authors: Aidan Healy, Bartek Klin

We determine the decidability of Karp's NP-complete problems on structures which are first-order definable over the theory of equality, also known as orbit-finite sets with atoms or nominal sets.

Parameterized Hardness of Mixed 2-Sided Orthant Depth

from arXiv: Computational Geometry

Authors: Michelle Döring, Georg Tennigkeit

We consider the maximum-depth problem for mixed 2-sided orthants in R^d: each region imposes one lower bound and one upper bound on distinct coordinates, and the task is to find a point contained in as many regions as possible. We show that the corresponding decision problem is W[1]-hard when parameterized by the dimension. Our reduction from MultiColoredClique uses two coordinates per color class and only polynomially many orthants.

Authors: Michelle Döring, Georg Tennigkeit

We consider the maximum-depth problem for mixed 2-sided orthants in R^d: each region imposes one lower bound and one upper bound on distinct coordinates, and the task is to find a point contained in as many regions as possible. We show that the corresponding decision problem is W[1]-hard when parameterized by the dimension. Our reduction from MultiColoredClique uses two coordinates per color class and only polynomially many orthants.

Flexible discrete translational surfaces

from arXiv: Computational Geometry

Authors: Georg Nawratil

We give a full list of translational nets which flex within their class of discrete surfaces of translation, by reducing the classification problem to the one of flexible complete bipartite frameworks on the sphere, for which the solution is known. We also obtained two novel classes which correspond to Bottema's spherical 16-bar mechanisms and the constant diagonal angle frameworks. Based on an algorithm for the construction of all flexible translational nets, we also discuss flexible translational tubes and toroids. Furthermore, we present novel results for both topologies which are implied by Bottema's spherical 16-bar mechanisms.

Authors: Georg Nawratil

We give a full list of translational nets which flex within their class of discrete surfaces of translation, by reducing the classification problem to the one of flexible complete bipartite frameworks on the sphere, for which the solution is known. We also obtained two novel classes which correspond to Bottema's spherical 16-bar mechanisms and the constant diagonal angle frameworks. Based on an algorithm for the construction of all flexible translational nets, we also discuss flexible translational tubes and toroids. Furthermore, we present novel results for both topologies which are implied by Bottema's spherical 16-bar mechanisms.

A Unified Dual Method for Matching Problems

from arXiv: Computational Geometry

Authors: Guillaume Houry, Ferdinand Genans, Jean Feydy, François-Xavier Vialard

Matching problems are ubiquitous in data science as they enable the alignment of structured objects and distributions. While existing solvers are often tailored to specific matching formulations, we unify a broad class of such problems within a common mathematical and optimization framework based on duality theory. Theoretically, we demonstrate that matching objectives decomposable as a difference of convex (DC) functions can be recast as implicit registration problems. This connection links matching to another well-studied class of objectives and yields a dual formulation amenable to natural optimization strategies. We then apply these findings to quadratic matching (QM) problems, which admit DC decompositions and for which we provide extensive convergence guarantees. Our framework applies to Gromov-Wasserstein (GW), as well as its unbalanced formulation and several variants, which are increasingly popular QM problems. Numerically, we implement our algorithms at scale for various data modalities such as graphs, point clouds, meshes, and word embeddings. Finally, our modular approach allows us to explore new formulations such as fracture matching, broadening the scope of problems that can be addressed within this framework..

Authors: Guillaume Houry, Ferdinand Genans, Jean Feydy, François-Xavier Vialard

Matching problems are ubiquitous in data science as they enable the alignment of structured objects and distributions. While existing solvers are often tailored to specific matching formulations, we unify a broad class of such problems within a common mathematical and optimization framework based on duality theory. Theoretically, we demonstrate that matching objectives decomposable as a difference of convex (DC) functions can be recast as implicit registration problems. This connection links matching to another well-studied class of objectives and yields a dual formulation amenable to natural optimization strategies. We then apply these findings to quadratic matching (QM) problems, which admit DC decompositions and for which we provide extensive convergence guarantees. Our framework applies to Gromov-Wasserstein (GW), as well as its unbalanced formulation and several variants, which are increasingly popular QM problems. Numerically, we implement our algorithms at scale for various data modalities such as graphs, point clouds, meshes, and word embeddings. Finally, our modular approach allows us to explore new formulations such as fracture matching, broadening the scope of problems that can be addressed within this framework..

A differentiability framework for zigzag persistent homology via linear interpolation

from arXiv: Computational Geometry

Authors: Enrico Maria Ferrari, Clemens Bannwart, Matteo Biagetti

Persistent homology can be differentiated and incorporated into learning pipelines, but no analogous framework exists for zigzag persistence, which is needed when the underlying topological structure evolves non-monotonically over time. We develop such a framework for sequences of simplicial complexes obtained by thresholding time-dependent filtering values on a fixed complex. By assigning persistence diagram endpoints the real-valued times at which linearly interpolated filtering values cross the threshold, we transfer the continuity of the filtering values to the diagram points. This yields smooth local lifts of the resulting persistence-diagram-valued map, from which we derive differentials almost everywhere under mild regularity conditions on the parametrization of the filtering values. We prove local Lipschitz continuity outside an explicit measure-zero exclusion set; standard stochastic subgradient convergence guarantees therefore do not apply directly. We argue that, even without such guarantees, this exclusion set is small enough in practice to allow effective optimization. We test this empirically in two experiments: sensor network coverage optimization and dynamic graph classification.

Authors: Enrico Maria Ferrari, Clemens Bannwart, Matteo Biagetti

Persistent homology can be differentiated and incorporated into learning pipelines, but no analogous framework exists for zigzag persistence, which is needed when the underlying topological structure evolves non-monotonically over time. We develop such a framework for sequences of simplicial complexes obtained by thresholding time-dependent filtering values on a fixed complex. By assigning persistence diagram endpoints the real-valued times at which linearly interpolated filtering values cross the threshold, we transfer the continuity of the filtering values to the diagram points. This yields smooth local lifts of the resulting persistence-diagram-valued map, from which we derive differentials almost everywhere under mild regularity conditions on the parametrization of the filtering values. We prove local Lipschitz continuity outside an explicit measure-zero exclusion set; standard stochastic subgradient convergence guarantees therefore do not apply directly. We argue that, even without such guarantees, this exclusion set is small enough in practice to allow effective optimization. We test this empirically in two experiments: sensor network coverage optimization and dynamic graph classification.

Gibbs Sampling in the Shattered Phase by Decoded Quantum Interferometry

from arXiv: Data Structures and Algorithms

Authors: Leo Zhou, Noah Shutty, Mark Sellke, Stephen P. Jordan

We apply Decoded Quantum Interferometry (DQI) to sample from the Gibbs measures of classical Ising spin Hamiltonians. We show that this Gibbs sampling problem reduces to a quantum decoding problem, and the temperature achievable by DQI is determined by the performance of decoding algorithms. We then focus on the task of Gibbs sampling for classical Ising $k$-spin glasses (or Max-$k$-XORSAT) on random Erdős-Rényi hypergraphs with average degree $D\ge k$. In a temperature range beginning asymptotically at the predicted dynamical phase transition, $β_{\rm dyn}(k,D) = \sqrt{(2\ln k)/D}\times [1+o_{k\to\infty}(1)]$, we show that shattering and disorder chaos form a topological barrier that obstructs many algorithms, including Glauber dynamics and any algorithm whose output distribution is "stable" under perturbations of the input. In contrast, we prove that this barrier can be broken both by a classical algorithm based on Prange's method, and by DQI equipped with a quantum decoder. For example, when $D=αk$ with fixed $α>1$, both Prange's algorithm and DQI can sample at any inverse temperature $β< \tanh^{-1}(1/α)$ for sufficiently large $k$, well beyond the dynamical threshold $β_{\rm dyn} \sim \sqrt{2\ln k / (αk)}$. Therefore, our results show that DQI can overcome topological barriers that obstruct stable algorithms.

Authors: Leo Zhou, Noah Shutty, Mark Sellke, Stephen P. Jordan

We apply Decoded Quantum Interferometry (DQI) to sample from the Gibbs measures of classical Ising spin Hamiltonians. We show that this Gibbs sampling problem reduces to a quantum decoding problem, and the temperature achievable by DQI is determined by the performance of decoding algorithms. We then focus on the task of Gibbs sampling for classical Ising $k$-spin glasses (or Max-$k$-XORSAT) on random Erdős-Rényi hypergraphs with average degree $D\ge k$. In a temperature range beginning asymptotically at the predicted dynamical phase transition, $β_{\rm dyn}(k,D) = \sqrt{(2\ln k)/D}\times [1+o_{k\to\infty}(1)]$, we show that shattering and disorder chaos form a topological barrier that obstructs many algorithms, including Glauber dynamics and any algorithm whose output distribution is "stable" under perturbations of the input. In contrast, we prove that this barrier can be broken both by a classical algorithm based on Prange's method, and by DQI equipped with a quantum decoder. For example, when $D=αk$ with fixed $α>1$, both Prange's algorithm and DQI can sample at any inverse temperature $β< \tanh^{-1}(1/α)$ for sufficiently large $k$, well beyond the dynamical threshold $β_{\rm dyn} \sim \sqrt{2\ln k / (αk)}$. Therefore, our results show that DQI can overcome topological barriers that obstruct stable algorithms.

Quantum Fine-Grained Lower Bounds for SetDisjointness via Sub-Linear Reductions from 3SUM

from arXiv: Data Structures and Algorithms

Authors: Jeremy Huang, Young Kun Ko, Chunhao Wang

In classical fine-grained complexity, the 3SUM Conjecture is used to prove a variety of conditional lower bounds on data structure and graph problems via an initial reduction to the SetDisjointness problem. However, there is an $\tilde{O}(n)$-time quantum algorithm for 3SUM and a direct application of Grover's algorithm to SetDisjointness queries beats the state-of-the-art classical conditional bound by Kopelowitz, Pettie, and Porat (SODA 2016); this shows that these classical bounds do not apply in the quantum setting. Thus establishing analogous conditional lower bounds in the quantum setting requires applying the quantum 3SUM Conjecture to a \emph{quantum} fine-grained reduction from 3SUM to SetDisjointness. We give the first sub-linear time quantum reductions from 3SUM to online SetDisjointness. Via our reduction, the quantum 3SUM conjecture implies a $p + 2q \geqslant 1$ tradeoff bound for quantum SetDisjointness algorithms with $O(N^p)$ preprocessing time and $O(N^q)$ query time. We also give an analogous reduction from 3XOR. These results are derived from a general framework for fine-grained reductions to SetDisjointness which applies to any Abelian 3-Orthogonal Array (3OA) problem with suitable almost-linear hash functions.

Authors: Jeremy Huang, Young Kun Ko, Chunhao Wang

In classical fine-grained complexity, the 3SUM Conjecture is used to prove a variety of conditional lower bounds on data structure and graph problems via an initial reduction to the SetDisjointness problem. However, there is an $\tilde{O}(n)$-time quantum algorithm for 3SUM and a direct application of Grover's algorithm to SetDisjointness queries beats the state-of-the-art classical conditional bound by Kopelowitz, Pettie, and Porat (SODA 2016); this shows that these classical bounds do not apply in the quantum setting. Thus establishing analogous conditional lower bounds in the quantum setting requires applying the quantum 3SUM Conjecture to a \emph{quantum} fine-grained reduction from 3SUM to SetDisjointness. We give the first sub-linear time quantum reductions from 3SUM to online SetDisjointness. Via our reduction, the quantum 3SUM conjecture implies a $p + 2q \geqslant 1$ tradeoff bound for quantum SetDisjointness algorithms with $O(N^p)$ preprocessing time and $O(N^q)$ query time. We also give an analogous reduction from 3XOR. These results are derived from a general framework for fine-grained reductions to SetDisjointness which applies to any Abelian 3-Orthogonal Array (3OA) problem with suitable almost-linear hash functions.

Verifiable quantum advantage based on polynomials with planted structures

from arXiv: Data Structures and Algorithms

Authors: Markus Bläser, Michael Gullans, Dominik Hangleiter, Yuxuan Liu, Youming Qiao

A central question in the theory of quantum advantage is whether there are quantum advantage protocols with similar resource requirements as random circuit sampling that are also verifiable just from the classical outputs of the quantum computation. Here, we develop the idea of simulation secrets for verifiable advantage. A verifier can use a simulation secret to evaluate a cross-entropy test faster than it would take a classical adversary to pass the test. We instantiate this idea using IQP circuits described by cubic polynomials with planted independent spaces. These correspond to the largest independent set in the orbit of a polynomial under the general linear group and yield a low-rank stabilizer decomposition of the corresponding state. We conjecture that large independent spaces are invisible to a computationally bounded adversary, and therefore they cannot exploit them to pass the protocol. A second conjecture regards the fine-grained complexity of producing samples that pass the cross-entropy test for uniformly random polynomials. Under these conjectures, our scheme results in a polynomial gap between the verification time and the time a classical adversary would need to pass the protocol---both are exponential. It has a potential application to generating classically certifiable randomness, since the output distributions have high min-entropy. We estimate that the planted polynomial scheme is implementable using 100 logical qubits at logical error rates around $10^{-6}$.

Authors: Markus Bläser, Michael Gullans, Dominik Hangleiter, Yuxuan Liu, Youming Qiao

A central question in the theory of quantum advantage is whether there are quantum advantage protocols with similar resource requirements as random circuit sampling that are also verifiable just from the classical outputs of the quantum computation. Here, we develop the idea of simulation secrets for verifiable advantage. A verifier can use a simulation secret to evaluate a cross-entropy test faster than it would take a classical adversary to pass the test. We instantiate this idea using IQP circuits described by cubic polynomials with planted independent spaces. These correspond to the largest independent set in the orbit of a polynomial under the general linear group and yield a low-rank stabilizer decomposition of the corresponding state. We conjecture that large independent spaces are invisible to a computationally bounded adversary, and therefore they cannot exploit them to pass the protocol. A second conjecture regards the fine-grained complexity of producing samples that pass the cross-entropy test for uniformly random polynomials. Under these conjectures, our scheme results in a polynomial gap between the verification time and the time a classical adversary would need to pass the protocol---both are exponential. It has a potential application to generating classically certifiable randomness, since the output distributions have high min-entropy. We estimate that the planted polynomial scheme is implementable using 100 logical qubits at logical error rates around $10^{-6}$.

Improved quantum volume estimation with transducers and amortized quantum walks

from arXiv: Data Structures and Algorithms

Authors: Arjan Cornelissen, Simon Apers, Sander Gribling

The volume estimation problem is a classic task in computational geometry. The development of randomized algorithms for this problem spurred the development of many influential algorithmic techniques related to Markov Chain Monte Carlo and simulated annealing, and the problem connects to several important geometrical results, like the recently-resolved KLS conjecture. In this work, we quantize the state-of-the-art $\widetilde{O}(d^{3.5}+d^3/\varepsilon^2)$-query randomized algorithm developed by Cousins and Vempala, and obtain a $\widetilde{O}(d^{3.5} + d^{1.75}/\varepsilon)$-query quantum algorithm, improving over the $\widetilde{O}(d^{3.5} + d^{2.25}/\varepsilon)$ state-of-the-art bound. Our key technical contribution is a framework for amortizing the cost of a quantum walk. The framework is based on the recent transducer toolkit introduced by Belovs, Jeffery and Yolcu. It is this amortized quantum walk framework that allows us to exploit the amortized analysis of the ball walk by Cousins and Vempala, thus overcoming the key barrier that previously barred its quantum implementation.

Authors: Arjan Cornelissen, Simon Apers, Sander Gribling

The volume estimation problem is a classic task in computational geometry. The development of randomized algorithms for this problem spurred the development of many influential algorithmic techniques related to Markov Chain Monte Carlo and simulated annealing, and the problem connects to several important geometrical results, like the recently-resolved KLS conjecture. In this work, we quantize the state-of-the-art $\widetilde{O}(d^{3.5}+d^3/\varepsilon^2)$-query randomized algorithm developed by Cousins and Vempala, and obtain a $\widetilde{O}(d^{3.5} + d^{1.75}/\varepsilon)$-query quantum algorithm, improving over the $\widetilde{O}(d^{3.5} + d^{2.25}/\varepsilon)$ state-of-the-art bound. Our key technical contribution is a framework for amortizing the cost of a quantum walk. The framework is based on the recent transducer toolkit introduced by Belovs, Jeffery and Yolcu. It is this amortized quantum walk framework that allows us to exploit the amortized analysis of the ball walk by Cousins and Vempala, thus overcoming the key barrier that previously barred its quantum implementation.

Polynomial-time algorithm for exact $(1,2)$-center problem under continuous Fréchet distance

from arXiv: Data Structures and Algorithms

Authors: Soumya Bhattacharya, Serene Rasheed, Sasanka Roy

In this paper, we explore the $(1,2)$-center problem for polygonal curves under continuous Fréchet distance. The $(k,\ell)$-center problem, in general, is known to be NP-hard. Aronov, Filtser, Horton, Katz, and Sheikhan (WADS'19) gave a polynomial-time algorithm for the $(1,2)$-center of curves in the plane under the discrete Fréchet distance. To the best of our knowledge, the $(1,2)$-center under continuous Fréchet distance has not been studied yet. We present a polynomial time algorithm to solve the problem exactly under both $\mathbb{L}_2$ and $\mathbb{L}_\infty$ norm, running in $O\bigr((n^2r+nr^2)^{2+ε}\bigl)$ time for curves in the plane where $r$ is the number of input curves and $n$ is the maximum complexity of any curve. Further, for curves in any dimension $d$, the expected time to compute the center using the algorithm is $O\bigr((n^2r+nr^2)^{2(d-1)+ε}\bigl)$. We have also shown that an $(1+\widetildeε)-$factor approximation of $(1,2)$-center can be computed in $O(n^2r+nr^2+1/ε^s)$ time for any $ε>\widetildeε>0$ and some constant $s$ for curves in the plane. For curves in the plane, we have shown that, with the center restricted to be horizontal, we can compute the exact center in $O(n^2r+nr^2)$ time. An algorithm has been introduced to find a $3$-factor approximation of the $(1,2)$-center in time linear in the number of curves. A formulation was introduced by de Berg, Mehrabi, and Ophelders (CCCG'17) to measure Fréchet distance between a curve and a query segment under $\mathbb{L}_2$ norm for curves in the plane. We have shown the formulation is valid under both $\mathbb{L}_2$ and $\mathbb{L}_\infty$ norm for curves in $\mathbb{R}^d$.

Authors: Soumya Bhattacharya, Serene Rasheed, Sasanka Roy

In this paper, we explore the $(1,2)$-center problem for polygonal curves under continuous Fréchet distance. The $(k,\ell)$-center problem, in general, is known to be NP-hard. Aronov, Filtser, Horton, Katz, and Sheikhan (WADS'19) gave a polynomial-time algorithm for the $(1,2)$-center of curves in the plane under the discrete Fréchet distance. To the best of our knowledge, the $(1,2)$-center under continuous Fréchet distance has not been studied yet. We present a polynomial time algorithm to solve the problem exactly under both $\mathbb{L}_2$ and $\mathbb{L}_\infty$ norm, running in $O\bigr((n^2r+nr^2)^{2+ε}\bigl)$ time for curves in the plane where $r$ is the number of input curves and $n$ is the maximum complexity of any curve. Further, for curves in any dimension $d$, the expected time to compute the center using the algorithm is $O\bigr((n^2r+nr^2)^{2(d-1)+ε}\bigl)$. We have also shown that an $(1+\widetildeε)-$factor approximation of $(1,2)$-center can be computed in $O(n^2r+nr^2+1/ε^s)$ time for any $ε>\widetildeε>0$ and some constant $s$ for curves in the plane. For curves in the plane, we have shown that, with the center restricted to be horizontal, we can compute the exact center in $O(n^2r+nr^2)$ time. An algorithm has been introduced to find a $3$-factor approximation of the $(1,2)$-center in time linear in the number of curves. A formulation was introduced by de Berg, Mehrabi, and Ophelders (CCCG'17) to measure Fréchet distance between a curve and a query segment under $\mathbb{L}_2$ norm for curves in the plane. We have shown the formulation is valid under both $\mathbb{L}_2$ and $\mathbb{L}_\infty$ norm for curves in $\mathbb{R}^d$.

Exponential quantum speedup for $\mathbb{F}_3^n$-Subset-Sum? Or, rigorous classical algorithms for Binary-Error LWE

from arXiv: Data Structures and Algorithms

Authors: Robin Kothari, Tony Metger, Ryan O'Donnell, Noah Shutty, Kewen Wu

We study vector subset sum over $\mathbb{F}_3^n$: given $m$ random vectors from $\mathbb{F}_3^n$, find a nonempty subset that sums to zero; the smaller $m$, the more difficult it is to find such a subset. Chen, Liu, and Zhandry (EUROCRYPT'22) introduced an efficient quantum algorithm that solves this problem when $m\approx n^2/2$, where a naive classical algorithm would require exponential time. Subsequently, Kothari, O'Donnell, and Wu (STOC'2026) gave an efficient classical algorithm that only requires $m \approx n^2/3$ vectors, thus removing the hope for an exponential quantum advantage in this parameter regime. Using the framework of Chen, Liu, and Zhandry, we give quantum algorithms that require much fewer input vectors, renewing the possibility of an exponential quantum speedup: for any fixed $ε>0$, our quantum algorithm solves $\mathbb{F}_3$-subset sum in polynomial time with $m=ε\cdot n^2$ vectors. More generally, we establish a full sample--time tradeoff that interpolates between exponential and polynomial runtime. The main ingredient is a deterministic classical algorithm for the binary-error Learning-with-Errors problem, which is of independent cryptographic interest. For this, we rigorously establish a sample--time tradeoff that was predicted by earlier algebraic heuristics. For vector subset sums over larger fields, we also significantly improve classical algorithms in Kothari, O'Donnell, and Wu (STOC'2026).

Authors: Robin Kothari, Tony Metger, Ryan O'Donnell, Noah Shutty, Kewen Wu

We study vector subset sum over $\mathbb{F}_3^n$: given $m$ random vectors from $\mathbb{F}_3^n$, find a nonempty subset that sums to zero; the smaller $m$, the more difficult it is to find such a subset. Chen, Liu, and Zhandry (EUROCRYPT'22) introduced an efficient quantum algorithm that solves this problem when $m\approx n^2/2$, where a naive classical algorithm would require exponential time. Subsequently, Kothari, O'Donnell, and Wu (STOC'2026) gave an efficient classical algorithm that only requires $m \approx n^2/3$ vectors, thus removing the hope for an exponential quantum advantage in this parameter regime. Using the framework of Chen, Liu, and Zhandry, we give quantum algorithms that require much fewer input vectors, renewing the possibility of an exponential quantum speedup: for any fixed $ε>0$, our quantum algorithm solves $\mathbb{F}_3$-subset sum in polynomial time with $m=ε\cdot n^2$ vectors. More generally, we establish a full sample--time tradeoff that interpolates between exponential and polynomial runtime. The main ingredient is a deterministic classical algorithm for the binary-error Learning-with-Errors problem, which is of independent cryptographic interest. For this, we rigorously establish a sample--time tradeoff that was predicted by earlier algebraic heuristics. For vector subset sums over larger fields, we also significantly improve classical algorithms in Kothari, O'Donnell, and Wu (STOC'2026).

Solving Sparse SDPs in Sublinear Time: A Classical Algorithm Inspired by the Quantum OR Lemma

from arXiv: Data Structures and Algorithms

Authors: Fernando G. S. L. Brandão, Alexander M. Dalzell, András Gilyén, Francisca Vasconcelos

We give the first sublinear-time classical solvers for sparse semidefinite programs in the bounded-radius regime, without low-rank assumptions or Frobenius norm dependence on the constraint matrices. For constant precision and bounded primal and dual radii, prior quantum algorithms of Brandão et al. (2019) and van Apeldoorn and Gilyén (2019) achieved $\widetilde{O}(\sqrt{n}+\sqrt{m})$ dependence on matrix dimension $n$ and constraint number $m$. Compared with the $\widetilde{O}(mn)$ runtime of existing classical methods, this suggests a quartic quantum speedup when $m \approx n$. Beyond a usual Grover speedup, this separation relies on the Quantum OR lemma, whose sample-reuse mechanism decouples the cost of Gibbs-state preparation from constraint search. We show that this reuse mechanism is classically realizable for sparse SDPs. Our main technical contribution is a classical procedure for simultaneously estimating many expectation values with respect to a sparse Hamiltonian's Gibbs state. This combines randomized Lánczos filtering with an efficient sampling-based estimator. We also introduce a stochastic online-learning framework for SDP solving, substantially improving accuracy-dependence over standard oracle-based MMWU approaches. Let $s$ denote the the input matrix sparsity and $γ:=Rr/\varepsilon$ capture dependence on the primal $(R)$ and dual $(r)$ radii as well as target accuracy $(\varepsilon)$. When $γ^2\leq\min\{m,n/s\}$, our solver runs in time $\widetilde{O}\left(nsγ^{4.5}+msγ^2\right)$. For $γ=O(1)$, this is $\widetilde{O}\left((n+m)s\right)$ and sublinear in the $O(mns)$ input size. Similar to the quantum algorithms, this matches known lower bounds with respect to $m$ and $n$, up to logarithmic factors. This implies that, with respect to dimensions $m$ and $n$, there is no super-quadratic quantum advantage for generic sparse SDP solving.

Authors: Fernando G. S. L. Brandão, Alexander M. Dalzell, András Gilyén, Francisca Vasconcelos

We give the first sublinear-time classical solvers for sparse semidefinite programs in the bounded-radius regime, without low-rank assumptions or Frobenius norm dependence on the constraint matrices. For constant precision and bounded primal and dual radii, prior quantum algorithms of Brandão et al. (2019) and van Apeldoorn and Gilyén (2019) achieved $\widetilde{O}(\sqrt{n}+\sqrt{m})$ dependence on matrix dimension $n$ and constraint number $m$. Compared with the $\widetilde{O}(mn)$ runtime of existing classical methods, this suggests a quartic quantum speedup when $m \approx n$. Beyond a usual Grover speedup, this separation relies on the Quantum OR lemma, whose sample-reuse mechanism decouples the cost of Gibbs-state preparation from constraint search. We show that this reuse mechanism is classically realizable for sparse SDPs. Our main technical contribution is a classical procedure for simultaneously estimating many expectation values with respect to a sparse Hamiltonian's Gibbs state. This combines randomized Lánczos filtering with an efficient sampling-based estimator. We also introduce a stochastic online-learning framework for SDP solving, substantially improving accuracy-dependence over standard oracle-based MMWU approaches. Let $s$ denote the the input matrix sparsity and $γ:=Rr/\varepsilon$ capture dependence on the primal $(R)$ and dual $(r)$ radii as well as target accuracy $(\varepsilon)$. When $γ^2\leq\min\{m,n/s\}$, our solver runs in time $\widetilde{O}\left(nsγ^{4.5}+msγ^2\right)$. For $γ=O(1)$, this is $\widetilde{O}\left((n+m)s\right)$ and sublinear in the $O(mns)$ input size. Similar to the quantum algorithms, this matches known lower bounds with respect to $m$ and $n$, up to logarithmic factors. This implies that, with respect to dimensions $m$ and $n$, there is no super-quadratic quantum advantage for generic sparse SDP solving.

Mixing FM-indexes and CSAs: backward search over an order-1 rank encoding

from arXiv: Data Structures and Algorithms

Authors: Travis Gagie

FM-indexes and compressed suffix arrays (CSAs) are often treated as interchangeable, but they behave differently as the alphabet grows. An FM-index step costs about one cache miss per level of a wavelet tree, so it gets slower with the alphabet size. A CSA step is a binary search whose range shrinks as characters get rarer. We describe a simple hybrid. Each character of the text is replaced by the rank of its frequency among the characters that follow the previous character. We backward-search on this encoding, which is over a small, skewed alphabet, and recover the one piece of information the encoding loses (the first character of the pattern) with a single CSA-like step on an array we call $\PsiE$. Counting is exact, and locating works with standard suffix-array sampling. A prototype on synthetic repetitive data shows that the hybrid is the fastest of the indexes we tried at intermediate alphabet sizes with 1\% noise, but even its compact version is 1.7 to 3.9 times larger than a compressed run-length CSA or FM-index of the original text, because the encoding and $\PsiE$ together have more runs than the original Burrows--Wheeler transform. Whether that changes on real data, such as parses and minimizer digests, is the main open question.

Authors: Travis Gagie

FM-indexes and compressed suffix arrays (CSAs) are often treated as interchangeable, but they behave differently as the alphabet grows. An FM-index step costs about one cache miss per level of a wavelet tree, so it gets slower with the alphabet size. A CSA step is a binary search whose range shrinks as characters get rarer. We describe a simple hybrid. Each character of the text is replaced by the rank of its frequency among the characters that follow the previous character. We backward-search on this encoding, which is over a small, skewed alphabet, and recover the one piece of information the encoding loses (the first character of the pattern) with a single CSA-like step on an array we call $\PsiE$. Counting is exact, and locating works with standard suffix-array sampling. A prototype on synthetic repetitive data shows that the hybrid is the fastest of the indexes we tried at intermediate alphabet sizes with 1\% noise, but even its compact version is 1.7 to 3.9 times larger than a compressed run-length CSA or FM-index of the original text, because the encoding and $\PsiE$ together have more runs than the original Burrows--Wheeler transform. Whether that changes on real data, such as parses and minimizer digests, is the main open question.

Conditioning-Free Non-Uniform Quantum Fourier and Chebyshev Transforms

from arXiv: Data Structures and Algorithms

Authors: Chaowen Guan, Akshit Katiyar

We present an efficient quantum algorithm for the non-uniform Chebyshev transform. It is defined as the projection of a function onto Chebyshev polynomials sampled at given nodes that are uniform in $x\in[-1,1]$, and hence non-uniform in the angle $θ=\arccos x$, a setting that QFT-based quantum Chebyshev transforms cannot handle. Our construction is based on an improvement of an existing Non-uniform Quantum Fourier Transform (NUQFT) whereby we remove the conditioning from non-uniform node sampling. Hence, error bounds are independent of the geometry-dependent parameter $κ$ of prior work. We use the fact that Chebyshev transform matrix is the average of two Type-II non-uniform DFTs, which we implement with a single controlled NUQFT circuit. We provide explicit oracle constructions, including the row-access oracle previously left as an assumption. The resulting $\varepsilon$-accurate block encoding has $O(1)$ normalization and uses $O(L)$ qubits and $\widetilde O(L^2)$ gates, where $L=\log N+\log(1/\varepsilon)$. We give an end-to-end implementation with complexity analysis, including the success probability and output-state error.

Authors: Chaowen Guan, Akshit Katiyar

We present an efficient quantum algorithm for the non-uniform Chebyshev transform. It is defined as the projection of a function onto Chebyshev polynomials sampled at given nodes that are uniform in $x\in[-1,1]$, and hence non-uniform in the angle $θ=\arccos x$, a setting that QFT-based quantum Chebyshev transforms cannot handle. Our construction is based on an improvement of an existing Non-uniform Quantum Fourier Transform (NUQFT) whereby we remove the conditioning from non-uniform node sampling. Hence, error bounds are independent of the geometry-dependent parameter $κ$ of prior work. We use the fact that Chebyshev transform matrix is the average of two Type-II non-uniform DFTs, which we implement with a single controlled NUQFT circuit. We provide explicit oracle constructions, including the row-access oracle previously left as an assumption. The resulting $\varepsilon$-accurate block encoding has $O(1)$ normalization and uses $O(L)$ qubits and $\widetilde O(L^2)$ gates, where $L=\log N+\log(1/\varepsilon)$. We give an end-to-end implementation with complexity analysis, including the success probability and output-state error.

Superlinear Quantum Query Lower Bounds for Subgraph Detection

from arXiv: Data Structures and Algorithms

Authors: Amin Shiraz Gilani, Xingyu Zhou

Subgraph detection asks whether an $n$-vertex graph, accessed through queries to its adjacency matrix, contains a copy of a fixed graph $H$. We prove the first unconditional superlinear lower bounds on the bounded-error quantum query complexity of this problem, answering a longstanding open question. A copy of $H$ is a certificate of constant size, so the adversary method with nonnegative weights cannot prove superlinear lower bounds. For every fixed $r\ge 4$, detecting the clique $K_r$ requires $n^{λ_r-o(1)}$ queries, where $λ_4=19/18$, the exponents $λ_r$ increase strictly with $r$, and $λ_r\ge 2-4\sqrt{2/r}+O(1/r)$. More generally, we prove superlinear lower bounds for detecting every fixed connected graph $H$ with chromatic number $c\ge 4$. These bounds approach quadratic as $c$ grows: for sufficiently large $c$, detection requires $n^{2-O(\sqrt{\log\log c/c})-o(1)}$ queries. Chromatic number alone does not characterize the quantum query complexity of subgraph detection: we show that detecting the complete bipartite graph $K_{r,r}$ requires $n^{β_r-o(1)}$ queries, where $β_{10}=181/180$ and $β_r\ge 2-O(1/\sqrt{r})$. Our main technical result is a lower bound for finding an all-ones certificate from a known family when the input bits are sampled independently. Its proof combines Zhandry's compressed oracle (CRYPTO 2019) with conditioning on a randomly planted certificate, adapting an argument of Belovs (FOCS 2026). Our hard instances are built from graphs containing many copies of the desired subgraph with limited overlap. For cliques, we use a construction of Gowers and Janzer (CPC 2021); for complete bipartite graphs, we use a random construction.

Authors: Amin Shiraz Gilani, Xingyu Zhou

Subgraph detection asks whether an $n$-vertex graph, accessed through queries to its adjacency matrix, contains a copy of a fixed graph $H$. We prove the first unconditional superlinear lower bounds on the bounded-error quantum query complexity of this problem, answering a longstanding open question. A copy of $H$ is a certificate of constant size, so the adversary method with nonnegative weights cannot prove superlinear lower bounds. For every fixed $r\ge 4$, detecting the clique $K_r$ requires $n^{λ_r-o(1)}$ queries, where $λ_4=19/18$, the exponents $λ_r$ increase strictly with $r$, and $λ_r\ge 2-4\sqrt{2/r}+O(1/r)$. More generally, we prove superlinear lower bounds for detecting every fixed connected graph $H$ with chromatic number $c\ge 4$. These bounds approach quadratic as $c$ grows: for sufficiently large $c$, detection requires $n^{2-O(\sqrt{\log\log c/c})-o(1)}$ queries. Chromatic number alone does not characterize the quantum query complexity of subgraph detection: we show that detecting the complete bipartite graph $K_{r,r}$ requires $n^{β_r-o(1)}$ queries, where $β_{10}=181/180$ and $β_r\ge 2-O(1/\sqrt{r})$. Our main technical result is a lower bound for finding an all-ones certificate from a known family when the input bits are sampled independently. Its proof combines Zhandry's compressed oracle (CRYPTO 2019) with conditioning on a randomly planted certificate, adapting an argument of Belovs (FOCS 2026). Our hard instances are built from graphs containing many copies of the desired subgraph with limited overlap. For cliques, we use a construction of Gowers and Janzer (CPC 2021); for complete bipartite graphs, we use a random construction.

Quantum oblique eigenprojection

from arXiv: Data Structures and Algorithms

Authors: Alexander M. Dalzell, Yuan Su

Every square matrix decomposes its underlying Hilbert space into generalized, nonorthogonal eigensubspaces. We show that a quantum computer can perform such an oblique eigenprojection $Π$ given block encoding access to the input matrix. Our approach has a query complexity nearly linear in the inverse gap and a normalization factor close to $\lVertΠ\rVert$ under a spectral-set condition on the input. This covers common assumptions on the numerical range or diagonalizability and matches known results for orthogonal eigenprojections. We achieve this with a two-sided block preconditioning that uses a discrete Fourier transform of the matrix resolvent. We describe applications to: (i) preparing eigenstates of matrices with complex eigenvalues, extending the quantum eigenvalue transformation algorithm of Low and Su beyond real spectra; (ii) solving continuous-time algebraic Riccati equations, cubically speeding up a prior solver of Rodenas-Ruiz, Zhao, and Lee; and (iii) solving ordinary Sylvester equations, quadratically improving a direct augmented method of Wang and Liu. Our result suggests a promising route to applying nonanalytic matrix functions on quantum computers.

Authors: Alexander M. Dalzell, Yuan Su

Every square matrix decomposes its underlying Hilbert space into generalized, nonorthogonal eigensubspaces. We show that a quantum computer can perform such an oblique eigenprojection $Π$ given block encoding access to the input matrix. Our approach has a query complexity nearly linear in the inverse gap and a normalization factor close to $\lVertΠ\rVert$ under a spectral-set condition on the input. This covers common assumptions on the numerical range or diagonalizability and matches known results for orthogonal eigenprojections. We achieve this with a two-sided block preconditioning that uses a discrete Fourier transform of the matrix resolvent. We describe applications to: (i) preparing eigenstates of matrices with complex eigenvalues, extending the quantum eigenvalue transformation algorithm of Low and Su beyond real spectra; (ii) solving continuous-time algebraic Riccati equations, cubically speeding up a prior solver of Rodenas-Ruiz, Zhao, and Lee; and (iii) solving ordinary Sylvester equations, quadratically improving a direct augmented method of Wang and Liu. Our result suggests a promising route to applying nonanalytic matrix functions on quantum computers.

Dynamic Time Warping in the Low-Distance Regime

from arXiv: Data Structures and Algorithms

Authors: Itai Boneh, Shay Golan, Tomasz Kociumaka

Dynamic Time Warping (DTW) is a classical similarity measure for strings and time series that allows local stretching. Given non-empty strings $S,T$ over an alphabet $Σ$ and a cost function $δ:Σ^2\to\mathbb{R}_{\ge0}$, $DTW_δ(S,T)$ is the minimum total cost of equal-length expansions of $S$ and $T$ obtained by duplicating characters. For strings of length at most $n$, DTW is computable in $O(n^2)$ time, and this is conditionally optimal under the Orthogonal Vectors Hypothesis (OVH). We study the low-distance regime, where an integer $k$ upper-bounds $DTW_δ(S,T)$, assuming $δ(a,a)=0$ and $δ(a,b)\ge1$ for $a\ne b$. For several classical similarity measures, this regime admits $O(n+\operatorname{poly}(k))$ algorithms, whereas for DTW with metric costs the best known bound is $O(nk)$. We show that this dependence is essentially optimal: assuming OVH, computing DTW requires $n^{1-o(1)}k$ time even for the discrete mismatch-cost function, which assigns cost $1$ to every mismatch. The lower bound applies to the whole spectrum of thresholds $k$ between constant and linear in $n$. Our reduction from Orthogonal Vectors encodes vector coordinates in the lengths of equal-character runs. The resulting instances are very structured: collapsing runs to single characters reveals long substrings with short periods. We complement the lower bound with a $\tilde O(n+\operatorname{poly}(k))$-time algorithm whenever, after collapsing runs in the inputs, every substring with period $O(k)$ has length $\operatorname{poly}(k)$. Finally, we extend this lower bound to DTW pattern matching, which asks whether any non-empty substring of a length-$n$ text has DTW distance at most $k$ from a length-$m$ pattern. We prove that the classic $O(nm)$-time dynamic-programming algorithm is near-optimal under OVH, even when $k=O(\log n)$.

Authors: Itai Boneh, Shay Golan, Tomasz Kociumaka

Dynamic Time Warping (DTW) is a classical similarity measure for strings and time series that allows local stretching. Given non-empty strings $S,T$ over an alphabet $Σ$ and a cost function $δ:Σ^2\to\mathbb{R}_{\ge0}$, $DTW_δ(S,T)$ is the minimum total cost of equal-length expansions of $S$ and $T$ obtained by duplicating characters. For strings of length at most $n$, DTW is computable in $O(n^2)$ time, and this is conditionally optimal under the Orthogonal Vectors Hypothesis (OVH). We study the low-distance regime, where an integer $k$ upper-bounds $DTW_δ(S,T)$, assuming $δ(a,a)=0$ and $δ(a,b)\ge1$ for $a\ne b$. For several classical similarity measures, this regime admits $O(n+\operatorname{poly}(k))$ algorithms, whereas for DTW with metric costs the best known bound is $O(nk)$. We show that this dependence is essentially optimal: assuming OVH, computing DTW requires $n^{1-o(1)}k$ time even for the discrete mismatch-cost function, which assigns cost $1$ to every mismatch. The lower bound applies to the whole spectrum of thresholds $k$ between constant and linear in $n$. Our reduction from Orthogonal Vectors encodes vector coordinates in the lengths of equal-character runs. The resulting instances are very structured: collapsing runs to single characters reveals long substrings with short periods. We complement the lower bound with a $\tilde O(n+\operatorname{poly}(k))$-time algorithm whenever, after collapsing runs in the inputs, every substring with period $O(k)$ has length $\operatorname{poly}(k)$. Finally, we extend this lower bound to DTW pattern matching, which asks whether any non-empty substring of a length-$n$ text has DTW distance at most $k$ from a length-$m$ pattern. We prove that the classic $O(nm)$-time dynamic-programming algorithm is near-optimal under OVH, even when $k=O(\log n)$.

Improved Quantum Query Bounds for Boolean Matrix Product Verification

from arXiv: Data Structures and Algorithms

Authors: Amin Shiraz Gilani, François Le Gall, Xingyu Zhou

We prove the first non-trivial upper bound for the quantum query complexity of Boolean Matrix Product Verification ($\mathsf{BMPV}$), answering a longstanding open question in quantum query complexity. For $n\times n$ matrices, our upper bound is $\widetilde O(n^{17/12})$, improving on the standard $O(n^{3/2})$ bound obtained using Grover search by Buhrman and Špalek (SODA 2006). We complement this result by showing an $Ω(n^{5/4})$ lower bound, which improves over the previous best known lower bound of $\widetildeΩ(n^{19/18})$ by Childs, Kimmel, and Kothari (ESA 2012). Our approach centers on a connection with Orthogonal Vectors ($\mathsf{OV}$), which asks whether an indexed list of $n$ Boolean vectors of dimension $n$ contains two vectors with disjoint supports. In particular, we prove equivalences between $\mathsf{OV}$ and $\mathsf{BMPV}$ and establish the above bounds for $\mathsf{OV}$. We also prove a tight $\widetilde Θ(n^{3/2})$ bound for a variant of $\mathsf{BMPV}$ that asks whether the product contains a given row vector. Together, these results imply a polynomial separation between the quantum query complexities of deciding whether a graph has radius at most two and whether it has diameter at most two.

Authors: Amin Shiraz Gilani, François Le Gall, Xingyu Zhou

We prove the first non-trivial upper bound for the quantum query complexity of Boolean Matrix Product Verification ($\mathsf{BMPV}$), answering a longstanding open question in quantum query complexity. For $n\times n$ matrices, our upper bound is $\widetilde O(n^{17/12})$, improving on the standard $O(n^{3/2})$ bound obtained using Grover search by Buhrman and Špalek (SODA 2006). We complement this result by showing an $Ω(n^{5/4})$ lower bound, which improves over the previous best known lower bound of $\widetildeΩ(n^{19/18})$ by Childs, Kimmel, and Kothari (ESA 2012). Our approach centers on a connection with Orthogonal Vectors ($\mathsf{OV}$), which asks whether an indexed list of $n$ Boolean vectors of dimension $n$ contains two vectors with disjoint supports. In particular, we prove equivalences between $\mathsf{OV}$ and $\mathsf{BMPV}$ and establish the above bounds for $\mathsf{OV}$. We also prove a tight $\widetilde Θ(n^{3/2})$ bound for a variant of $\mathsf{BMPV}$ that asks whether the product contains a given row vector. Together, these results imply a polynomial separation between the quantum query complexities of deciding whether a graph has radius at most two and whether it has diameter at most two.

Shadow Quantum Singular Value Transformation with Shallow Quantum Circuits

from arXiv: Data Structures and Algorithms

Authors: Nai-Hui Chia, Hyunseong Kim, Chia-Ying Lin

We introduce shadow quantum singular value transformation (Shadow QSVT): given an initial state $|ψ\rangle$, a Hermitian matrix $H$, a polynomial $f$, and a set of observables $\{O_1,\dots,O_m\}$, the goal is to estimate $\langleψ|f(H)^{\dagger}O_j f(H)|ψ\rangle$ for all $j\in\{1,\dots,m\}$. Shadow QSVT provides a systematic route to reduce the quantum resources required by standard QSVT, which constructs a unitary block-encoding of $f(H)$. It uses structure in the input state and observables, together with the fact that many applications require only observable estimates rather than synthesizing the full unitary. We present three algorithms that exploit structure in the initial state and observables to reduce quantum circuit depth. First, we develop a state-aware QSVT algorithm that prepares the target state with low circuit depth when the Krylov subspace associated with $H$ and $|ψ\rangle$ is low-dimensional or admits an accurate low-dimensional approximation. Second, we introduce an observable-aware Shadow QSVT algorithm that combines a new observable-aware Krylov subspace with history states to further reduce circuit depth and gate complexity. Finally, we develop Classical Shadow QSVT, which constructs a classical representation from $H$, $f$, and $|ψ\rangle$ without prior knowledge of the observables or explicit preparation of the target state proportional to $f(H)|ψ\rangle$. This representation enables estimation of the target quantities for observables specified after the quantum computation. Together, these three algorithms provide tools for reducing the circuit depth of QSVT-based computations across a range of settings.

Authors: Nai-Hui Chia, Hyunseong Kim, Chia-Ying Lin

We introduce shadow quantum singular value transformation (Shadow QSVT): given an initial state $|ψ\rangle$, a Hermitian matrix $H$, a polynomial $f$, and a set of observables $\{O_1,\dots,O_m\}$, the goal is to estimate $\langleψ|f(H)^{\dagger}O_j f(H)|ψ\rangle$ for all $j\in\{1,\dots,m\}$. Shadow QSVT provides a systematic route to reduce the quantum resources required by standard QSVT, which constructs a unitary block-encoding of $f(H)$. It uses structure in the input state and observables, together with the fact that many applications require only observable estimates rather than synthesizing the full unitary. We present three algorithms that exploit structure in the initial state and observables to reduce quantum circuit depth. First, we develop a state-aware QSVT algorithm that prepares the target state with low circuit depth when the Krylov subspace associated with $H$ and $|ψ\rangle$ is low-dimensional or admits an accurate low-dimensional approximation. Second, we introduce an observable-aware Shadow QSVT algorithm that combines a new observable-aware Krylov subspace with history states to further reduce circuit depth and gate complexity. Finally, we develop Classical Shadow QSVT, which constructs a classical representation from $H$, $f$, and $|ψ\rangle$ without prior knowledge of the observables or explicit preparation of the target state proportional to $f(H)|ψ\rangle$. This representation enables estimation of the target quantities for observables specified after the quantum computation. Together, these three algorithms provide tools for reducing the circuit depth of QSVT-based computations across a range of settings.

Policy Iteration Is Not Strongly Polynomial for Deterministic Markov Decision Processes: The Price of Algorithmic Anarchy

from arXiv: Data Structures and Algorithms

Authors: Han Zhong, Yinyu Ye

We establish an exponential iteration lower bound in the number of states for Howard's policy iteration on deterministic discounted Markov decision processes, with at most two actions per state. This rules out strong polynomiality of Howard's policy iteration when the discount factor is part of the input and yields an exponential separation from the simplex method with Dantzig's pivoting rule, which is proved to be strongly polynomial on this class. Even when each reward is restricted to logarithmic bit length, we obtain a stretched-exponential iteration lower bound. The gap between Howard's decentralized and simultaneous selfish improvements and Dantzig's coordinated selection of a single action with the largest gain across all states reveals a ``price'' of algorithmic anarchy.

Authors: Han Zhong, Yinyu Ye

We establish an exponential iteration lower bound in the number of states for Howard's policy iteration on deterministic discounted Markov decision processes, with at most two actions per state. This rules out strong polynomiality of Howard's policy iteration when the discount factor is part of the input and yields an exponential separation from the simplex method with Dantzig's pivoting rule, which is proved to be strongly polynomial on this class. Even when each reward is restricted to logarithmic bit length, we obtain a stretched-exponential iteration lower bound. The gap between Howard's decentralized and simultaneous selfish improvements and Dantzig's coordinated selection of a single action with the largest gain across all states reveals a ``price'' of algorithmic anarchy.

Sharper Gaussian Covers in the Bansal--Huang--Lee Coloring Framework

from arXiv: Data Structures and Algorithms

Authors: Konstantin Makarychev, Yury Makarychev

Bansal, Huang, and Lee recently gave a polynomial-time algorithm that colors every $3$-colorable graph on $n$ vertices with $O(n^{0.19539})$ colors. We improve their bound to $O(n^{0.17794})$ colors. The improvement comes from a sharper rule for combining Gaussian covers. We prove the rule from Ehrhard's inequality by comparing the probability that a Gaussian vector misses a polyhedron at two thresholds. At the farther threshold a union bound records how many halfspaces define the polyhedron, and this information reduces the loss in the combination. We then organize the successive neighborhood steps of the Bansal-Huang-Lee argument into a single recursion. The smaller loss lets the recursion run one step farther than before, and at that step it would require more distinct vertices than the graph contains. This rules out the remaining case. Combining the resulting bounded-degree guarantee with the dense-graph algorithm of Kawarabayashi, Thorup, and Yoneda gives a randomized polynomial-time algorithm that colors every $3$-colorable graph on $n$ vertices with $O(n^{0.17794})$ colors.

Authors: Konstantin Makarychev, Yury Makarychev

Bansal, Huang, and Lee recently gave a polynomial-time algorithm that colors every $3$-colorable graph on $n$ vertices with $O(n^{0.19539})$ colors. We improve their bound to $O(n^{0.17794})$ colors. The improvement comes from a sharper rule for combining Gaussian covers. We prove the rule from Ehrhard's inequality by comparing the probability that a Gaussian vector misses a polyhedron at two thresholds. At the farther threshold a union bound records how many halfspaces define the polyhedron, and this information reduces the loss in the combination. We then organize the successive neighborhood steps of the Bansal-Huang-Lee argument into a single recursion. The smaller loss lets the recursion run one step farther than before, and at that step it would require more distinct vertices than the graph contains. This rules out the remaining case. Combining the resulting bounded-degree guarantee with the dense-graph algorithm of Kawarabayashi, Thorup, and Yoneda gives a randomized polynomial-time algorithm that colors every $3$-colorable graph on $n$ vertices with $O(n^{0.17794})$ colors.

Super-Quadratic Quantum Speedups for Combinatorial Optimization via Tilted Walks

from arXiv: Data Structures and Algorithms

Authors: Guneykan Ozgul, Shouvanik Chakrabarti

We introduce quantum tilted walks, a quantum algorithmic framework for solving exact combinatorial optimization problems. The framework applies an average of powers of a tilted Hamiltonian that biases the discriminant matrix of a base Markov chain (mixer) with the objective function. Our starting point is quantum short-path algorithms, which prepare the ground state of such a Hamiltonian and obtain super-quadratic speedups over exhaustive search for certain combinatorial optimization problems. Recently, Le Gall and Tamaki~(arXiv:2604.12131) developed a classical conditioning-and-search algorithm for weighted MAX-E$k$-LIN2 and weighted MAX-$k$-CSP. Under the same assumptions, their algorithm is only sub-quadratically slower than quantum short-path algorithms. Consequently, existing short-path algorithms do not establish a super-quadratic speedup over this stronger classical baseline. For maximization problems, we give conditions under which tilted walks increase the amplitude on the target state with high objective value when initialized from a starting state with lower objective value. This framework captures conditioning-and-search and yields super-quadratic speedups over it for the same problems. While our framework recovers quantum short-path algorithms as a special case, it neither requires ground-state preparation nor initialization in the ground state of the base mixer. We demonstrate these advantages on a synthetic optimization problem for which tilted walks achieve a super-quadratic speedup whereas the short-path algorithms do not.

Authors: Guneykan Ozgul, Shouvanik Chakrabarti

We introduce quantum tilted walks, a quantum algorithmic framework for solving exact combinatorial optimization problems. The framework applies an average of powers of a tilted Hamiltonian that biases the discriminant matrix of a base Markov chain (mixer) with the objective function. Our starting point is quantum short-path algorithms, which prepare the ground state of such a Hamiltonian and obtain super-quadratic speedups over exhaustive search for certain combinatorial optimization problems. Recently, Le Gall and Tamaki~(arXiv:2604.12131) developed a classical conditioning-and-search algorithm for weighted MAX-E$k$-LIN2 and weighted MAX-$k$-CSP. Under the same assumptions, their algorithm is only sub-quadratically slower than quantum short-path algorithms. Consequently, existing short-path algorithms do not establish a super-quadratic speedup over this stronger classical baseline. For maximization problems, we give conditions under which tilted walks increase the amplitude on the target state with high objective value when initialized from a starting state with lower objective value. This framework captures conditioning-and-search and yields super-quadratic speedups over it for the same problems. While our framework recovers quantum short-path algorithms as a special case, it neither requires ground-state preparation nor initialization in the ground state of the base mixer. We demonstrate these advantages on a synthetic optimization problem for which tilted walks achieve a super-quadratic speedup whereas the short-path algorithms do not.

Robust and Learned Online Matching in Growing Trees

from arXiv: Data Structures and Algorithms

Authors: Marek Gałązka, Hanna Wdowicka

We study irrevocable maximum-cardinality matching in trees revealed by successive leaf attachments, with a known horizon and an exogenous growth law that is misspecified or unknown. For deterministic affine attachment forecasts with nonnegative degree reinforcement, the optimal threshold policy loses at most twice the cumulative expected conditional total-variation error relative to an online oracle knowing the actual growth law. This follows from a unit-span property of the Bellman continuation score and has no additional horizon factor. A four-vertex example attains the coefficient two for the specified deterministic policy, and a two-model argument gives a lower bound linear in the model-error budget for arbitrary policies under general misspecification. For uniform-preferential attachment, the local error has an exact expression through the leaf count. When its constant mixture parameter is unknown, we estimate it from the same growing tree and update the threshold policy at geometric times. A parameter-sensitivity bound for individual Bellman prices and uniform degree-moment estimates yield expected regret $O(\sqrt{n}\log^2 n)$, using $O(n^2\log n)$ arithmetic operations and $O(n)$ stored entries. The exact minimax rate remains open.

Authors: Marek Gałązka, Hanna Wdowicka

We study irrevocable maximum-cardinality matching in trees revealed by successive leaf attachments, with a known horizon and an exogenous growth law that is misspecified or unknown. For deterministic affine attachment forecasts with nonnegative degree reinforcement, the optimal threshold policy loses at most twice the cumulative expected conditional total-variation error relative to an online oracle knowing the actual growth law. This follows from a unit-span property of the Bellman continuation score and has no additional horizon factor. A four-vertex example attains the coefficient two for the specified deterministic policy, and a two-model argument gives a lower bound linear in the model-error budget for arbitrary policies under general misspecification. For uniform-preferential attachment, the local error has an exact expression through the leaf count. When its constant mixture parameter is unknown, we estimate it from the same growing tree and update the threshold policy at geometric times. A parameter-sensitivity bound for individual Bellman prices and uniform degree-moment estimates yield expected regret $O(\sqrt{n}\log^2 n)$, using $O(n^2\log n)$ arithmetic operations and $O(n)$ stored entries. The exact minimax rate remains open.

Component-Weighted Centroid Search for Exact Incremental BPE

from arXiv: Data Structures and Algorithms

Authors: Harshit Verma, Rex Ying

Exact incremental BPE maintains the canonical tokenization state after every appended byte. The recent algorithm of Jiang and Gong (2026) does this in $O(\log^2 t)$ worst-case time, where $t$ is the maximum canonical token length. Its centroid search visits $O(\log t)$ components and can pay another $O(\log t)$ for ordered point location at each one. Within Jiang and Gong's normalized/proper merge-stage model, we change only that local search. Each interval is weighted by the size of the recursive component it selects, so a move from size $m$ to size $m'$ costs $O(1+\log(m/m'))$. These charges telescope, giving $O(\log t)$ time per append and $O(n\log t)$ over an $n$-byte stream, with the same BPE semantics and asymptotic space. We also construct a normalized proper BPE family over a fixed alphabet where count-balanced search uses $Θ(\log^2 t)$ probes on a reachable update, while the weighted search uses $Θ(\log t)$. A Rust implementation matches the predicted probe counts on every tested instance. On ordinary vocabularies the queried degrees are small, however, and the improvement is a worst-case guarantee rather than an average-speed result.

Authors: Harshit Verma, Rex Ying

Exact incremental BPE maintains the canonical tokenization state after every appended byte. The recent algorithm of Jiang and Gong (2026) does this in $O(\log^2 t)$ worst-case time, where $t$ is the maximum canonical token length. Its centroid search visits $O(\log t)$ components and can pay another $O(\log t)$ for ordered point location at each one. Within Jiang and Gong's normalized/proper merge-stage model, we change only that local search. Each interval is weighted by the size of the recursive component it selects, so a move from size $m$ to size $m'$ costs $O(1+\log(m/m'))$. These charges telescope, giving $O(\log t)$ time per append and $O(n\log t)$ over an $n$-byte stream, with the same BPE semantics and asymptotic space. We also construct a normalized proper BPE family over a fixed alphabet where count-balanced search uses $Θ(\log^2 t)$ probes on a reachable update, while the weighted search uses $Θ(\log t)$. A Rust implementation matches the predicted probe counts on every tested instance. On ordinary vocabularies the queried degrees are small, however, and the improvement is a worst-case guarantee rather than an average-speed result.

Learning Random Quantum Circuits and the Emergence of Pseudorandomness

from arXiv: Data Structures and Algorithms

Authors: Srinivasan Arunachalam, Qizhao Huang, Makrand Sinha

We give an efficient algorithm for learning $k$-dimensional brickwork random quantum circuits using only copies of the output state obtained by applying $U$ to the all-zero input. For a depth-$d$ circuit on $n$ sites with random $2\ell$-qubit gates, the algorithm learns the original circuit $U$ with high probability in $\text{poly}(n,2^{\ell d})$ time for every constant dimensional lattice. In particular, the algorithm is polynomial time as long as $\ell d = O(\log n)$. In one dimension, this reaches the natural boundary suggested by pseudorandomness: pseudorandom states require $\ell d=ω(\log n)$, and structured cryptographic constructions suggest that this scale may be achievable from above. In higher dimensions, it remains plausible that the same $\ell d=ω(\log n)$ scale remains the threshold for ancilla-free pseudorandomness, and our results help clarify the conditions under which pseudorandomness can arise in this setting. The main idea behind our algorithm is a local correlation criterion that identifies gates in the final layer without learning their entire backward light cones, avoiding a bottleneck in the previous approaches. A key technical ingredient is a Carbery-Wright type anticoncentration inequality for low-degree polynomials of Haar random unitaries whose small-ball exponent is independent of the matrix dimension. The dimension-independent exponent is crucial for handling gates of growing locality. This anticoncentration result may also be of independent interest.

Authors: Srinivasan Arunachalam, Qizhao Huang, Makrand Sinha

We give an efficient algorithm for learning $k$-dimensional brickwork random quantum circuits using only copies of the output state obtained by applying $U$ to the all-zero input. For a depth-$d$ circuit on $n$ sites with random $2\ell$-qubit gates, the algorithm learns the original circuit $U$ with high probability in $\text{poly}(n,2^{\ell d})$ time for every constant dimensional lattice. In particular, the algorithm is polynomial time as long as $\ell d = O(\log n)$. In one dimension, this reaches the natural boundary suggested by pseudorandomness: pseudorandom states require $\ell d=ω(\log n)$, and structured cryptographic constructions suggest that this scale may be achievable from above. In higher dimensions, it remains plausible that the same $\ell d=ω(\log n)$ scale remains the threshold for ancilla-free pseudorandomness, and our results help clarify the conditions under which pseudorandomness can arise in this setting. The main idea behind our algorithm is a local correlation criterion that identifies gates in the final layer without learning their entire backward light cones, avoiding a bottleneck in the previous approaches. A key technical ingredient is a Carbery-Wright type anticoncentration inequality for low-degree polynomials of Haar random unitaries whose small-ball exponent is independent of the matrix dimension. The dimension-independent exponent is crucial for handling gates of growing locality. This anticoncentration result may also be of independent interest.

Connected Dominating Set on Semi-Ladder-Free Graphs

from arXiv: Data Structures and Algorithms

Authors: Sobyasachi Chatterjee, Sushmita Gupta, Saket Saurabh, Sanjay Seetharaman, Anannya Upasana

We study \textsc{Connected Dominating Set} on graphs whose closed-neighborhood set systems are $d$-semi-ladder-free. This structural condition strictly generalizes the biclique-free setting and provides a natural regime for connectivity-constrained domination. We obtain both a fixed-parameter algorithm and an approximate kernelization framework for the problem on this class. Our algorithmic result is based on a new compact representation theorem for inclusion-wise minimal set covers in $d$-semi-ladder-free set systems. Although the number of minimal set covers of size at most $k$ may be as large as $n^{Ω(k)}$, we show that all such set covers can nevertheless be encoded by a family of at most $k^{kd+1}$ tuples, and that this family can be enumerated in time $\Oh(k^{kd+2}\cdot nm)$. Combining this representation with a \textsc{Group Steiner Tree} subroutine, we obtain an algorithm for \textsc{Connected Set Cover}, which in turn yields an algorithm for \textsc{Connected Dominating Set} running in time $k^{kd+2}\cdot 2^k \cdot n^{\Oh(1)}$ and polynomial space. For the preprocessing result, we introduce grouped domination cores and dominator cores, and prove polynomial upper bounds on their sizes in $d$-semi-ladder-free graphs. Using these structures, we obtain, for every fixed $d$ and $\varepsilon>0$, a polynomial-time $(1+\varepsilon)$-lossy compression for \textsc{Connected Dominating Set} to an equivalent reduced instance of size $k^{\Oh(d^2/\varepsilon)}$. The reduced instance is a \textsc{Connected Dominating Set} instance on a $(d+2)$-semi-ladder-free graph.

Authors: Sobyasachi Chatterjee, Sushmita Gupta, Saket Saurabh, Sanjay Seetharaman, Anannya Upasana

We study \textsc{Connected Dominating Set} on graphs whose closed-neighborhood set systems are $d$-semi-ladder-free. This structural condition strictly generalizes the biclique-free setting and provides a natural regime for connectivity-constrained domination. We obtain both a fixed-parameter algorithm and an approximate kernelization framework for the problem on this class. Our algorithmic result is based on a new compact representation theorem for inclusion-wise minimal set covers in $d$-semi-ladder-free set systems. Although the number of minimal set covers of size at most $k$ may be as large as $n^{Ω(k)}$, we show that all such set covers can nevertheless be encoded by a family of at most $k^{kd+1}$ tuples, and that this family can be enumerated in time $\Oh(k^{kd+2}\cdot nm)$. Combining this representation with a \textsc{Group Steiner Tree} subroutine, we obtain an algorithm for \textsc{Connected Set Cover}, which in turn yields an algorithm for \textsc{Connected Dominating Set} running in time $k^{kd+2}\cdot 2^k \cdot n^{\Oh(1)}$ and polynomial space. For the preprocessing result, we introduce grouped domination cores and dominator cores, and prove polynomial upper bounds on their sizes in $d$-semi-ladder-free graphs. Using these structures, we obtain, for every fixed $d$ and $\varepsilon>0$, a polynomial-time $(1+\varepsilon)$-lossy compression for \textsc{Connected Dominating Set} to an equivalent reduced instance of size $k^{\Oh(d^2/\varepsilon)}$. The reduced instance is a \textsc{Connected Dominating Set} instance on a $(d+2)$-semi-ladder-free graph.

Testing Induced-Subgraph Freeness in Outerplanar Graphs under the Random-Neighbor Oracle

from arXiv: Data Structures and Algorithms

Authors: Pan Peng, Kefan Yu

We prove that, for every fixed nonempty graph $H$, induced-$H$-freeness is testable with $\varepsilon^{-O_H(1)}$ queries on outerplanar graphs with no maximum-degree bound in the $\textit{random-neighbor model}$, where each query at a vertex returns a uniformly random neighbor. Thus, the query complexity is polynomial in $1/\varepsilon$ and independent of the number $n$ of vertices. Previously, the best bound known for this problem was the $\operatorname{poly}(\log n)$-query guarantee that follows from the general outerplanar-graph tester of Babu, Khoury, and Newman (2016) in the stronger $\textit{adjacency-list model}$, which provides exact degree queries and indexed access to neighbors. Our tester has $\textit{two-sided error}$, which is necessary in general: induced-$P_3$-freeness has no one-sided constant-query tester in the random-neighbor model, even on outerplanar graphs of maximum degree two.

Authors: Pan Peng, Kefan Yu

We prove that, for every fixed nonempty graph $H$, induced-$H$-freeness is testable with $\varepsilon^{-O_H(1)}$ queries on outerplanar graphs with no maximum-degree bound in the $\textit{random-neighbor model}$, where each query at a vertex returns a uniformly random neighbor. Thus, the query complexity is polynomial in $1/\varepsilon$ and independent of the number $n$ of vertices. Previously, the best bound known for this problem was the $\operatorname{poly}(\log n)$-query guarantee that follows from the general outerplanar-graph tester of Babu, Khoury, and Newman (2016) in the stronger $\textit{adjacency-list model}$, which provides exact degree queries and indexed access to neighbors. Our tester has $\textit{two-sided error}$, which is necessary in general: induced-$P_3$-freeness has no one-sided constant-query tester in the random-neighbor model, even on outerplanar graphs of maximum degree two.

$(α, β)$ Spanners and Hybrid Spanners with Nearly Tight Bounds

from arXiv: Data Structures and Algorithms

Authors: Shiri Chechik, Gur Lifshitz

For an $n$-vertex undirected, unweighted graph $G=(V,E)$ and a positive integer $k$, we present new spanner constructions with $O_k(n^{1+1/k})$ edges that achieve nearly optimal guarantees for all distances $d\le k$. Specifically, we construct a spanner $H\subseteq G$ with $O(n^{1+1/k}+(k+d\log d)n)$ edges, ensuring that any pair at original distance at most $d$ satisfies $\mathrm{dist}_H(u,v)\le 2k+O(d\log d)$. Equivalently, the multiplicative stretch for pairs at distance $d$ is $2k/d+O(\log d)$. In particular, setting $d=k/\log k$ yields an $(O(\log k),O(k))$-spanner with $O(n^{1+1/k}+kn)$ edges. For comparison, Ben-Levy and Parter (SODA'20) obtained, for every fixed $\varepsilon>0$ and sufficiently large $k$, an $(O(k^\varepsilon),O_\varepsilon(k))$-spanner with $O_{\varepsilon,k}(n^{1+1/k})$ edges. Our result improves the multiplicative stretch from $O(k^\varepsilon)$ to $O(\log k)$ while keeping the additive term linear in $k$, bringing us closer to the goal of $(O(1),O(k))$-spanners. Furthermore, Ben-Levy and Parter obtained multiplicative stretch $O_\varepsilon(k/d)$ for distances $d\le k^{1-\varepsilon}$, for every fixed $\varepsilon>0$, and an explicit bound of $7k/d$ for $d\le\sqrt{k}/2$. We achieve $2k/d+O(\log d)$, which is $(2+o(1))k/d$ whenever $d=o(k/\log k)$. Our second result is an improved construction of $k$-hybrid spanners, which guarantee stretch $2k-1$ for adjacent pairs and $k$ for non-adjacent pairs. Parter's original construction uses $O(k^2 n^{1+1/k})$ edges; we achieve the same guarantees with $O(n^{1+1/k}+kn)$ edges, removing the $k^2$ factor from the $n^{1+1/k}$ term. For every fixed $k$, our edge bound is optimal up to a constant factor under Erdős' girth conjecture.

Authors: Shiri Chechik, Gur Lifshitz

For an $n$-vertex undirected, unweighted graph $G=(V,E)$ and a positive integer $k$, we present new spanner constructions with $O_k(n^{1+1/k})$ edges that achieve nearly optimal guarantees for all distances $d\le k$. Specifically, we construct a spanner $H\subseteq G$ with $O(n^{1+1/k}+(k+d\log d)n)$ edges, ensuring that any pair at original distance at most $d$ satisfies $\mathrm{dist}_H(u,v)\le 2k+O(d\log d)$. Equivalently, the multiplicative stretch for pairs at distance $d$ is $2k/d+O(\log d)$. In particular, setting $d=k/\log k$ yields an $(O(\log k),O(k))$-spanner with $O(n^{1+1/k}+kn)$ edges. For comparison, Ben-Levy and Parter (SODA'20) obtained, for every fixed $\varepsilon>0$ and sufficiently large $k$, an $(O(k^\varepsilon),O_\varepsilon(k))$-spanner with $O_{\varepsilon,k}(n^{1+1/k})$ edges. Our result improves the multiplicative stretch from $O(k^\varepsilon)$ to $O(\log k)$ while keeping the additive term linear in $k$, bringing us closer to the goal of $(O(1),O(k))$-spanners. Furthermore, Ben-Levy and Parter obtained multiplicative stretch $O_\varepsilon(k/d)$ for distances $d\le k^{1-\varepsilon}$, for every fixed $\varepsilon>0$, and an explicit bound of $7k/d$ for $d\le\sqrt{k}/2$. We achieve $2k/d+O(\log d)$, which is $(2+o(1))k/d$ whenever $d=o(k/\log k)$. Our second result is an improved construction of $k$-hybrid spanners, which guarantee stretch $2k-1$ for adjacent pairs and $k$ for non-adjacent pairs. Parter's original construction uses $O(k^2 n^{1+1/k})$ edges; we achieve the same guarantees with $O(n^{1+1/k}+kn)$ edges, removing the $k^2$ factor from the $n^{1+1/k}$ term. For every fixed $k$, our edge bound is optimal up to a constant factor under Erdős' girth conjecture.

Consensus for Compressed Static Functions

from arXiv: Data Structures and Algorithms

Authors: Dominik Rosch, Jonatan Ziegler

The Consensus technique marked a breakthrough in the construction of minimal perfect hash functions (MPHFs), reaching a linear tradeoff between construction time and space overhead relative to the optimum. Consensus provides a clever scheme to search for and encode seeds of tasks in random data structures. We apply Consensus to the related field of compressed static functions (CSFs). These data structures store a function $f: S \to Σ$ such that querying a key $x \in S$ returns $f(x)$ and querying $x \not \in S$ returns an arbitrary value. CSFs do not need to store the keys $S$ and only need space close to the zeroth-order empirical entropy of the multiset of values. Often, some values are much more common than others. In these cases, CSFs can use less space than their non-compressed counterparts. CSFs are a useful building block, for example in database design and bioinformatics. We introduce Consensus-CSF, which can reach arbitrarily close to the empirical entropy $n H_0$, with a construction time of $n \exp(\tilde{\cal{O}} (\sqrt{1 / δ}))$ for space usage of $n H_0 (1 + δ)$ when assuming some parameters of the value distribution to be constants. This tradeoff beats previously implemented approaches that can only reach some fixed threshold above the entropy lower bound. We enable Consensus in the setting of CSFs, which is less structured than MPHFs, with the introduction of task insertions. Our approach randomly distributes the keys into one-bit Consensus tasks and then strategically inserts additional tasks in places where the construction would get stuck otherwise. We provide an implemented version of our algorithm which reaches the same order of magnitude in space overhead as competitors but is not competitive in practice. Beyond these results, we present a new way to think and reason about Consensus, which may also be applied to other problems.

Authors: Dominik Rosch, Jonatan Ziegler

The Consensus technique marked a breakthrough in the construction of minimal perfect hash functions (MPHFs), reaching a linear tradeoff between construction time and space overhead relative to the optimum. Consensus provides a clever scheme to search for and encode seeds of tasks in random data structures. We apply Consensus to the related field of compressed static functions (CSFs). These data structures store a function $f: S \to Σ$ such that querying a key $x \in S$ returns $f(x)$ and querying $x \not \in S$ returns an arbitrary value. CSFs do not need to store the keys $S$ and only need space close to the zeroth-order empirical entropy of the multiset of values. Often, some values are much more common than others. In these cases, CSFs can use less space than their non-compressed counterparts. CSFs are a useful building block, for example in database design and bioinformatics. We introduce Consensus-CSF, which can reach arbitrarily close to the empirical entropy $n H_0$, with a construction time of $n \exp(\tilde{\cal{O}} (\sqrt{1 / δ}))$ for space usage of $n H_0 (1 + δ)$ when assuming some parameters of the value distribution to be constants. This tradeoff beats previously implemented approaches that can only reach some fixed threshold above the entropy lower bound. We enable Consensus in the setting of CSFs, which is less structured than MPHFs, with the introduction of task insertions. Our approach randomly distributes the keys into one-bit Consensus tasks and then strategically inserts additional tasks in places where the construction would get stuck otherwise. We provide an implemented version of our algorithm which reaches the same order of magnitude in space overhead as competitors but is not competitive in practice. Beyond these results, we present a new way to think and reason about Consensus, which may also be applied to other problems.

A Quantum Scaling Algorithm for Maximum-Weight Perfect Matching in General Graphs

from arXiv: Data Structures and Algorithms

Authors: Kourosh Mirsohi, Sandy Irani, Michael T. Goodrich

Quantum speed-ups have been obtained for many fundamental graph problems, including most variants of matching. A notable exception, however, is the maximum-weight perfect matching (MWPM) problem in general graphs with integer edge weights, which is arguably the most challenging variant of matching. We present a quantum algorithm for MWPM in general graphs that runs in \( \widetilde{O}(n m^{2/3}\log W) \) time, where $W$ is an upper bound on the magnitude of the edge weights. This is an improvement over the best known classical combinatorial bound of \( \widetilde{O}(m\sqrt n \log W) \) in the dense regime, where $m\ge n^{3/2}$. To the best of our knowledge, this is the first quantum algorithm to obtain an asymptotic improvement over the best classical combinatorial algorithm for the MWPM problem in general graphs. The running time of our method accounts for QRAM initialization and access overheads up to polylogarithmic factors, as well as all classical updates to the data structures. At a high level, our algorithm is based on a classical framework due to Duan, Pettie, and Su, but our algorithm requires replacing certain classical tasks with quantum methods, alternative analysis of classical procedures, and the use of alternative data structures that can be implemented effectively in the QRAM model.

Authors: Kourosh Mirsohi, Sandy Irani, Michael T. Goodrich

Quantum speed-ups have been obtained for many fundamental graph problems, including most variants of matching. A notable exception, however, is the maximum-weight perfect matching (MWPM) problem in general graphs with integer edge weights, which is arguably the most challenging variant of matching. We present a quantum algorithm for MWPM in general graphs that runs in \( \widetilde{O}(n m^{2/3}\log W) \) time, where $W$ is an upper bound on the magnitude of the edge weights. This is an improvement over the best known classical combinatorial bound of \( \widetilde{O}(m\sqrt n \log W) \) in the dense regime, where $m\ge n^{3/2}$. To the best of our knowledge, this is the first quantum algorithm to obtain an asymptotic improvement over the best classical combinatorial algorithm for the MWPM problem in general graphs. The running time of our method accounts for QRAM initialization and access overheads up to polylogarithmic factors, as well as all classical updates to the data structures. At a high level, our algorithm is based on a classical framework due to Duan, Pettie, and Su, but our algorithm requires replacing certain classical tasks with quantum methods, alternative analysis of classical procedures, and the use of alternative data structures that can be implemented effectively in the QRAM model.

Multidimensional Resource Scheduling with Small Demands

from arXiv: Data Structures and Algorithms

Authors: Yossi Azar, Rathish Das, Hao Sun

We study multidimensional resource scheduling. Each job $i$ has a $d$-dimensional resource-demand vector $v_i$ and a processing time $s_i$. The scheduler assigns a start time to each job, subject to the constraint that, at every time, the total demand of the jobs being processed does not exceed $1$ in any resource dimension. The objective is to minimize the makespan. We focus on the regime in which every individual resource demand is small. We ask whether the favorable \emph{small-vector phenomenon} known for multidimensional vector packing, which corresponds to the special case of unit processing times, extends to jobs with heterogeneous processing times. The key difficulty is that arbitrary processing times create temporal interactions across multiple duration scales: a long job may overlap many shorter jobs, while feasibility must be maintained throughout every job's execution interval. We prove that the small-vector phenomenon persists in this temporal setting. For any $0<ε<1/4$, after normalizing the maximum processing time to $1$, if every coordinate of every demand vector is at most $O(ε^2/\log(d/ε))$, we give a randomized offline algorithm that produces a schedule with expected makespan at most $ (1+6ε)\mathrm{OPT}+3$. Thus, sufficiently small resource demands admit asymptotically near-optimal schedules in arbitrary dimension, despite heterogeneous processing times. We also obtain constant competitive ratios in the online setting. Let $T$ denote the ratio between the maximum and minimum processing times. If every coordinate is at most $O(1/(\log d\log T))$, we give a randomized $O(1)$-competitive algorithm, with a competitive ratio independent of both $d$ and $T$. We further derandomize our approach, obtaining a deterministic $O(1)$-competitive algorithm under a comparable smallness assumption.

Authors: Yossi Azar, Rathish Das, Hao Sun

We study multidimensional resource scheduling. Each job $i$ has a $d$-dimensional resource-demand vector $v_i$ and a processing time $s_i$. The scheduler assigns a start time to each job, subject to the constraint that, at every time, the total demand of the jobs being processed does not exceed $1$ in any resource dimension. The objective is to minimize the makespan. We focus on the regime in which every individual resource demand is small. We ask whether the favorable \emph{small-vector phenomenon} known for multidimensional vector packing, which corresponds to the special case of unit processing times, extends to jobs with heterogeneous processing times. The key difficulty is that arbitrary processing times create temporal interactions across multiple duration scales: a long job may overlap many shorter jobs, while feasibility must be maintained throughout every job's execution interval. We prove that the small-vector phenomenon persists in this temporal setting. For any $0<ε<1/4$, after normalizing the maximum processing time to $1$, if every coordinate of every demand vector is at most $O(ε^2/\log(d/ε))$, we give a randomized offline algorithm that produces a schedule with expected makespan at most $ (1+6ε)\mathrm{OPT}+3$. Thus, sufficiently small resource demands admit asymptotically near-optimal schedules in arbitrary dimension, despite heterogeneous processing times. We also obtain constant competitive ratios in the online setting. Let $T$ denote the ratio between the maximum and minimum processing times. If every coordinate is at most $O(1/(\log d\log T))$, we give a randomized $O(1)$-competitive algorithm, with a competitive ratio independent of both $d$ and $T$. We further derandomize our approach, obtaining a deterministic $O(1)$-competitive algorithm under a comparable smallness assumption.

MultiTable: A Faster Hash Table at any Physical Load Factor up to and Including One

from arXiv: Data Structures and Algorithms

Authors: Maksym Petkus

We present \emph{multitable} and its Rust reference implementation: a stable hash table both materially faster at equal physical memory and more flexible than the SwissTable in its Rust's hashbrown implementation. As an arithmetic mean over 84 configurations it delivers $\mathbf{2.1\times}$ hashbrown's throughput when both hash the same raw bytes and $\mathbf{1.9\times}$ when hashbrown is keyed on native integers, its best case; on negative lookups alone, $3.2\times$ and $2.9\times$. Multitable reaches \textbf{any physical load factor} up to and \textbf{including one} ($0.9999$ demonstrated), exactly for the requested capacity, compared to hashbrown which doubles at $0.777$ for 4-byte keys and values. At $75\%$ saturation of hashbrown (assumed average case of its rigid ladder) and multitable sized to $0.97$ physical load factor, hashbrown takes $66\%$ more space. The lookup probe count has no cliff as the load factor approaches one. Bucket size, physical load factor, and failure budget are parameters, and the multitable can be grown without rehashing. We implement two variants of multitable: plain and filtered. At equal physical memory on an Apple M2 Pro the filtered multitable leads hashbrown in all $84$ insert, hit, and miss configurations. Multitable is more \textbf{memory-efficient}, at equal mixed-lookup throughput on the map of $4$-byte keys and values the filtered multitable needs up to $12\%$ fewer bytes than hashbrown, and the plain multitable is $18\%$ smaller, holding $\mathbf{22\%}$ more keys in the same memory.

Authors: Maksym Petkus

We present \emph{multitable} and its Rust reference implementation: a stable hash table both materially faster at equal physical memory and more flexible than the SwissTable in its Rust's hashbrown implementation. As an arithmetic mean over 84 configurations it delivers $\mathbf{2.1\times}$ hashbrown's throughput when both hash the same raw bytes and $\mathbf{1.9\times}$ when hashbrown is keyed on native integers, its best case; on negative lookups alone, $3.2\times$ and $2.9\times$. Multitable reaches \textbf{any physical load factor} up to and \textbf{including one} ($0.9999$ demonstrated), exactly for the requested capacity, compared to hashbrown which doubles at $0.777$ for 4-byte keys and values. At $75\%$ saturation of hashbrown (assumed average case of its rigid ladder) and multitable sized to $0.97$ physical load factor, hashbrown takes $66\%$ more space. The lookup probe count has no cliff as the load factor approaches one. Bucket size, physical load factor, and failure budget are parameters, and the multitable can be grown without rehashing. We implement two variants of multitable: plain and filtered. At equal physical memory on an Apple M2 Pro the filtered multitable leads hashbrown in all $84$ insert, hit, and miss configurations. Multitable is more \textbf{memory-efficient}, at equal mixed-lookup throughput on the map of $4$-byte keys and values the filtered multitable needs up to $12\%$ fewer bytes than hashbrown, and the plain multitable is $18\%$ smaller, holding $\mathbf{22\%}$ more keys in the same memory.

Breaking the $2^n$ barrier for directed hamiltonicity

from arXiv: Data Structures and Algorithms

Authors: Tomohiro Koana, Soh Kumabe

We give a randomized algorithm for Directed Hamiltonian Cycle on $n$-vertex directed graphs that runs in time $O^*((375/196)^n)=O^*(1.9133^n)$. For general directed graphs, this is the first improvement in the exponential base over the classical $O^*(2^n)$-time algorithms of Bellman and Held--Karp (1962). To obtain this improvement, we first give a $(2-2^{-d})^n \, \text{poly}(n,W)$-time algorithm for counting Hamiltonian paths modulo two at each total weight when at most $d$ distinct weights from $\{1,\ldots,W\}$ enter each vertex. The algorithm combines the Laplacian determinant method of Björklund, Kaski, and Koutis (ICALP 2017) with a random linearization also used by Arvind and Guruswami (IPEC 2021). To apply the isolation lemma while keeping $d$ small, we randomly delete and duplicate arcs, partitioning the incoming copies at each vertex into $d$ groups, where $d\ge2$. We show that, if the input graph has a Hamiltonian path from $s$ to $t$, then with probability at least $\left(1-\frac{1}{1+(2^d-1)^2}\right)^{n-1}$ one can select one group at each vertex other than $s$ so that the selected arcs contain an odd number of such paths.

Authors: Tomohiro Koana, Soh Kumabe

We give a randomized algorithm for Directed Hamiltonian Cycle on $n$-vertex directed graphs that runs in time $O^*((375/196)^n)=O^*(1.9133^n)$. For general directed graphs, this is the first improvement in the exponential base over the classical $O^*(2^n)$-time algorithms of Bellman and Held--Karp (1962). To obtain this improvement, we first give a $(2-2^{-d})^n \, \text{poly}(n,W)$-time algorithm for counting Hamiltonian paths modulo two at each total weight when at most $d$ distinct weights from $\{1,\ldots,W\}$ enter each vertex. The algorithm combines the Laplacian determinant method of Björklund, Kaski, and Koutis (ICALP 2017) with a random linearization also used by Arvind and Guruswami (IPEC 2021). To apply the isolation lemma while keeping $d$ small, we randomly delete and duplicate arcs, partitioning the incoming copies at each vertex into $d$ groups, where $d\ge2$. We show that, if the input graph has a Hamiltonian path from $s$ to $t$, then with probability at least $\left(1-\frac{1}{1+(2^d-1)^2}\right)^{n-1}$ one can select one group at each vertex other than $s$ so that the selected arcs contain an odd number of such paths.

A Robustified Greedy Algorithm for Online Transportation with Improved Competitive Guarantees

from arXiv: Data Structures and Algorithms

Authors: Ritesh Seth, Syamantak Das, Sharath Raghvendra

We study the \emph{online transportation problem}, in which $n$ requests arriving sequentially in a metric space must be irrevocably assigned to $k$ capacitated facilities. Beyond classical logistics applications, this problem models resource-allocation tasks arising in machine learning, including online facility assignments, recommender systems, and mixture-of-experts routing. We introduce \emph{Robustified Greedy} (RG), a deterministic generalization of the Robust Matching algorithm that achieves a competitive ratio of $6.6604k-2.89$, improving upon the state-of-the-art bounds of $8k-7$ (Arndt et al., SOSA 2026) and $8k-5$ (Harada and Itoh, ICALP 2025). RG also retains the metric-sensitive guarantee established for Robust Matching (RM) (Nayyar and Raghvendra, FOCS 2017), achieving a competitive ratio of $O(k^{1-1/d}\log^2 n)$ in $d$-dimensional Euclidean spaces for fixed $d>1$. No comparable metric-sensitive guarantee is known for the transportation algorithms of Arndt et al.\ or Harada and Itoh. Beyond these competitive guarantees, RG provides a simple explanation for its decisions. It favors the natural nearest-neighbor assignment and, for suitable parameters, departs from this choice only when it identifies a reassignment that reduces the cost of its maintained auxiliary matching, thereby correcting accumulated assignment costs. We also prove that nearest-neighbor assignments account for a guaranteed fraction of RG's total cost, approaching one-half for appropriate parameters, even under adversarial arrivals. Experiments on real-world datasets corroborate the theory: RG achieves lower cost-to-\textsc{Opt} ratios than the competing algorithms while retaining a substantial nearest-neighbor component in its cost.

Authors: Ritesh Seth, Syamantak Das, Sharath Raghvendra

We study the \emph{online transportation problem}, in which $n$ requests arriving sequentially in a metric space must be irrevocably assigned to $k$ capacitated facilities. Beyond classical logistics applications, this problem models resource-allocation tasks arising in machine learning, including online facility assignments, recommender systems, and mixture-of-experts routing. We introduce \emph{Robustified Greedy} (RG), a deterministic generalization of the Robust Matching algorithm that achieves a competitive ratio of $6.6604k-2.89$, improving upon the state-of-the-art bounds of $8k-7$ (Arndt et al., SOSA 2026) and $8k-5$ (Harada and Itoh, ICALP 2025). RG also retains the metric-sensitive guarantee established for Robust Matching (RM) (Nayyar and Raghvendra, FOCS 2017), achieving a competitive ratio of $O(k^{1-1/d}\log^2 n)$ in $d$-dimensional Euclidean spaces for fixed $d>1$. No comparable metric-sensitive guarantee is known for the transportation algorithms of Arndt et al.\ or Harada and Itoh. Beyond these competitive guarantees, RG provides a simple explanation for its decisions. It favors the natural nearest-neighbor assignment and, for suitable parameters, departs from this choice only when it identifies a reassignment that reduces the cost of its maintained auxiliary matching, thereby correcting accumulated assignment costs. We also prove that nearest-neighbor assignments account for a guaranteed fraction of RG's total cost, approaching one-half for appropriate parameters, even under adversarial arrivals. Experiments on real-world datasets corroborate the theory: RG achieves lower cost-to-\textsc{Opt} ratios than the competing algorithms while retaining a substantial nearest-neighbor component in its cost.

Provable Classical and Quantum Local Algorithms for Max-$k$-Cut and Quantum Advantage at Moderate Girth

from arXiv: Data Structures and Algorithms

Authors: Anuj Apte, Abid Khan, Edward Farhi, Kunal Marwaha, Sami Boulebnane, Ruslan Shaydulin

Broadening the study of quantum optimization algorithms from binary to $k$-element alphabets has been shown to open new avenues for potential quantum advantage. A quantum advantage claim for approximate optimization requires showing that, under the same assumptions, an efficient quantum algorithm provably achieves a better performance than can be proven for the best known efficient classical algorithms. We study local classical and quantum algorithms for Max-$k$-Cut on $d$-regular graphs of girth $g$. We advance classical algorithms for this problem by developing a local vector algorithm based on the explicit vector construction of Thompson, Parekh, and Marwaha (TPM). Our algorithm gives the best provable cut fraction guarantee among known efficient classical algorithms on regular graphs for $k\geq3$. To evaluate the performance of QAOA under identical assumptions of girth and regularity, we develop tensor network techniques for general $d$ and $k \geq 2$. In addition, using an equivalence to a coupled qudit--boson system, we compute the QAOA performance in the infinite-degree limit. Together, these techniques give provable guarantees on QAOA performance on large graphs. Despite the improvements we introduce to the classical algorithm, QAOA achieves a better cut fraction guarantee for depths $p\geq 9$, corresponding to girth $g\geq 20$, for both finite- and infinite-degree regimes. Thus, we obtain an apparent quantum advantage from applying QAOA to the Max-$k$-Cut problem.

Authors: Anuj Apte, Abid Khan, Edward Farhi, Kunal Marwaha, Sami Boulebnane, Ruslan Shaydulin

Broadening the study of quantum optimization algorithms from binary to $k$-element alphabets has been shown to open new avenues for potential quantum advantage. A quantum advantage claim for approximate optimization requires showing that, under the same assumptions, an efficient quantum algorithm provably achieves a better performance than can be proven for the best known efficient classical algorithms. We study local classical and quantum algorithms for Max-$k$-Cut on $d$-regular graphs of girth $g$. We advance classical algorithms for this problem by developing a local vector algorithm based on the explicit vector construction of Thompson, Parekh, and Marwaha (TPM). Our algorithm gives the best provable cut fraction guarantee among known efficient classical algorithms on regular graphs for $k\geq3$. To evaluate the performance of QAOA under identical assumptions of girth and regularity, we develop tensor network techniques for general $d$ and $k \geq 2$. In addition, using an equivalence to a coupled qudit--boson system, we compute the QAOA performance in the infinite-degree limit. Together, these techniques give provable guarantees on QAOA performance on large graphs. Despite the improvements we introduce to the classical algorithm, QAOA achieves a better cut fraction guarantee for depths $p\geq 9$, corresponding to girth $g\geq 20$, for both finite- and infinite-degree regimes. Thus, we obtain an apparent quantum advantage from applying QAOA to the Max-$k$-Cut problem.

Optimal VC Dimension of Contrastive Learning with Margin

from arXiv: Data Structures and Algorithms

Authors: Dionysis Arvanitakis, Vaggos Chatziafratis, Yiyuan Luo, Konstantin Makarychev

Contrastive learning is a successful paradigm for learning $d$-dimensional geometric representations from a collection of ``anchor--positive--negative'' triplets $(i,j^{+},k^{-})$, indicating that ``item $i$ is closer to $j$ than to $k$.'' Despite its success, understanding why contrastive learning leads to representations of high \textit{generalization} quality---beyond the often pessimistic predictions from PAC-learning---remains a central question. Recently, \citet*{alon2024optimal} proved that, for PAC-learning $d$-dimensional Euclidean representations of $n$-point datasets, $Θ(\min(nd, n^2))$ triplets are necessary and sufficient, while they posed as an open question whether their VC dimension bounds for the more realistic setting of \textit{contrastive learning with a margin} can be improved. For a margin parameter $α>0$, a triplet $(i,j^{+},k^{-})_α$ is satisfied by the embedding $φ:[n]\rightarrow \mathbb{R}^{d}$, if $\|φ(i)-φ(k)\|_2>(1+α)\cdot\|φ(i)-φ(j)\|_2$. In this work, we resolve their question by proving that the VC dimension of contrastive learning under any margin $α\in(0,1)$ is in fact $O(n/α^2)$, improving on the previous bound of $O(n\log(n)/α^2)$. We also establish that the bounds are optimal up to constant factors, by providing a matching lower bound of $Ω(\frac{n}{α^2})$ (the previously known lower bound was $Ω(\frac{n}α)$), for $α\geq \max(n^{-1/2},d^{-1/2})$.

Authors: Dionysis Arvanitakis, Vaggos Chatziafratis, Yiyuan Luo, Konstantin Makarychev

Contrastive learning is a successful paradigm for learning $d$-dimensional geometric representations from a collection of ``anchor--positive--negative'' triplets $(i,j^{+},k^{-})$, indicating that ``item $i$ is closer to $j$ than to $k$.'' Despite its success, understanding why contrastive learning leads to representations of high \textit{generalization} quality---beyond the often pessimistic predictions from PAC-learning---remains a central question. Recently, \citet*{alon2024optimal} proved that, for PAC-learning $d$-dimensional Euclidean representations of $n$-point datasets, $Θ(\min(nd, n^2))$ triplets are necessary and sufficient, while they posed as an open question whether their VC dimension bounds for the more realistic setting of \textit{contrastive learning with a margin} can be improved. For a margin parameter $α>0$, a triplet $(i,j^{+},k^{-})_α$ is satisfied by the embedding $φ:[n]\rightarrow \mathbb{R}^{d}$, if $\|φ(i)-φ(k)\|_2>(1+α)\cdot\|φ(i)-φ(j)\|_2$. In this work, we resolve their question by proving that the VC dimension of contrastive learning under any margin $α\in(0,1)$ is in fact $O(n/α^2)$, improving on the previous bound of $O(n\log(n)/α^2)$. We also establish that the bounds are optimal up to constant factors, by providing a matching lower bound of $Ω(\frac{n}{α^2})$ (the previously known lower bound was $Ω(\frac{n}α)$), for $α\geq \max(n^{-1/2},d^{-1/2})$.

Local decisions, diffusive influence, and lower bounds for graphical balanced allocation

from arXiv: Data Structures and Algorithms

Authors: Obinna Okechukwu

In graphical two-choice allocation, each arriving ball is assigned to one endpoint of a random edge. We study rules whose decision is a monotone function of the two endpoint loads, allowing edge-dependent thresholds and fresh randomization. Such a rule has an exact unit-discrepancy coupling: adding one ball to the initial state produces one tagged discrepancy at every later time. We represent the tag by conditional-expectation projections on the marked edge space and obtain diffusive displacement bounds. A transport-volume inequality then converts slow propagation of influence into lower bounds for the load gap. On the cycle with $n$ vertices, from every initial distribution and at every physical time $t\ge 1/n$, the expected gap is at least a constant times $\min\{\sqrt n,t^{1/4}\}$, and the gap exceeds this scale with probability at least $1/8$. After exactly $k\ge1$ allocations, the corresponding scale is $\min\{\sqrt n,(k/n)^{1/4}\}$. No stationarity, symmetry, recurrence, or moment assumption is used. The general inequality also yields a lower bound of order $\sqrt{L/K+\log K}$ on the $L\times K$ discrete torus $C_L\square C_K$. These results separate endpoint-local rules from global-information strategies that achieve polylogarithmic gaps on cycles.

Authors: Obinna Okechukwu

In graphical two-choice allocation, each arriving ball is assigned to one endpoint of a random edge. We study rules whose decision is a monotone function of the two endpoint loads, allowing edge-dependent thresholds and fresh randomization. Such a rule has an exact unit-discrepancy coupling: adding one ball to the initial state produces one tagged discrepancy at every later time. We represent the tag by conditional-expectation projections on the marked edge space and obtain diffusive displacement bounds. A transport-volume inequality then converts slow propagation of influence into lower bounds for the load gap. On the cycle with $n$ vertices, from every initial distribution and at every physical time $t\ge 1/n$, the expected gap is at least a constant times $\min\{\sqrt n,t^{1/4}\}$, and the gap exceeds this scale with probability at least $1/8$. After exactly $k\ge1$ allocations, the corresponding scale is $\min\{\sqrt n,(k/n)^{1/4}\}$. No stationarity, symmetry, recurrence, or moment assumption is used. The general inequality also yields a lower bound of order $\sqrt{L/K+\log K}$ on the $L\times K$ discrete torus $C_L\square C_K$. These results separate endpoint-local rules from global-information strategies that achieve polylogarithmic gaps on cycles.

High-accuracy simulation of Picard HMC, part I: Gaussian cloud correction

from arXiv: Data Structures and Algorithms

Authors: Fan Chen, Sinho Chewi, Jianfeng Lu, Matthew S. Zhang

We study the problem of sampling from a continuous density $π\propto \exp(-V)$ on $\mathbb R^d$, where $V\in C^2(\mathbb R^d)$ has a $β$-Lipschitz gradient and $π$ satisfies a logarithmic Sobolev inequality with constant $α^{-1}$, and write $κ= β/α$. We introduce the Gaussian cloud sampler, which achieves total variation accuracy $\varepsilon$ using $\widetilde O(κd^{1/5}\,\text{polylog}(1/\varepsilon))$ gradient queries in expectation. The algorithm uses first-order rejection sampling (FORS) to correct the law of smoothed Picard HMC trajectories. To do so, we represent the iterates of the ideal Picard iteration via Gaussian clouds, whose centers are never evaluated, and we develop a suite of likelihood correction gadgets which only use samples from this indirect cloud representation.

Authors: Fan Chen, Sinho Chewi, Jianfeng Lu, Matthew S. Zhang

We study the problem of sampling from a continuous density $π\propto \exp(-V)$ on $\mathbb R^d$, where $V\in C^2(\mathbb R^d)$ has a $β$-Lipschitz gradient and $π$ satisfies a logarithmic Sobolev inequality with constant $α^{-1}$, and write $κ= β/α$. We introduce the Gaussian cloud sampler, which achieves total variation accuracy $\varepsilon$ using $\widetilde O(κd^{1/5}\,\text{polylog}(1/\varepsilon))$ gradient queries in expectation. The algorithm uses first-order rejection sampling (FORS) to correct the law of smoothed Picard HMC trajectories. To do so, we represent the iterates of the ideal Picard iteration via Gaussian clouds, whose centers are never evaluated, and we develop a suite of likelihood correction gadgets which only use samples from this indirect cloud representation.

Faster network motif discovery by counting isomorphic subtrees

from arXiv: Data Structures and Algorithms

Authors: Tarek Tohme, Joshua A. Grochow

We develop a new algorithm for counting the number of subgraphs of a network isomorphic to a given query graph (#SubgraphIsomorphism), motivated by network motif search. High-degree vertices (hubs), common in real-world networks, contribute to a combinatorial explosion in the number of subgraphs, making existing motif search algorithms intractable for motif sizes greater than $\approx 8$ on a wide variety of networks of interest. Our procedure leverages the $k$-core decomposition and a novel subtree-counting technique to quickly scan the periphery of a network. These two innovations allow our algorithm to significantly speed up its predecessors in practice, especially as most real-world networks have a relatively large periphery. We prove that #RootedSubtreeIsomorphism, a key subroutine in our algorithm, is #P-complete via a reduction from counting bipartite matchings. We provide analytic upper bounds on our algorithm's execution time, and evaluate its performance on 11 real-world networks of varying topologies.

Authors: Tarek Tohme, Joshua A. Grochow

We develop a new algorithm for counting the number of subgraphs of a network isomorphic to a given query graph (#SubgraphIsomorphism), motivated by network motif search. High-degree vertices (hubs), common in real-world networks, contribute to a combinatorial explosion in the number of subgraphs, making existing motif search algorithms intractable for motif sizes greater than $\approx 8$ on a wide variety of networks of interest. Our procedure leverages the $k$-core decomposition and a novel subtree-counting technique to quickly scan the periphery of a network. These two innovations allow our algorithm to significantly speed up its predecessors in practice, especially as most real-world networks have a relatively large periphery. We prove that #RootedSubtreeIsomorphism, a key subroutine in our algorithm, is #P-complete via a reduction from counting bipartite matchings. We provide analytic upper bounds on our algorithm's execution time, and evaluate its performance on 11 real-world networks of varying topologies.

Auction-Based Algorithms for Matroid Intersection: Near-Linear Query Complexity and Constant-Pass Semi-Streaming

from arXiv: Data Structures and Algorithms

Authors: Chien-Chung Huang, Yusuke Kobayashi

In this paper, we develop a new auction-based framework for matroid intersection and use it to obtain improved approximation algorithms in several computational settings. Our framework is inspired by Fleiner's generalized stable matching algorithm and extends the semi-streaming auction algorithm for bipartite matching due to Assadi, Liu, and Tarjan. Using this framework, for any $\varepsilon > 0$, we present a simple $(1-\varepsilon)$-approximation algorithm in the rank-oracle model whose query complexity matches that of the current fastest algorithm. Furthermore, by extending this result, we obtain the first $(1-\varepsilon)$-approximation algorithm for the weighted problem that requires only a near-linear number of rank-oracle queries, achieving the best known rank-oracle query complexity for the problem. We also obtain a $(1-\varepsilon)$-approximation semi-streaming algorithm for matroid intersection in the multi-pass streaming model, where the elements of the ground set arrive sequentially. It is the first algorithm achieving this approximation ratio using a constant number of passes and nearly linear space in the ranks of the matroids. When viewed in the standard offline setting, the same algorithm yields the first deterministic $(1-\varepsilon)$-approximation algorithm for matroid intersection that requires only a near-linear number of independence-oracle queries.

Authors: Chien-Chung Huang, Yusuke Kobayashi

In this paper, we develop a new auction-based framework for matroid intersection and use it to obtain improved approximation algorithms in several computational settings. Our framework is inspired by Fleiner's generalized stable matching algorithm and extends the semi-streaming auction algorithm for bipartite matching due to Assadi, Liu, and Tarjan. Using this framework, for any $\varepsilon > 0$, we present a simple $(1-\varepsilon)$-approximation algorithm in the rank-oracle model whose query complexity matches that of the current fastest algorithm. Furthermore, by extending this result, we obtain the first $(1-\varepsilon)$-approximation algorithm for the weighted problem that requires only a near-linear number of rank-oracle queries, achieving the best known rank-oracle query complexity for the problem. We also obtain a $(1-\varepsilon)$-approximation semi-streaming algorithm for matroid intersection in the multi-pass streaming model, where the elements of the ground set arrive sequentially. It is the first algorithm achieving this approximation ratio using a constant number of passes and nearly linear space in the ranks of the matroids. When viewed in the standard offline setting, the same algorithm yields the first deterministic $(1-\varepsilon)$-approximation algorithm for matroid intersection that requires only a near-linear number of independence-oracle queries.

Certifiable Near-Optimality: A Simple Framework for Unifying Search and Refutation for (Semi)random CSPs

from arXiv: Data Structures and Algorithms

Authors: Prashanti Anderson, Peter Manohar, Jeff Xu

A classical problem in average-case complexity is the study of random constraint satisfaction problems (CSPs). Random CSPs are traditionally studied in two different settings: refutation, where the instances are uniformly random and thus unsatisfiable with high probability, and search, where the instances are drawn from a planted model so that they are satisfiable. While there is no formal relationship between the refutation and search variants of random CSPs, known algorithms are strikingly similar with near-identical computational thresholds. In this work, we establish a formal relationship between the known algorithms for refutation and search by showing that in either case they achieve a stronger guarantee: they output an assignment $x$ along with a certificate $π$ that the fraction of constraints satisfied by $x$ is within some small $\varepsilon$ of the optimal assignment. We call this guarantee certifiable $\varepsilon$-optimality. As an application, we design new algorithms for a model of semirandom CSPs where the instance hypergraph (or scopes) is random, but the literal negation patterns are adversarially chosen and may depend on the hypergraph. For such CSPs, we give a family of algorithms that output certifiably $\varepsilon$-optimal solutions. We additionally study such semirandom CSPs in the "strong contamination model", where an adversary is allowed to corrupt an $O(δ)$-fraction of constraints after seeing the initial CSP. For such CSPs, we give an algorithm to output a certifiably $O(δ)$-optimal solution.

Authors: Prashanti Anderson, Peter Manohar, Jeff Xu

A classical problem in average-case complexity is the study of random constraint satisfaction problems (CSPs). Random CSPs are traditionally studied in two different settings: refutation, where the instances are uniformly random and thus unsatisfiable with high probability, and search, where the instances are drawn from a planted model so that they are satisfiable. While there is no formal relationship between the refutation and search variants of random CSPs, known algorithms are strikingly similar with near-identical computational thresholds. In this work, we establish a formal relationship between the known algorithms for refutation and search by showing that in either case they achieve a stronger guarantee: they output an assignment $x$ along with a certificate $π$ that the fraction of constraints satisfied by $x$ is within some small $\varepsilon$ of the optimal assignment. We call this guarantee certifiable $\varepsilon$-optimality. As an application, we design new algorithms for a model of semirandom CSPs where the instance hypergraph (or scopes) is random, but the literal negation patterns are adversarially chosen and may depend on the hypergraph. For such CSPs, we give a family of algorithms that output certifiably $\varepsilon$-optimal solutions. We additionally study such semirandom CSPs in the "strong contamination model", where an adversary is allowed to corrupt an $O(δ)$-fraction of constraints after seeing the initial CSP. For such CSPs, we give an algorithm to output a certifiably $O(δ)$-optimal solution.

Faster Stable Numerical Polynomial Multiplication

from arXiv: Data Structures and Algorithms

Authors: Hong Duc Bui

In the preprint [vdH08], van der Hoeven considers the problem of multiplying two polynomials with floating-point coefficients, and give an algorithm to compute the product with small relative Newton error in time $O(np \log(np))$, where $n$ is the degree and $p$ is the required precision. In this paper, we describe a significantly simpler algorithm with the same time complexity and error bound. Independently, Bringmann and Cassis considered the near-convex min-plus convolution problem in [BC23a], and presented an algorithm to solve that problem. We observe that our algorithm can be adapted to that problem to speed up Bringmann and Cassis' algorithm by a logarithmic factor.

Authors: Hong Duc Bui

In the preprint [vdH08], van der Hoeven considers the problem of multiplying two polynomials with floating-point coefficients, and give an algorithm to compute the product with small relative Newton error in time $O(np \log(np))$, where $n$ is the degree and $p$ is the required precision. In this paper, we describe a significantly simpler algorithm with the same time complexity and error bound. Independently, Bringmann and Cassis considered the near-convex min-plus convolution problem in [BC23a], and presented an algorithm to solve that problem. We observe that our algorithm can be adapted to that problem to speed up Bringmann and Cassis' algorithm by a logarithmic factor.

Quantum Query Complexity for List Search

from arXiv: Data Structures and Algorithms

Authors: Niranka Banerjee, Akinori Kawachi

Searching in a linked list is one of the most basic problems in classical algorithms. Although the nodes of the list come with memory addresses, classically those addresses play no role in the cost of search: one simply starts at the head and follows successor pointers to search for an element. In this paper, we show that the quantum setting is different. Here, the ambient address space from which the list vertices are drawn can itself affect the query complexity. We study the following problem analogous to search in a linked list in the query complexity model: the input consists of an address universe $[N]$, a public start symbol $s$, a successor oracle $f$ whose non-$\perp$ values trace a hidden simple path $s \to a_1 \to a_2 \to \cdots \to a_\ell \to \perp$, and a marking oracle $g$ that marks at most one list vertex. The task is to decide whether the list contains a marked vertex. We prove that both the decision and search versions of this problem for all $N \ge \ell \ge 1$ have quantum query complexity $Θ\!\bigl(\min\{\ell,(N\ell)^{1/4}\}\bigr)$. Thus, quite surprisingly, when $N < \ell^3$ the optimal quantum complexity is $(N\ell)^{1/4}$, which is strictly smaller than the $Θ(\ell)$ cost of ordinary linked-list traversal. This gives a precise characterization of when the ambient address space yields a genuine quantum advantage for linked-list search. We extend our results and give the same tight asymptotic bounds for the natural double linked-list version as well.

Authors: Niranka Banerjee, Akinori Kawachi

Searching in a linked list is one of the most basic problems in classical algorithms. Although the nodes of the list come with memory addresses, classically those addresses play no role in the cost of search: one simply starts at the head and follows successor pointers to search for an element. In this paper, we show that the quantum setting is different. Here, the ambient address space from which the list vertices are drawn can itself affect the query complexity. We study the following problem analogous to search in a linked list in the query complexity model: the input consists of an address universe $[N]$, a public start symbol $s$, a successor oracle $f$ whose non-$\perp$ values trace a hidden simple path $s \to a_1 \to a_2 \to \cdots \to a_\ell \to \perp$, and a marking oracle $g$ that marks at most one list vertex. The task is to decide whether the list contains a marked vertex. We prove that both the decision and search versions of this problem for all $N \ge \ell \ge 1$ have quantum query complexity $Θ\!\bigl(\min\{\ell,(N\ell)^{1/4}\}\bigr)$. Thus, quite surprisingly, when $N < \ell^3$ the optimal quantum complexity is $(N\ell)^{1/4}$, which is strictly smaller than the $Θ(\ell)$ cost of ordinary linked-list traversal. This gives a precise characterization of when the ambient address space yields a genuine quantum advantage for linked-list search. We extend our results and give the same tight asymptotic bounds for the natural double linked-list version as well.

Wednesday, September 30

Tenure-Track Faculty at Northeastern University (apply by November 15, 2026)

from CCI: jobs

The Khoury College of Computer Sciences at Northeastern University has multiple open faculty positions at all ranks (Assistant Professor, Associate Professor, Full Professor) for candidates across all areas of computer science, whose research connects to the following sub-areas of AI: Physical AI, Human-centered AI, AI and software engineering, AI safety, security and alignment. Website: www.khoury.northeastern.edu/careers/tenure-track/ […]

The Khoury College of Computer Sciences at Northeastern University has multiple open faculty positions at all ranks (Assistant Professor, Associate Professor, Full Professor) for candidates across all areas of computer science, whose research connects to the following sub-areas of AI: Physical AI, Human-centered AI, AI and software engineering, AI safety, security and alignment.

Website: https://www.khoury.northeastern.edu/careers/tenure-track/
Email: khoury-hiring@northeastern.edu

By shacharlovett

Linkage

from David Eppstein

The 3D Heesch Hemiobelisk: a nonzero-Dehn-invariant hexahedron with a complete corona (\(\mathbb{M}\)).

By David Eppstein

Does Programming Help You Understand Complexity?

from Computational Complexity

I got the following question in an email.

My nephew is currently in high school in China and has developed a strong interest in computer science and mathematics. Recently, we've been talking about how computers can solve incredibly complex problems, yet there are still some problems that even the most powerful computers struggle to handle efficiently. I work in backend data maintenance myself, so I often think about how much a system's performance depends on the way a problem is approached, rather than simply how powerful the technology is.

From your experience, what do you think is the most important thing a young student should understand about the limits of what computers can do? Would you encourage someone like my nephew to begin exploring these questions through mathematics and logical reasoning, or to start by writing programs and discovering the challenges through practice?

A great question with no perfect answer. For me, I did considerable programming in high school and college in the early days of personal computers. Computers were slow so you really had to optimize. I learned new algorithmic techniques mostly from talking to other programmers and there were some problems that the computers just didn't have enough time to solve. Those experiences definitely helped me have a good understanding of the power of good algorithms and the limitations of computing when I later went into theoretical computer science.

But many of my colleagues successfully went into the field from the mathematical side without much programming experience and did fine. 

My experience comes from a different time. Now computers are much faster and AI can just give you the best known algorithms. We have much better algorithms for NP-complete problems like Satisfiability. Unless you go looking for hard problems, you'll rarely hit one you can't solve.

Reminds me of when I took my daughter driving during a snow storm to teach her how to handle a skid. In an empty parking lot I told her to move fast and hit hard on the brakes. The car just stopped. She failed to learn because of anti-lock brakes.

Ultimately, to learn why some problems are hard, you need to understand some theoretical computer science, particularly why the halting problem is impossible to solve and why NP-complete problems are likely hard. How you get there depends on the person.

If your nephew likes to program, you can give him the challenge of trying to solve some SAT competition problems or solving some of the hard LeetCode problems. If he is more into mathematics, ask him to try to come up with a SAT algorithm mathematically so he understands how hard the problem is. The difference today is that you have to search out hard problems as it's just less likely to run into them naturally.

By Lance Fortnow

I got the following question in an email.

My nephew is currently in high school in China and has developed a strong interest in computer science and mathematics. Recently, we've been talking about how computers can solve incredibly complex problems, yet there are still some problems that even the most powerful computers struggle to handle efficiently. I work in backend data maintenance myself, so I often think about how much a system's performance depends on the way a problem is approached, rather than simply how powerful the technology is.

From your experience, what do you think is the most important thing a young student should understand about the limits of what computers can do? Would you encourage someone like my nephew to begin exploring these questions through mathematics and logical reasoning, or to start by writing programs and discovering the challenges through practice?

A great question with no perfect answer. For me, I did considerable programming in high school and college in the early days of personal computers. Computers were slow so you really had to optimize. I learned new algorithmic techniques mostly from talking to other programmers and there were some problems that the computers just didn't have enough time to solve. Those experiences definitely helped me have a good understanding of the power of good algorithms and the limitations of computing when I later went into theoretical computer science.

But many of my colleagues successfully went into the field from the mathematical side without much programming experience and did fine. 

My experience comes from a different time. Now computers are much faster and AI can just give you the best known algorithms. We have much better algorithms for NP-complete problems like Satisfiability. Unless you go looking for hard problems, you'll rarely hit one you can't solve.

Reminds me of when I took my daughter driving during a snow storm to teach her how to handle a skid. In an empty parking lot I told her to move fast and hit hard on the brakes. The car just stopped. She failed to learn because of anti-lock brakes.

Ultimately, to learn why some problems are hard, you need to understand some theoretical computer science, particularly why the halting problem is impossible to solve and why NP-complete problems are likely hard. How you get there depends on the person.

If your nephew likes to program, you can give him the challenge of trying to solve some SAT competition problems or solving some of the hard LeetCode problems. If he is more into mathematics, ask him to try to come up with a SAT algorithm mathematically so he understands how hard the problem is. The difference today is that you have to search out hard problems as it's just less likely to run into them naturally.

By Lance Fortnow

You Should Always Go for It

from Ben Recht

Nothing shows off the absurdity of cost-benefit analysis better than football analytics.

Hi there, argmin readers! Today brings one of my polarizing football posts. But it is nicely timed with the classes on cost-benefit analysis!

Football season is well underway, and my favorite absurd sports-analytical caricature, Seth Walder, has been tweeting out his hilarious takes. Walder is an ESPN NFL analyst who built a model that he claims can predict the probability of winning a game for every play to three digits of precision.

Though he maintains his sauce must remain secret, Walder claims he can determine the optimal move for any NFL play at any point in any game. His sophisticated model accounts for all aspects of the game, including teams’ strengths, the nuances of route trees, and the fluid dynamics of the air. Apparently, by looking at vast amounts of historical data and using the most cutting-edge probabilistic calculations, he can compute the probability that a team wins given their next play call. On Sundays, he loves to post graphics from his model like this:

What does this all mean? At the end of the first quarter, the Chargers had the ball on the 13-yard line. It was fourth and three, and they chose to kick a field goal rather than attempt to get a first down in hopes of continuing the drive for a touchdown. Walder’s model said the probability of winning was 53.7% if they kicked the field goal and 55.6% if they attempted the fourth-down conversion.

Where do these precise probabilities come from? He’ll never tell. On Bluesky, a commenter appropriately named “Serious Professional, PhD,” wrote, “In my experience it’s “gotta round somewhere.” Fair enough. I’ve looked into a lot of these probability calculations before, and they always end up being hacky conventions. They are definitely not precise enough to let you accurately judge whether 55.6% or 53.7% in the first quarter accurately translates into anything about winning the game.

Walder, of course, begs to differ: “This is one of those where people say “two-possession lead” when you’re way too early to think that way. It’s about point maximization and kicking, isn’t it.”

I guess he knows best.

Now, look, I’m not going to defend the Chargers, who look to have a total mess of an offense. But it is funny that Walder gets more attention than a random anon account who yells at their screen on Sundays. His cost-benefit analysis is not objective in any way. As we discussed in class, a major part of the cost-benefit framework is transparency so all stakeholders can examine the assumptions that go into computing the ratio of costs to benefits. A secret computer program that argues for maximum aggression at all times doesn’t fit the bill. It’s just a computer program he’s thrown together that says what he wants it to say.

Moreover, I will never get over how analytics people have convinced themselves that they can narrow down the infinite complexity of “going for it” to a single number. The relative complexity of attempting a kick versus running a play is vast. Kicking a field goal is the only play in football without complexity! The snapper snaps it, the holder places the ball on the ground, the kicker kicks. Sure, the field position and weather change the conditions a bit, but compare this to the complexity of what you do when you “go for it” on fourth and three. Do you run or pass? What personnel do you bring in? What route combination do you use? What does the defense do in response? It’s a complex web of decisions that Walder can’t know. But he can tell you the right answer to three digits of precision.

And speaking of the complexity of going for it versus kicking, the Jets did the go-for-2-down-8 thing again on Sunday. This is the third time under Aaron Glenn’s short coaching tenure that the Jets have found themselves down by fourteen points in the fourth quarter. They might want to work on getting leads in the fourth quarter, but that doesn’t seem part of the team’s identity.

Aaron Glenn, who spent four years as an assistant coach under Detroit Lions madman Dan Campbell, has internalized the strategy that maximizes the analytics cost-benefit analysis: score a touchdown to get within 8 and then attempt a two-point conversion. It’s certainly more aggressive, which is why Campbell’s coaching tree loves it.

Indeed, Glenn is one of the three coaches ever to have success with this desperation move. And it almost worked for him this time! He failed the first two-point conversion, but managed to get a second one with the luck of (a) foiling Campbell’s predictable run play on fourth down on the Jets’ 45, (b) driving back 55 yards for a touchdown against the Lions’ pathetic secondary, and (c) throwing a challenge flag the NFL agreed with after the conversion had apparently failed. Of course, the Lions just went and scored another touchdown on the subsequent drive, and the Jets fumbled on their final possession. So it goes.

The Jets are now 4 and 16 under Glenn. They are 1 and 3 when going for two down 8. I suppose you could argue that they are a lot better than last year, but they’ve got a long way to go before they are a good team. But Glenn gets analytics kudos because his losing moves match the model. On the broadcast, announcer Adam Amin declared, “This is an analytically advantageous place to go for two, even if you don’t get it.” Walder of course tweeted out that Glenn had passed the down 8 test. It doesn’t matter if you win. All that matters is that you adhere to the cost-benefit analysis.

I just don’t know how you watch these games and think they are more fun when you ask your computer what the right move is on every down. There’s a market for it. Some people love turning sports into escalating wars of bureaucracy. But we should at least call those people weirdos.

Subscribe now

By Ben Recht

TR26-222 | Discrepancy for Random Linear Codes | Joao Ribeiro, Nicolas Resch, Dean Doron, Jonathan Mosheiff, Tal Leonov, Henrique Navas

from ECCC Papers

We show that random linear codes possess nearly optimal discrepancy-type properties in a broad range of settings. Our main results are two general discrepancy theorems: one controls all translates of a fixed test, and the other controls large families of Fourier-pseudorandom tests. Two motivating applications follow: First, random linear codes behave essentially like unstructured random codes for list-decoding from errors above capacity. More precisely, a random linear code $C\subseteq \mathbb{F}_q^n$ of rate $1 - \frac{1}{n}\log_q|B_\rho| + \varepsilon$, where $|B_\rho|$ is the volume of a radius-$\rho$ Hamming ball in $\mathbb{F}_q^n$, satisfies $|C \cap B| = (1\pm o(1)) \frac{|C|\cdot |B|}{q^n}$ simultaneously for all radius-$\rho$ Hamming balls $B$ in $\mathbb{F}_q^n$ with high probability. This vastly generalizes the previously best known fact that random linear codes of this rate have covering radius at most $\rho n$ with high probability (Blinovsky, 1987). Second, over prime fields, random linear codes behave essentially like unstructured random codes for zero-error list-recovery, and list-recovery from erasures, above capacity. More precisely, for a prime $q>2$ and input list size $2\leq \ell\leq q-1$, a random linear code $C\subseteq \mathbb{F}_q^n$ of rate $1-\log_q \ell+\varepsilon$ will satisfy $|C \cap S| = (1\pm o(1)) \frac{|C|\cdot \ell^n}{q^n}$ simultaneously for all combinatorial rectangles $S=S_1\times S_2\times\cdots\times S_n$, where $|S_i|=\ell$ for all $i$, with high probability. An analogous result also holds when we can bound $|S_i|$ only for some of the $i$'s. In particular, we use this to show the abundance of locally leakage-resilient $n$-party linear ramp secret sharing schemes with any linear reconstruction threshold and sublinear threshold gap $O(n/\log n)$ over fields $\mathbb{F}_q$ of polynomial size $q=\Theta(n^\gamma)$ for a constant $\gamma\in(0,1/5)$. Prior work on the existence of leakage-resilient linear secret sharing was stuck at reconstruction thresholds above $n/2$ for both threshold and ramp schemes. The translate-family result, and hence the list-decoding application, applies over arbitrary finite fields, even when the field size grows with $n$. The list-recovery and leakage applications over prime fields hold under moderate-growth conditions on $q$, for example $q\le n^{1/5-o(1)}$. Our results are obtained through a careful second-moment analysis of the evolution of intersection sizes as random generators are added to $C$ one by one.
We show that random linear codes possess nearly optimal discrepancy-type properties in a broad range of settings. Our main results are two general discrepancy theorems: one controls all translates of a fixed test, and the other controls large families of Fourier-pseudorandom tests. Two motivating applications follow: First, random linear codes behave essentially like unstructured random codes for list-decoding from errors above capacity. More precisely, a random linear code $C\subseteq \mathbb{F}_q^n$ of rate $1 - \frac{1}{n}\log_q|B_\rho| + \varepsilon$, where $|B_\rho|$ is the volume of a radius-$\rho$ Hamming ball in $\mathbb{F}_q^n$, satisfies $|C \cap B| = (1\pm o(1)) \frac{|C|\cdot |B|}{q^n}$ simultaneously for all radius-$\rho$ Hamming balls $B$ in $\mathbb{F}_q^n$ with high probability. This vastly generalizes the previously best known fact that random linear codes of this rate have covering radius at most $\rho n$ with high probability (Blinovsky, 1987). Second, over prime fields, random linear codes behave essentially like unstructured random codes for zero-error list-recovery, and list-recovery from erasures, above capacity. More precisely, for a prime $q>2$ and input list size $2\leq \ell\leq q-1$, a random linear code $C\subseteq \mathbb{F}_q^n$ of rate $1-\log_q \ell+\varepsilon$ will satisfy $|C \cap S| = (1\pm o(1)) \frac{|C|\cdot \ell^n}{q^n}$ simultaneously for all combinatorial rectangles $S=S_1\times S_2\times\cdots\times S_n$, where $|S_i|=\ell$ for all $i$, with high probability. An analogous result also holds when we can bound $|S_i|$ only for some of the $i$'s. In particular, we use this to show the abundance of locally leakage-resilient $n$-party linear ramp secret sharing schemes with any linear reconstruction threshold and sublinear threshold gap $O(n/\log n)$ over fields $\mathbb{F}_q$ of polynomial size $q=\Theta(n^\gamma)$ for a constant $\gamma\in(0,1/5)$. Prior work on the existence of leakage-resilient linear secret sharing was stuck at reconstruction thresholds above $n/2$ for both threshold and ramp schemes. The translate-family result, and hence the list-decoding application, applies over arbitrary finite fields, even when the field size grows with $n$. The list-recovery and leakage applications over prime fields hold under moderate-growth conditions on $q$, for example $q\le n^{1/5-o(1)}$. Our results are obtained through a careful second-moment analysis of the evolution of intersection sizes as random generators are added to $C$ one by one.

Information Checking Protocols

from Decentralized Thoughts

The siege is over. The colony’s archivist brings its captain a recorded message from their founder, Hari. He predicted the attack, she says, and the choices that would let them survive. Hari is long dead. Did he foresee the siege, or did the archivist fake the recording last week? This story is inspired by Hari Seldon’s recorded messages in Isaac Asimov’s Foundation. Before the expedition leaves Earth, Hari records his...

By Ittai Abraham, Gilad Asharov, Gilad Stern

The siege is over. The colony’s archivist brings its captain a recorded message from their founder, Hari. He predicted the attack, she says, and the choices that would let them survive. Hari is long dead. Did he foresee the siege, or did the archivist fake the recording last week? This story is inspired by Hari Seldon’s recorded messages in Isaac Asimov’s Foundation. Before the expedition leaves Earth, Hari records his...

By Ittai Abraham, Gilad Asharov, Gilad Stern

TR26-221 | Top-Down Lower Bounds for All Depths | Oliver Korten

from ECCC Papers

We prove that Parity requires $2^{n^{\Omega(1)}}$ size De Morgan circuits of constant depth using a new method which is completely “top-down” in the sense of [HJP95]. The proof relies crucially on the core ideas developed in a line of work [HJP95, PPZ99, MW19, GRSS24] which previously established top-down lower bounds for circuits of depth 3 and 4. We first present a proof of a lower bound $\exp(n^{3^{-d}})$. In this case, nearly all of the relevant combinatorial ideas necessary for the proof are already present in some form in [GRSS24]. We then present two extensions of this argument, the first achieving a lower bound $\exp(\epsilon_d n^{1/(2d-2)})$ for some $\epsilon_d>0$ depending only on $d$, and the second achieving the essentially tight lower bound $\exp(\epsilon_d n^{1/(d-1)})$. These improved results each hinge on establishing a key lemma which quantifies the extent to which a high entropy random variable in $\{0,1\}^n$ will look close to uniform after projecting it onto a random small set of coordinates $R \subseteq [n]$.
We prove that Parity requires $2^{n^{\Omega(1)}}$ size De Morgan circuits of constant depth using a new method which is completely “top-down” in the sense of [HJP95]. The proof relies crucially on the core ideas developed in a line of work [HJP95, PPZ99, MW19, GRSS24] which previously established top-down lower bounds for circuits of depth 3 and 4. We first present a proof of a lower bound $\exp(n^{3^{-d}})$. In this case, nearly all of the relevant combinatorial ideas necessary for the proof are already present in some form in [GRSS24]. We then present two extensions of this argument, the first achieving a lower bound $\exp(\epsilon_d n^{1/(2d-2)})$ for some $\epsilon_d>0$ depending only on $d$, and the second achieving the essentially tight lower bound $\exp(\epsilon_d n^{1/(d-1)})$. These improved results each hinge on establishing a key lemma which quantifies the extent to which a high entropy random variable in $\{0,1\}^n$ will look close to uniform after projecting it onto a random small set of coordinates $R \subseteq [n]$.

Ramanujan quantum expanders from the Weil representation

from arXiv: Computational Complexity

Authors: Siddhartha Jain

For every odd prime power $q$ and $D=q+1$, we construct an infinite family of Ramanujan quantum expanders of degree $D$. The construction transfers Morgenstern's Ramanujan Cayley graphs on $\operatorname{PSL}_2(\mathbb{F}_{p})$ for $p$ which is an even power of $q$, through the odd irreducible subrepresentation of the Weil representation of $\operatorname{SL}_2(\mathbb{F}_{p})$. For a quantum expander of dimension $N$, our implementation uses $O(\log^2 N)$ elementary gates, and $O(\log N)$ ancilla qudits, using a fixed finite gate set depending on $q$. An advantage compared to the previous work of Iyer, Jain, Jordan, and Somma (FOCS 2026) is that assuming the quantum circuit is implemented exactly, we satisfy the $2\sqrt{D-1}/D$ singular value bound exactly without any additive error.

Authors: Siddhartha Jain

For every odd prime power $q$ and $D=q+1$, we construct an infinite family of Ramanujan quantum expanders of degree $D$. The construction transfers Morgenstern's Ramanujan Cayley graphs on $\operatorname{PSL}_2(\mathbb{F}_{p})$ for $p$ which is an even power of $q$, through the odd irreducible subrepresentation of the Weil representation of $\operatorname{SL}_2(\mathbb{F}_{p})$. For a quantum expander of dimension $N$, our implementation uses $O(\log^2 N)$ elementary gates, and $O(\log N)$ ancilla qudits, using a fixed finite gate set depending on $q$. An advantage compared to the previous work of Iyer, Jain, Jordan, and Somma (FOCS 2026) is that assuming the quantum circuit is implemented exactly, we satisfy the $2\sqrt{D-1}/D$ singular value bound exactly without any additive error.

Optimal Quantum-Classical Separations for Exact Learning

from arXiv: Computational Complexity

Authors: Srinivasan Arunachalam, Amin Shiraz Gilani, Nikhil S. Mande

We study exact learning with membership queries for concept classes $\mathcal C\subseteq\{0,1\}^N$, focusing on the relationships among their deterministic, randomized, and quantum query complexities, denoted $\mathsf{D}(\mathcal C)$, $\mathsf{R}(\mathcal C)$, and $\mathsf{Q}(\mathcal C)$, respectively. The two canonical quantum speedups in this model are witnessed by Grover search and Bernstein-Vazirani, leading to the longstanding conjecture $$ \mathsf{R}(\mathcal C)=O(\mathsf{Q}(\mathcal C)^2+\mathsf{Q}(\mathcal C)\log N). $$ We first refute this conjecture by constructing concept classes $\mathcal C$ and $\mathcal C'$ satisfying \[ \mathsf{R}(\mathcal C)=Ω\!\left(\frac{\mathsf{Q}(\mathcal C)^3\log N}{\log \mathsf{Q}(\mathcal C)}\right) \qquad\text{and}\qquad \mathsf{D}(\mathcal C')=Ω(\mathsf{Q}(\mathcal C')^3\log N). \] The first bound matches the upper bound of Arunachalam et al.~[Quantum'21] up to constant factors, while the second matches the upper bound of Servedio and Gortler~[SICOMP'04]. In particular, this shows that the saving in the randomized upper bound of Arunachalam et al. fundamentally relies on randomness. Apart from characterizing the optimal relationship between classical and quantum query complexity, our results are the first to show that quantum speedups for learning can go beyond the Grover and Bernstein-Vazirani paradigms.

Authors: Srinivasan Arunachalam, Amin Shiraz Gilani, Nikhil S. Mande

We study exact learning with membership queries for concept classes $\mathcal C\subseteq\{0,1\}^N$, focusing on the relationships among their deterministic, randomized, and quantum query complexities, denoted $\mathsf{D}(\mathcal C)$, $\mathsf{R}(\mathcal C)$, and $\mathsf{Q}(\mathcal C)$, respectively. The two canonical quantum speedups in this model are witnessed by Grover search and Bernstein-Vazirani, leading to the longstanding conjecture $$ \mathsf{R}(\mathcal C)=O(\mathsf{Q}(\mathcal C)^2+\mathsf{Q}(\mathcal C)\log N). $$ We first refute this conjecture by constructing concept classes $\mathcal C$ and $\mathcal C'$ satisfying \[ \mathsf{R}(\mathcal C)=Ω\!\left(\frac{\mathsf{Q}(\mathcal C)^3\log N}{\log \mathsf{Q}(\mathcal C)}\right) \qquad\text{and}\qquad \mathsf{D}(\mathcal C')=Ω(\mathsf{Q}(\mathcal C')^3\log N). \] The first bound matches the upper bound of Arunachalam et al.~[Quantum'21] up to constant factors, while the second matches the upper bound of Servedio and Gortler~[SICOMP'04]. In particular, this shows that the saving in the randomized upper bound of Arunachalam et al. fundamentally relies on randomness. Apart from characterizing the optimal relationship between classical and quantum query complexity, our results are the first to show that quantum speedups for learning can go beyond the Grover and Bernstein-Vazirani paradigms.

Resonance Breaking in Noisy Shor's Algorithm

from arXiv: Computational Complexity

Authors: Zhengwei Liu

We argue that the exponential quantum advantage in Shor's algorithm is broken under one-layer depolarizing noise of arbitrary small error rate, by analyzing a resonance breaking phenomenon in noisy quantum circuits. First, we express the distribution of measurements on $n$-qubit strings as the superposition of the $4^n$ wave functions in the Pauli path integral. In the noiseless case, the bit-strings achieving resonant peaks guaranteed a constant rate of successful measurements to factor a large number. Secondly, when the middle layer of the quantum circuit has independent depolarizing noise of rate $λ$, for any Hamming weight $d$, we obtain corresponding measurement rate bounded by $2(1-λ)^d$ for higher frequency terms and by $O(n^d)/2^{n/2}$ for low frequency terms. The rate approaches to zero as $n$ and $d$ approaches to infinity. Resonance breaking destroys the exponential quantum advantage. Furthermore, we design a classical factoring algorithm in polynomial time $O(n^d)$ to substitute the low frequency contribution.

Authors: Zhengwei Liu

We argue that the exponential quantum advantage in Shor's algorithm is broken under one-layer depolarizing noise of arbitrary small error rate, by analyzing a resonance breaking phenomenon in noisy quantum circuits. First, we express the distribution of measurements on $n$-qubit strings as the superposition of the $4^n$ wave functions in the Pauli path integral. In the noiseless case, the bit-strings achieving resonant peaks guaranteed a constant rate of successful measurements to factor a large number. Secondly, when the middle layer of the quantum circuit has independent depolarizing noise of rate $λ$, for any Hamming weight $d$, we obtain corresponding measurement rate bounded by $2(1-λ)^d$ for higher frequency terms and by $O(n^d)/2^{n/2}$ for low frequency terms. The rate approaches to zero as $n$ and $d$ approaches to infinity. Resonance breaking destroys the exponential quantum advantage. Furthermore, we design a classical factoring algorithm in polynomial time $O(n^d)$ to substitute the low frequency contribution.

Robust Approximation and the Arity Barrier at Width Two

from arXiv: Computational Complexity

Authors: Nathan Benedetto Proença, Koppány István Encz, Monaldo Mastrolilli

Zwick's algorithm for Horn satisfiability shows that constraints of unbounded arity can admit a robust approximation guarantee independent of the arity. For finite constraint languages, Barto and Kozik proved that robust approximability is characterized by bounded width. We ask whether arity-independent robustness persists beyond width one and show that it fails already at width two. For Majority-closed Boolean linear constraints of maximum arity $k \ge 2$, we give a randomized polynomial-time algorithm that, without knowing $\varepsilon$, violates an expected $O(\sqrt{\varepsilon \log k})$ fraction of the constraint weight on $(1-\varepsilon)$-satisfiable instances. Under the Unique Games Conjecture (UGC), a matching NP-hardness lower bound of $Ω(\sqrt{\varepsilon \log k})$ holds in an explicit parameter regime. Under this assumption, the bounded-width characterization therefore does not extend uniformly to constraints of unbounded arity. The algorithm rounds a degree-eight Sum-of-Squares (SoS) relaxation with a single Gaussian threshold. Since polynomial size alone does not make SoS solvable in polynomial bit complexity, we construct complete truncated Groebner bases for the soft Majority ideal and show that, at every fixed degree $2d$, the relaxation can be optimized to any rational accuracy in polynomial time with exactly feasible solutions, and degree-$2d$ SoS proofs can be found after an additive perturbation at degree at most $4d+10$. The lower bound combines Raghavendra's gap-to-hardness theorem with an integrality gap on a Gaussian star, analyzed via the Isaksson-Mossel theorem that parallel halfspaces maximize the joint membership probability of exchangeable Gaussians.

Authors: Nathan Benedetto Proença, Koppány István Encz, Monaldo Mastrolilli

Zwick's algorithm for Horn satisfiability shows that constraints of unbounded arity can admit a robust approximation guarantee independent of the arity. For finite constraint languages, Barto and Kozik proved that robust approximability is characterized by bounded width. We ask whether arity-independent robustness persists beyond width one and show that it fails already at width two. For Majority-closed Boolean linear constraints of maximum arity $k \ge 2$, we give a randomized polynomial-time algorithm that, without knowing $\varepsilon$, violates an expected $O(\sqrt{\varepsilon \log k})$ fraction of the constraint weight on $(1-\varepsilon)$-satisfiable instances. Under the Unique Games Conjecture (UGC), a matching NP-hardness lower bound of $Ω(\sqrt{\varepsilon \log k})$ holds in an explicit parameter regime. Under this assumption, the bounded-width characterization therefore does not extend uniformly to constraints of unbounded arity. The algorithm rounds a degree-eight Sum-of-Squares (SoS) relaxation with a single Gaussian threshold. Since polynomial size alone does not make SoS solvable in polynomial bit complexity, we construct complete truncated Groebner bases for the soft Majority ideal and show that, at every fixed degree $2d$, the relaxation can be optimized to any rational accuracy in polynomial time with exactly feasible solutions, and degree-$2d$ SoS proofs can be found after an additive perturbation at degree at most $4d+10$. The lower bound combines Raghavendra's gap-to-hardness theorem with an integrality gap on a Gaussian star, analyzed via the Isaksson-Mossel theorem that parallel halfspaces maximize the joint membership probability of exchangeable Gaussians.

Rational Identity Testing for Noncommutative Circuits is in Polynomial Space

from arXiv: Computational Complexity

Authors: V. Arvind, Pushkar S. Joglekar

The rational identity testing problem, RIT, asks whether an input \emph{rational circuit}, computes the zero element of the free skew field. For rational \emph{formulas} the problem is known to be in deterministic polynomial time. For rational circuits of size $s$ the associated linear pencil has dimension $2^{O(s)}$, and the known algorithms use exponential time and exponential space. We show that RIT for rational circuits over $Q$ and finite fields is in PSPACE. As a consequence, noncommutative PIT for degree unrestricted noncommutative circuits is also in PSPACE. The proof is based on the following three observations Analyzing the Hrubes-Wigderson construction of a linear pencil for an input rational formula, we obtain a succinctly represented linear pencil $A$ of size $2^{O(s)}$ for the input rational \emph{circuit} of size $s$. More precisely, given indices $i$ and $j$ for the linear pencil we can compute $A_{i,j}$ in space polynomial in $s$ and length of $i,j$. This algorithm essentially gives a succinct presentation of the linear pencil $A$ obtained by the reduction. By the recent theorem of Chatterjee, Ghosh, Gurjar, Raj and Thierauf that deciding whether a symbolic matrix $\sum_i A_ix_i$ has full noncommutative rank is in NC. We note that their algorithm is logspace-uniform NC and we use it as a black-box for computing rank of succinctly represented pencil $A$. Finally, we note a folklore simulation: If a problem is solved by a logspace-uniform family of \emph{deterministic} Boolean circuits of polylogarithmic depth, and its input is not written down but is presented by a polylogspace subroutine that returns any requested input bit, then the circuit can be evaluated in polylogarithmic space. Hence, simulating a depth $O(\log^{i}M)$ circuit for succinctly presented inputs of length $M = 2^{Θ(s)}$ yields a polynomial space algorithm.

Authors: V. Arvind, Pushkar S. Joglekar

The rational identity testing problem, RIT, asks whether an input \emph{rational circuit}, computes the zero element of the free skew field. For rational \emph{formulas} the problem is known to be in deterministic polynomial time. For rational circuits of size $s$ the associated linear pencil has dimension $2^{O(s)}$, and the known algorithms use exponential time and exponential space. We show that RIT for rational circuits over $Q$ and finite fields is in PSPACE. As a consequence, noncommutative PIT for degree unrestricted noncommutative circuits is also in PSPACE. The proof is based on the following three observations Analyzing the Hrubes-Wigderson construction of a linear pencil for an input rational formula, we obtain a succinctly represented linear pencil $A$ of size $2^{O(s)}$ for the input rational \emph{circuit} of size $s$. More precisely, given indices $i$ and $j$ for the linear pencil we can compute $A_{i,j}$ in space polynomial in $s$ and length of $i,j$. This algorithm essentially gives a succinct presentation of the linear pencil $A$ obtained by the reduction. By the recent theorem of Chatterjee, Ghosh, Gurjar, Raj and Thierauf that deciding whether a symbolic matrix $\sum_i A_ix_i$ has full noncommutative rank is in NC. We note that their algorithm is logspace-uniform NC and we use it as a black-box for computing rank of succinctly represented pencil $A$. Finally, we note a folklore simulation: If a problem is solved by a logspace-uniform family of \emph{deterministic} Boolean circuits of polylogarithmic depth, and its input is not written down but is presented by a polylogspace subroutine that returns any requested input bit, then the circuit can be evaluated in polylogarithmic space. Hence, simulating a depth $O(\log^{i}M)$ circuit for succinctly presented inputs of length $M = 2^{Θ(s)}$ yields a polynomial space algorithm.

Efficiently Approximating Attention Is Hard

from arXiv: Computational Complexity

Authors: Lukas Haverbeck, Carmen Amo Alonso, Andres Felipe Posada-Moreno, Sebastian Trimpe, Marco Pavone

Softmax attention is ubiquitous in modern machine learning, but its quadratic scaling with sequence length makes it costly. To reduce this cost, attention is often approximated with fast algorithms, which incur error but can still perform well in practice and on some inputs. At the same time, the growing diversity of attention applications makes approximation guarantees that do not depend on particular input structure a compelling target. For such uniform guarantees over all inputs, known runtime lower bounds rule out fast algorithms for near-exact attention, but leave open the practically important regime: is there an efficient algorithm with even a modest uniform approximation guarantee? We answer this question negatively. Under standard complexity-theoretic assumptions, no truly subquadratic algorithm can approximate attention with any nontrivial additive or relative guarantee uniformly over all inputs. This impossibility holds in the mildest parameter regime for which known algorithms do not already achieve strong approximation guarantees in near-linear time, and extends to practically relevant relaxations: even after polynomial preprocessing of the KV cache, no efficient algorithm can obtain a nontrivial uniform approximation guarantee, or identify a small set of keys receiving substantial attention under sparsity. Overall, our results settle the computational limits of uniform attention approximation.

Authors: Lukas Haverbeck, Carmen Amo Alonso, Andres Felipe Posada-Moreno, Sebastian Trimpe, Marco Pavone

Softmax attention is ubiquitous in modern machine learning, but its quadratic scaling with sequence length makes it costly. To reduce this cost, attention is often approximated with fast algorithms, which incur error but can still perform well in practice and on some inputs. At the same time, the growing diversity of attention applications makes approximation guarantees that do not depend on particular input structure a compelling target. For such uniform guarantees over all inputs, known runtime lower bounds rule out fast algorithms for near-exact attention, but leave open the practically important regime: is there an efficient algorithm with even a modest uniform approximation guarantee? We answer this question negatively. Under standard complexity-theoretic assumptions, no truly subquadratic algorithm can approximate attention with any nontrivial additive or relative guarantee uniformly over all inputs. This impossibility holds in the mildest parameter regime for which known algorithms do not already achieve strong approximation guarantees in near-linear time, and extends to practically relevant relaxations: even after polynomial preprocessing of the KV cache, no efficient algorithm can obtain a nontrivial uniform approximation guarantee, or identify a small set of keys receiving substantial attention under sparsity. Overall, our results settle the computational limits of uniform attention approximation.

Sample Complexity of Equivariant Reinforcement Learning

from arXiv: Computational Complexity

Authors: Rayan Mazouz, Haibo Zhao, Chris Hillar, Christian Shewmake

Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is costly. While the world often exhibits geometric and physical symmetries, standard RL algorithms typically fail to exploit this structure. In this paper, we demonstrate that exploiting group symmetries significantly reduces the sample complexity of RL. Focusing on finite-horizon Markov decision processes, we find that leveraging homomorphisms induced by group symmetries significantly reduces the theoretical upper and lower bounds on the number of environment interactions required to reach an optimal return. We further extend these bounds to continuous state and action spaces, providing corresponding sample-complexity guarantees under appropriate regularity assumptions. Beyond theory, we validate our findings through controlled experiments and demonstrate the advantages of symmetry-aware policy learning on high-dimensional continuous robotic simulations. Our results show that integrating symmetry into the learning pipeline yields substantial gains in sample efficiency and performance, offering a principled path toward more data-efficient robotics.

Authors: Rayan Mazouz, Haibo Zhao, Chris Hillar, Christian Shewmake

Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is costly. While the world often exhibits geometric and physical symmetries, standard RL algorithms typically fail to exploit this structure. In this paper, we demonstrate that exploiting group symmetries significantly reduces the sample complexity of RL. Focusing on finite-horizon Markov decision processes, we find that leveraging homomorphisms induced by group symmetries significantly reduces the theoretical upper and lower bounds on the number of environment interactions required to reach an optimal return. We further extend these bounds to continuous state and action spaces, providing corresponding sample-complexity guarantees under appropriate regularity assumptions. Beyond theory, we validate our findings through controlled experiments and demonstrate the advantages of symmetry-aware policy learning on high-dimensional continuous robotic simulations. Our results show that integrating symmetry into the learning pipeline yields substantial gains in sample efficiency and performance, offering a principled path toward more data-efficient robotics.

Spectral Methods for the Complexity of Planar Graph Homomorphisms

from arXiv: Computational Complexity

Authors: Ashwin Maran, Jin-Yi Cai, Zhuxiao Tang

We explore the frontier beyond the recently discovered barrier represented by the \emph{quantum automorphism group} $qut(M)$ in the classification theory of planar graph homomorphisms $PlGH(M)$. We show that analyzing the spectral relations of $M$ can prove \#P-hardness when traditional vertex separation and domain-reduction methods with planar edge gadgets provably fail due to the $\qut(M)$ barrier. We prove two criteria of \#P-hardness for $PlGH(M)$: a spectral criterion and a determinant criterion. It is known that the core problem for the classification of $PlGH(M)$ for nonnegative matrices $M$ is for positive definite entry-wise positive matrices. We use the spectral criterion to show that $PlGH(M)$ is \#P-hard for all circulant matrices of prime order $q \ge 3$, while for $q=2$ it is precisely the matchgate case and is P-time computable by the FKT algorithm (for planar perfect matching). We also prove a complexity dichotomy for $\PlGH$ problems defined by tensor products of 2 by 2 matrices. This gives a complete complexity classification for this class of matrices, and the FKT algorithm together with a holographic transformation is \emph{universal}---every $PlGH(M)$ is either (1) P-time computable over all graphs, or (2) \#P-hard in general but P-time computable over planar graphs, or (3) \#P-hard over planar graphs; furthermore, $PlGH(M)$ in (2) consists of precisely those computable by FKT with a holographic transformation.

Authors: Ashwin Maran, Jin-Yi Cai, Zhuxiao Tang

We explore the frontier beyond the recently discovered barrier represented by the \emph{quantum automorphism group} $qut(M)$ in the classification theory of planar graph homomorphisms $PlGH(M)$. We show that analyzing the spectral relations of $M$ can prove \#P-hardness when traditional vertex separation and domain-reduction methods with planar edge gadgets provably fail due to the $\qut(M)$ barrier. We prove two criteria of \#P-hardness for $PlGH(M)$: a spectral criterion and a determinant criterion. It is known that the core problem for the classification of $PlGH(M)$ for nonnegative matrices $M$ is for positive definite entry-wise positive matrices. We use the spectral criterion to show that $PlGH(M)$ is \#P-hard for all circulant matrices of prime order $q \ge 3$, while for $q=2$ it is precisely the matchgate case and is P-time computable by the FKT algorithm (for planar perfect matching). We also prove a complexity dichotomy for $\PlGH$ problems defined by tensor products of 2 by 2 matrices. This gives a complete complexity classification for this class of matrices, and the FKT algorithm together with a holographic transformation is \emph{universal}---every $PlGH(M)$ is either (1) P-time computable over all graphs, or (2) \#P-hard in general but P-time computable over planar graphs, or (3) \#P-hard over planar graphs; furthermore, $PlGH(M)$ in (2) consists of precisely those computable by FKT with a holographic transformation.

From Weak to Strong Testing in Gaussian Models

from arXiv: Computational Complexity

Authors: Ansh Nagda, Alexander S. Wein

We study the computational complexity of hypothesis testing in the spiked Wigner model, a prototypical model for detecting low-rank structure in a large random matrix. Below the "BBP" eigenvalue transition, it is expected that strong detection --- with both type I and II errors vanishing --- requires exponential time. Assuming this as a conjecture, we determine the limits of polynomial-time weak detection, exactly characterizing the possible tradeoffs between type I and II errors. Specifically, the optimal tradeoff is achieved by a particular linear spectral statistic. Thus, the question of weak detection is entirely reduced to that of strong detection. The proof builds on ideas of Nagda-Raghavendra (2025) and Moitra-Wein (2025). The low-degree likelihood ratio (LDLR) plays a key role: any test that slightly beats the LDLR can be boosted to have an even higher success probability. This leads us to establish a computational analogue of the Neyman-Pearson lemma for a subclass of additive Gaussian models: for a given super-polynomial runtime, the best possible tradeoff between type I and II errors is either the one achieved by thresholding the LDLR, or the trivial tradeoff that results from strong detection.

Authors: Ansh Nagda, Alexander S. Wein

We study the computational complexity of hypothesis testing in the spiked Wigner model, a prototypical model for detecting low-rank structure in a large random matrix. Below the "BBP" eigenvalue transition, it is expected that strong detection --- with both type I and II errors vanishing --- requires exponential time. Assuming this as a conjecture, we determine the limits of polynomial-time weak detection, exactly characterizing the possible tradeoffs between type I and II errors. Specifically, the optimal tradeoff is achieved by a particular linear spectral statistic. Thus, the question of weak detection is entirely reduced to that of strong detection. The proof builds on ideas of Nagda-Raghavendra (2025) and Moitra-Wein (2025). The low-degree likelihood ratio (LDLR) plays a key role: any test that slightly beats the LDLR can be boosted to have an even higher success probability. This leads us to establish a computational analogue of the Neyman-Pearson lemma for a subclass of additive Gaussian models: for a given super-polynomial runtime, the best possible tradeoff between type I and II errors is either the one achieved by thresholding the LDLR, or the trivial tradeoff that results from strong detection.

Sparse cubical complexes for efficient topology-preservation in image data

from arXiv: Computational Geometry

Authors: Alexander H. Berger, Marco Fontana, Daniel Rueckert, Johannes C. Paetzold, Laurin Lux, Ulrich Bauer

Persistent homology (PH) is a frequently used tool for extracting and preserving topological information from image data, particularly in image segmentation, where preservation of topological structures is important. However, despite its general applicability across dimensionality, domains, and target structures, the runtime cost of PH-based methods often makes their practical use infeasible. In this work, we argue that this runtime cost is largely driven by processing information that is unimportant for downstream application (e.g. as optimization objective). We propose sparse cubical filtrations as an alternative foundation for PH computation, reducing subsequent computational costs by factors of up to 100 on real datasets. We show close agreement with the optimization signal of the dense counterpart and empirically evaluate our solution's effectiveness as an optimization objective in realistic training regimes where other PH-based objectives can practically not operate (i.e., 3D data with large patch sizes). We show how our solution improves topological accuracy by up to 80\% across six diverse datasets while maintaining pixel- and region-based accuracy.

Authors: Alexander H. Berger, Marco Fontana, Daniel Rueckert, Johannes C. Paetzold, Laurin Lux, Ulrich Bauer

Persistent homology (PH) is a frequently used tool for extracting and preserving topological information from image data, particularly in image segmentation, where preservation of topological structures is important. However, despite its general applicability across dimensionality, domains, and target structures, the runtime cost of PH-based methods often makes their practical use infeasible. In this work, we argue that this runtime cost is largely driven by processing information that is unimportant for downstream application (e.g. as optimization objective). We propose sparse cubical filtrations as an alternative foundation for PH computation, reducing subsequent computational costs by factors of up to 100 on real datasets. We show close agreement with the optimization signal of the dense counterpart and empirically evaluate our solution's effectiveness as an optimization objective in realistic training regimes where other PH-based objectives can practically not operate (i.e., 3D data with large patch sizes). We show how our solution improves topological accuracy by up to 80\% across six diverse datasets while maintaining pixel- and region-based accuracy.

Solving Linear Systems in $\widetilde{O}(mn \log \fracκε)$ Bit Operations

from arXiv: Data Structures and Algorithms

Authors: Jonathan A. Kelner

We give a deterministic algorithm that solves a nonsingular linear system $Ax=b$, where $A\in\mathbb{R}^{n\times n}$ has $m$ nonzero entries and condition number $κ$, to any relative residual tolerance $0<ε\le1/2$ using $\widetilde{O(}mn\log(κ/ε))$ bit operations for inputs with logarithmically many bits per entry. For sparse, polynomially conditioned systems with $m=\widetilde{O}(n)$, this gives an $\widetilde{O}(n^2)$ algorithm for any inverse-polynomial accuracy, improving on the algorithm of Peng and Vempala, as sharpened by Nie, whose running time in this regime is approximately $O(n^{2.2707})$ with the best current matrix multiplication exponent, and largely closing a gap between the idealized performance of the conjugate gradient method in exact arithmetic and the running time achievable with finite-precision computation that has persisted for over 70 years. The algorithm is surprisingly simple. For integer inputs, we apply Dixon's lifting algorithm to the perturbed system $(A+I/R)x=b$ for a suitable integer $R$. After scaling, the matrix of this system is $RA+I\equiv I\pmod R$, so its modular inverse is trivial, and each lifting step needs only a sparse matrix-vector product with $A$ on $O(\log R)$-bit numbers. Fast rational reconstruction then recovers the exact solution of the perturbed system, which is an $ε$-accurate solution of the original one. Normalization and rounding extend the result to fixed-point and floating-point inputs, with floating-point outputs represented using short integer significands and a common encoded exponent. A computable certificate removes the need for prior knowledge of $κ$.

Authors: Jonathan A. Kelner

We give a deterministic algorithm that solves a nonsingular linear system $Ax=b$, where $A\in\mathbb{R}^{n\times n}$ has $m$ nonzero entries and condition number $κ$, to any relative residual tolerance $0<ε\le1/2$ using $\widetilde{O(}mn\log(κ/ε))$ bit operations for inputs with logarithmically many bits per entry. For sparse, polynomially conditioned systems with $m=\widetilde{O}(n)$, this gives an $\widetilde{O}(n^2)$ algorithm for any inverse-polynomial accuracy, improving on the algorithm of Peng and Vempala, as sharpened by Nie, whose running time in this regime is approximately $O(n^{2.2707})$ with the best current matrix multiplication exponent, and largely closing a gap between the idealized performance of the conjugate gradient method in exact arithmetic and the running time achievable with finite-precision computation that has persisted for over 70 years. The algorithm is surprisingly simple. For integer inputs, we apply Dixon's lifting algorithm to the perturbed system $(A+I/R)x=b$ for a suitable integer $R$. After scaling, the matrix of this system is $RA+I\equiv I\pmod R$, so its modular inverse is trivial, and each lifting step needs only a sparse matrix-vector product with $A$ on $O(\log R)$-bit numbers. Fast rational reconstruction then recovers the exact solution of the perturbed system, which is an $ε$-accurate solution of the original one. Normalization and rounding extend the result to fixed-point and floating-point inputs, with floating-point outputs represented using short integer significands and a common encoded exponent. A computable certificate removes the need for prior knowledge of $κ$.

Query Complexity of Testing Structured Parenthesis Languages

from arXiv: Data Structures and Algorithms

Authors: Tim Jackman, Diptaksho Palit, Sofya Raskhodnikova

We study the query complexity of testing membership in structured string languages, focusing on Dyck languages and natural generalizations. A tester receives query access to a word and must distinguish valid inputs from words that are far in Hamming distance, while inspecting only a sublinear number of positions. Our results sharpen the boundary between constant-query testability and polynomial query complexity. First, we prove an $Ω(n^{2/5})$ lower bound for testing Dyck languages $D_m$ with any fixed number $m\ge2$ of parenthesis types, improving the previous $Ω(n^{1/5})$ lower bound of Fischer, Magniez, and Starikovskaya (SODA `18) and nearly matching their upper bound of $O(n^{2/5+o(1)})$. Furthermore, we show that all nonadaptive algorithms for these problems require $Ω(n^{1/2})$ queries. Our lower bounds use a Pólya-urn process to construct the hard distribution; Second, we identify a broad class of weighted-parenthesis languages, which we call {\em excursion languages,} that remain constant-query testable. These languages encode bounded-step walks that stay nonnegative and return to zero. For every fixed excursion language, we give a nonadaptive tester with query complexity $O(1/\varepsilon^2)$, and we prove this dependence on $\varepsilon$ is optimal, even for adaptive algorithms. As a special case, we obtain the tight $Θ(1/\varepsilon^2)$ query complexity of testing $D_1$, improving the previous $O(\log(1/\varepsilon)/\varepsilon^2)$ upper bound and giving the first matching two-sided-error lower bound. Third, we construct a simple hard language, Hidden String, that is generated by a deterministic linear grammar but nevertheless requires $Ω(n^{2/5})$ adaptive queries and $Ω(n^{1/2})$ nonadaptive queries to test. This shows that polynomial query complexity appears even for highly restricted string languages.

Authors: Tim Jackman, Diptaksho Palit, Sofya Raskhodnikova

We study the query complexity of testing membership in structured string languages, focusing on Dyck languages and natural generalizations. A tester receives query access to a word and must distinguish valid inputs from words that are far in Hamming distance, while inspecting only a sublinear number of positions. Our results sharpen the boundary between constant-query testability and polynomial query complexity. First, we prove an $Ω(n^{2/5})$ lower bound for testing Dyck languages $D_m$ with any fixed number $m\ge2$ of parenthesis types, improving the previous $Ω(n^{1/5})$ lower bound of Fischer, Magniez, and Starikovskaya (SODA `18) and nearly matching their upper bound of $O(n^{2/5+o(1)})$. Furthermore, we show that all nonadaptive algorithms for these problems require $Ω(n^{1/2})$ queries. Our lower bounds use a Pólya-urn process to construct the hard distribution; Second, we identify a broad class of weighted-parenthesis languages, which we call {\em excursion languages,} that remain constant-query testable. These languages encode bounded-step walks that stay nonnegative and return to zero. For every fixed excursion language, we give a nonadaptive tester with query complexity $O(1/\varepsilon^2)$, and we prove this dependence on $\varepsilon$ is optimal, even for adaptive algorithms. As a special case, we obtain the tight $Θ(1/\varepsilon^2)$ query complexity of testing $D_1$, improving the previous $O(\log(1/\varepsilon)/\varepsilon^2)$ upper bound and giving the first matching two-sided-error lower bound. Third, we construct a simple hard language, Hidden String, that is generated by a deterministic linear grammar but nevertheless requires $Ω(n^{2/5})$ adaptive queries and $Ω(n^{1/2})$ nonadaptive queries to test. This shows that polynomial query complexity appears even for highly restricted string languages.

Collision Detection is Instance $\widetilde{O}$ptimal Under the Birthday Threshold

from arXiv: Data Structures and Algorithms

Authors: Omri Ben-Eliezer, Tomer Grossman, Václav Rozhoň, Jakub Tětek

Can structural knowledge about a hash function help accelerate the (black box) detection of collisions in it? This question is fundamental to cryptography theory given the importance of collision-resistant hash functions, and in this paper we tackle it from the angle of instance optimality, an ultimate notion of beyond worst case algorithm analysis that has gained significant traction in recent years. Instance optimality asks for a single algorithm that, on every input, performs nearly as well as the best correct algorithm that ``knows the structure'' of that specific input. Here we measure algorithms by the number of queries they make to the hash function $f\colon [n]\to [n]$, and we say that an algorithm ``knows the structure'' of the input if, in addition to query access to $f$, it has free access to an unlabeled copy $π^{-1}\circ f\circπ$ of $f$, for an unknown permutation $π$ on $[n]$. We prove the existence of an (almost) instance-optimal algorithm for collision detection in the regime most interesting from a cryptographic perspective: among functions where finding a collision takes significantly less than $\sqrt{n}$ queries. Specifically, we prove the existence of a single algorithm $A$ that, for any input $f$ in which a structure-aware algorithm can find a collision using $q\leq O(\sqrt{n/\log n})$ queries in expectation, $A$ can find a collision in at most $O(q\log n)$ queries. The $O(\log n)$ multiplicative overhead is tight, matching a lower bound of Ben-Eliezer, Grossman, and Naor [ICALP'25], and partially resolving their main open question. Our result implies, in particular, that it is impossible for a cryptographic designer to plant purely structural backdoors for collision finding (for this unlabeled notion of structure): whatever collisions the designer's secret knowledge finds, the public can find with a multiplicative overhead of $O(\log n)$.

Authors: Omri Ben-Eliezer, Tomer Grossman, Václav Rozhoň, Jakub Tětek

Can structural knowledge about a hash function help accelerate the (black box) detection of collisions in it? This question is fundamental to cryptography theory given the importance of collision-resistant hash functions, and in this paper we tackle it from the angle of instance optimality, an ultimate notion of beyond worst case algorithm analysis that has gained significant traction in recent years. Instance optimality asks for a single algorithm that, on every input, performs nearly as well as the best correct algorithm that ``knows the structure'' of that specific input. Here we measure algorithms by the number of queries they make to the hash function $f\colon [n]\to [n]$, and we say that an algorithm ``knows the structure'' of the input if, in addition to query access to $f$, it has free access to an unlabeled copy $π^{-1}\circ f\circπ$ of $f$, for an unknown permutation $π$ on $[n]$. We prove the existence of an (almost) instance-optimal algorithm for collision detection in the regime most interesting from a cryptographic perspective: among functions where finding a collision takes significantly less than $\sqrt{n}$ queries. Specifically, we prove the existence of a single algorithm $A$ that, for any input $f$ in which a structure-aware algorithm can find a collision using $q\leq O(\sqrt{n/\log n})$ queries in expectation, $A$ can find a collision in at most $O(q\log n)$ queries. The $O(\log n)$ multiplicative overhead is tight, matching a lower bound of Ben-Eliezer, Grossman, and Naor [ICALP'25], and partially resolving their main open question. Our result implies, in particular, that it is impossible for a cryptographic designer to plant purely structural backdoors for collision finding (for this unlabeled notion of structure): whatever collisions the designer's secret knowledge finds, the public can find with a multiplicative overhead of $O(\log n)$.

Three-Color Free-Flood-It on Fixed-Height Grids Is Polynomial-Time Solvable

from arXiv: Data Structures and Algorithms

Authors: Yuxuan Zhou

We give a deterministic algorithm for \textsc{Free-Flood-It} on rectangular grids $P_k\square P_n$ with at most three colors. For every fixed height $k$, it computes the minimum number of moves and an optimal sequence in $N^{O(k^2)}$ time, where $N=kn$. This resolves the previously open three-color case on complete $3\times n$ boards. The proof uses a representation of flooding strategies as paintings by connected regions. We show that an optimal painting can be chosen so that its regions form a rooted tree with strong restrictions on the colors along ancestral paths. These restrictions make the part of the tree visible in any fixed-size connected window admit only polynomially many descriptions. The remaining common ancestors may form an arbitrarily long chain. We retain that chain on a stack and use a finite context-free recurrence to minimize the total painting cost without enumerating all possible stack contents. Connectivity information at the boundary ensures that the resulting local descriptions assemble into one valid global painting. The argument also applies to graphs supplied with an ordering into connected layers of bounded size, provided every edge lies within one layer or joins consecutive layers. A family of three-row boards shows that the unbounded ancestor chains handled by the algorithm are necessary even for optimal paintings. For any fixed height and palette size, we also give a deterministic EPTAS and a randomized sampling variant with an explicit failure-probability bound. These approximation results use a separate algorithm with single-exponential dependence on the move budget.

Authors: Yuxuan Zhou

We give a deterministic algorithm for \textsc{Free-Flood-It} on rectangular grids $P_k\square P_n$ with at most three colors. For every fixed height $k$, it computes the minimum number of moves and an optimal sequence in $N^{O(k^2)}$ time, where $N=kn$. This resolves the previously open three-color case on complete $3\times n$ boards. The proof uses a representation of flooding strategies as paintings by connected regions. We show that an optimal painting can be chosen so that its regions form a rooted tree with strong restrictions on the colors along ancestral paths. These restrictions make the part of the tree visible in any fixed-size connected window admit only polynomially many descriptions. The remaining common ancestors may form an arbitrarily long chain. We retain that chain on a stack and use a finite context-free recurrence to minimize the total painting cost without enumerating all possible stack contents. Connectivity information at the boundary ensures that the resulting local descriptions assemble into one valid global painting. The argument also applies to graphs supplied with an ordering into connected layers of bounded size, provided every edge lies within one layer or joins consecutive layers. A family of three-row boards shows that the unbounded ancestor chains handled by the algorithm are necessary even for optimal paintings. For any fixed height and palette size, we also give a deterministic EPTAS and a randomized sampling variant with an explicit failure-probability bound. These approximation results use a separate algorithm with single-exponential dependence on the move budget.

Degree Balance as a Fine-Grained Complexity Boundary for Quantum SAT

from arXiv: Data Structures and Algorithms

Authors: Atsuya Hasegawa, Jonas Kamminga, François Le Gall, Suguru Tamaki

The local Hamiltonian problem is the canonical $\mathsf{QMA}$-complete problem, and $O(2^n)$ time classical algorithms and $O(2^{n/2})$ time quantum algorithms are known to solve the problem in the worst case. It is not clear how to improve these brute force strategies for a broad class of the problem because ground states are highly entangled in general, and we cannot directly apply known strategies for classical CSPs. In this work, we present exponentially faster classical and quantum algorithms under two mild assumptions: (1) the Hamiltonian is frustration-free on YES instances, and (2) it is approximately regular, meaning that every qubit is acted upon by approximately the same number of constraints. We complement these upper bounds by showing that, assuming (Q)SETH, quantum 5-SAT admits no non-trivial worst-case speedup. Our lower bound further demonstrates that the dependence of our algorithms on regularity is in some sense nearly optimal. Specifically, quantum 5-SAT remains (Q)SETH-hard even for Hamiltonians in which all but $O(\sqrt{n})$ qubits participate in only constantly many constraints, while the remaining $O(\sqrt{n})$ qubits each participate in $O(\sqrt{n})$ constraints. By contrast, if either the size of this high-degree subset or the degrees of its qubits is reduced by a factor of $n^δ$, for any $δ>0$, our algorithm solves the problem in time $O(2^{(1- \varepsilon)n})$ for some $\varepsilon>0$. Together, our upper and lower bounds establish a fine-grained complexity dichotomy for quantum satisfiability.

Authors: Atsuya Hasegawa, Jonas Kamminga, François Le Gall, Suguru Tamaki

The local Hamiltonian problem is the canonical $\mathsf{QMA}$-complete problem, and $O(2^n)$ time classical algorithms and $O(2^{n/2})$ time quantum algorithms are known to solve the problem in the worst case. It is not clear how to improve these brute force strategies for a broad class of the problem because ground states are highly entangled in general, and we cannot directly apply known strategies for classical CSPs. In this work, we present exponentially faster classical and quantum algorithms under two mild assumptions: (1) the Hamiltonian is frustration-free on YES instances, and (2) it is approximately regular, meaning that every qubit is acted upon by approximately the same number of constraints. We complement these upper bounds by showing that, assuming (Q)SETH, quantum 5-SAT admits no non-trivial worst-case speedup. Our lower bound further demonstrates that the dependence of our algorithms on regularity is in some sense nearly optimal. Specifically, quantum 5-SAT remains (Q)SETH-hard even for Hamiltonians in which all but $O(\sqrt{n})$ qubits participate in only constantly many constraints, while the remaining $O(\sqrt{n})$ qubits each participate in $O(\sqrt{n})$ constraints. By contrast, if either the size of this high-degree subset or the degrees of its qubits is reduced by a factor of $n^δ$, for any $δ>0$, our algorithm solves the problem in time $O(2^{(1- \varepsilon)n})$ for some $\varepsilon>0$. Together, our upper and lower bounds establish a fine-grained complexity dichotomy for quantum satisfiability.

Local Search for Fair Max-Min Diversification

from arXiv: Data Structures and Algorithms

Authors: Sepideh Mahabadi, Shyam Narayanan, Varun Sivashankar

Given $n$ points in a metric space, Max-Min diversification asks for a subset of $k$ points maximizing the minimum pairwise distance between the selected points. This is arguably the most fundamental notion of diversity with applications across a wide range of domains. We consider this problem under partition constraints, previously studied as Fair Max-Min Diversification (FMMD). Here, each point has a color in $[m]$, and a feasible solution must contain exactly $k_i$ points of color $i$, where $k_1,\ldots,k_m$ are prescribed parameters satisfying $\sum_i k_i=k$. We give the first constant factor approximation for the problem using local search, that runs in time $f(m)\cdot \operatorname{poly}(n)$, in which all constraints are satisfied exactly. All previously known algorithms either provided an $\widetilde Θ(m)$ approximation factor, had running times exponential in the solution size $k$, or satisfied the fairness constraints only approximately or in expectation. We further generalize our result to the problem where each point may belong to an arbitrary subset of colors. Given lower and upper bounds $\ell_i$ and $u_i$ for every color $i$, the goal is to find $k$ points whose color counts satisfy all these bounds while maximizing their diversity.

Authors: Sepideh Mahabadi, Shyam Narayanan, Varun Sivashankar

Given $n$ points in a metric space, Max-Min diversification asks for a subset of $k$ points maximizing the minimum pairwise distance between the selected points. This is arguably the most fundamental notion of diversity with applications across a wide range of domains. We consider this problem under partition constraints, previously studied as Fair Max-Min Diversification (FMMD). Here, each point has a color in $[m]$, and a feasible solution must contain exactly $k_i$ points of color $i$, where $k_1,\ldots,k_m$ are prescribed parameters satisfying $\sum_i k_i=k$. We give the first constant factor approximation for the problem using local search, that runs in time $f(m)\cdot \operatorname{poly}(n)$, in which all constraints are satisfied exactly. All previously known algorithms either provided an $\widetilde Θ(m)$ approximation factor, had running times exponential in the solution size $k$, or satisfied the fairness constraints only approximately or in expectation. We further generalize our result to the problem where each point may belong to an arbitrary subset of colors. Given lower and upper bounds $\ell_i$ and $u_i$ for every color $i$, the goal is to find $k$ points whose color counts satisfy all these bounds while maximizing their diversity.

A Tale of Two Walks: Kipnis, Marchioro and Presutti Meet Kac in a Quantum World

from arXiv: Data Structures and Algorithms

Authors: Qian Chen, Jingcheng Liu, Minglong Qin, Leonard Schulman, Fang Song, Penghui Yao, Mingnan Zhao

We reveal an unexpected connection between the parallel Kac's walk and the Kipnis-Marchioro-Presutti (KMP) process. The twirling channel induced by the parallel Kac's walk on the symmetric subspace is exactly encoded by a classical Markov chain on partitions, which lifts to a parallel KMP process on complete graphs. This correspondence reduces the analysis of the twirling channel to the mixing of the parallel KMP process. We prove that $O(\log d+\log(1/\varepsilon))$ repetitions suffice to approximate Haar twirling on the symmetric subspace of $(\mathbb C^d)^{\otimes t}$ to error $\varepsilon$, uniformly in the number of copies $t$. For the standard KMP process on general graphs, we prove a mixing-time analogue of Aldous's conjecture: at fixed accuracy, the mixing time of the $t$-particle process is at most a constant times the single-particle mixing time multiplied by the logarithm of the number of vertices, uniformly in $t$. As an application, we improve the total variation mixing-time bound for coordinate hit-and-run on the $n$-dimensional standard simplex from $\widetilde O(n^3)$ (Kook and Vempala, 2026) to $\widetilde O(n)$, while removing the dependence on the initial distribution. Our main technical contribution is conditional product structure for both parallel and standard KMP processes. Conditioned on suitable auxiliary randomness, the labeled particles evolve independently. Combining this structure with an exact coupling yields mixing bounds uniform in the number of particles for both unlabeled KMP models. These bounds are sharp up to logarithmic factors and imply rapid convergence of the parallel Kac twirling channel on the symmetric subspace.

Authors: Qian Chen, Jingcheng Liu, Minglong Qin, Leonard Schulman, Fang Song, Penghui Yao, Mingnan Zhao

We reveal an unexpected connection between the parallel Kac's walk and the Kipnis-Marchioro-Presutti (KMP) process. The twirling channel induced by the parallel Kac's walk on the symmetric subspace is exactly encoded by a classical Markov chain on partitions, which lifts to a parallel KMP process on complete graphs. This correspondence reduces the analysis of the twirling channel to the mixing of the parallel KMP process. We prove that $O(\log d+\log(1/\varepsilon))$ repetitions suffice to approximate Haar twirling on the symmetric subspace of $(\mathbb C^d)^{\otimes t}$ to error $\varepsilon$, uniformly in the number of copies $t$. For the standard KMP process on general graphs, we prove a mixing-time analogue of Aldous's conjecture: at fixed accuracy, the mixing time of the $t$-particle process is at most a constant times the single-particle mixing time multiplied by the logarithm of the number of vertices, uniformly in $t$. As an application, we improve the total variation mixing-time bound for coordinate hit-and-run on the $n$-dimensional standard simplex from $\widetilde O(n^3)$ (Kook and Vempala, 2026) to $\widetilde O(n)$, while removing the dependence on the initial distribution. Our main technical contribution is conditional product structure for both parallel and standard KMP processes. Conditioned on suitable auxiliary randomness, the labeled particles evolve independently. Combining this structure with an exact coupling yields mixing bounds uniform in the number of particles for both unlabeled KMP models. These bounds are sharp up to logarithmic factors and imply rapid convergence of the parallel Kac twirling channel on the symmetric subspace.

Can We Break Fine-Grained and NP-Hardness Barriers if We've Seen the Graph Before? The Isomorphic-Priors Model

from arXiv: Data Structures and Algorithms

Authors: Dani Dorfmann, Simon Döring, Martin G. Herold, Danupon Nanongkai, Daniel Neuen, Joachim Spoerhase, Zihang Wu

If we run a heavy-duty computation on prior data, can we avoid repeated computation for similar future inputs? Inspired by this question, we introduce a new computational model for graph problems called algorithms with isomorphic priors. Solving a graph problem $Π$ in this model involves two phases: (i) The preprocessing phase quickly analyzes prior graphs $G_1, ..., G_k$ along with the (previously computed) exact optimal values OPT$(G_i)$. (ii) Subsequently, given a new graph H, a fast query phase must either (a) output the exact solution OPT(H), or (b) correctly report that H is not isomorphic to any $G_i$. Can we avoid computing OPT(H) from scratch when H is isomorphic to some $G_i$? We show that this is the case for a number of problems; for many others, we establish conditional lower bounds. $\textbf{(1)}$ Some NP-hard problems, including Constrained Shortest Path and $\ell_p$-Shortest Path and Constrained Spanning Tree, admit polynomial preprocessing and query times in our model. In contrast, almost all of Karp's 21 NP-complete problems and $(2-\varepsilon)$-approximate $k$-Center, for every fixed $\varepsilon>0$, admit no such algorithms unless Graph Isomorphism (GI) is in P, even with O(1) priors. $\textbf{(2)}$ In contrast to conditional $n^{3-o(1)}$ fine-grained lower bounds, our framework achieves an $O(n^ω)$ query time for Negative Triangle and a near-linear query time for Replacement Path. It also achieves near-linear query time for Maximum Flow. $\textbf{(3)}$ While it remains a major open problem whether infinite-duration games (Paritiy Game, Mean Payoff Game, Energy Game, and Stochastic Game) admit polynomial-time algorithms, they can be easily solved in near-linear time within our model. Our proofs rely on a simple combination of existing tools and are accessible to readers without specialized background.

Authors: Dani Dorfmann, Simon Döring, Martin G. Herold, Danupon Nanongkai, Daniel Neuen, Joachim Spoerhase, Zihang Wu

If we run a heavy-duty computation on prior data, can we avoid repeated computation for similar future inputs? Inspired by this question, we introduce a new computational model for graph problems called algorithms with isomorphic priors. Solving a graph problem $Π$ in this model involves two phases: (i) The preprocessing phase quickly analyzes prior graphs $G_1, ..., G_k$ along with the (previously computed) exact optimal values OPT$(G_i)$. (ii) Subsequently, given a new graph H, a fast query phase must either (a) output the exact solution OPT(H), or (b) correctly report that H is not isomorphic to any $G_i$. Can we avoid computing OPT(H) from scratch when H is isomorphic to some $G_i$? We show that this is the case for a number of problems; for many others, we establish conditional lower bounds. $\textbf{(1)}$ Some NP-hard problems, including Constrained Shortest Path and $\ell_p$-Shortest Path and Constrained Spanning Tree, admit polynomial preprocessing and query times in our model. In contrast, almost all of Karp's 21 NP-complete problems and $(2-\varepsilon)$-approximate $k$-Center, for every fixed $\varepsilon>0$, admit no such algorithms unless Graph Isomorphism (GI) is in P, even with O(1) priors. $\textbf{(2)}$ In contrast to conditional $n^{3-o(1)}$ fine-grained lower bounds, our framework achieves an $O(n^ω)$ query time for Negative Triangle and a near-linear query time for Replacement Path. It also achieves near-linear query time for Maximum Flow. $\textbf{(3)}$ While it remains a major open problem whether infinite-duration games (Paritiy Game, Mean Payoff Game, Energy Game, and Stochastic Game) admit polynomial-time algorithms, they can be easily solved in near-linear time within our model. Our proofs rely on a simple combination of existing tools and are accessible to readers without specialized background.

Vanishing Ideals and the Computational Tractability of Sum-of-Squares over Boolean Domains

from arXiv: Data Structures and Algorithms

Authors: Monaldo Mastrolilli

Building on the bit-complexity framework of Raghavendra-Weitz and the moment-SOS criteria of Gribling-Polak-Slot, we study the effective use of truncated vanishing identities over Boolean polynomial systems. A complete, constructible identity space gives an augmented moment SDP that can be optimized with exact rational feasibility and arbitrary additive accuracy. We give a self-contained geometric implementation: an explicit affine reduction and a simplex of Boolean evaluations supply the radius bounds required by rational ellipsoids. The identity space can be constructed directly or extracted from a supplied graded Groebner basis, including one supplied at a higher truncation degree. An explicit transfer theorem connects this augmented formulation to the original system. Two-sided SoS derivations eliminate the added equality axioms from certificates, with controlled degree and coefficient growth, and imply containment of a projected higher-level moment relaxation in the augmented body. Together with spectral coefficient bounds, this gives polynomial-time search for rational proofs with an additive perturbation. For Min-closed linear systems, propagation constructs the identity space in polynomial time at fixed degree and gives degree-(4t+4) certificates for both signs of each degree-t basis element. Consequently, a rational degree-(8d+4) proof of f + epsilon >= 0 can be found in polynomial time for fixed d whenever f >= 0 has a degree-2d proof. The augmented degree-2d moment SDP can be optimized in polynomial time with exact rational feasibility and a comparison to the original degree-(8d+4) relaxation. Boolean complementation gives the same results for Max-closed systems, including generalized packing and covering. Min-closed systems thus provide a concrete application of the general criteria. All complexity bounds are in the Turing model.

Authors: Monaldo Mastrolilli

Building on the bit-complexity framework of Raghavendra-Weitz and the moment-SOS criteria of Gribling-Polak-Slot, we study the effective use of truncated vanishing identities over Boolean polynomial systems. A complete, constructible identity space gives an augmented moment SDP that can be optimized with exact rational feasibility and arbitrary additive accuracy. We give a self-contained geometric implementation: an explicit affine reduction and a simplex of Boolean evaluations supply the radius bounds required by rational ellipsoids. The identity space can be constructed directly or extracted from a supplied graded Groebner basis, including one supplied at a higher truncation degree. An explicit transfer theorem connects this augmented formulation to the original system. Two-sided SoS derivations eliminate the added equality axioms from certificates, with controlled degree and coefficient growth, and imply containment of a projected higher-level moment relaxation in the augmented body. Together with spectral coefficient bounds, this gives polynomial-time search for rational proofs with an additive perturbation. For Min-closed linear systems, propagation constructs the identity space in polynomial time at fixed degree and gives degree-(4t+4) certificates for both signs of each degree-t basis element. Consequently, a rational degree-(8d+4) proof of f + epsilon >= 0 can be found in polynomial time for fixed d whenever f >= 0 has a degree-2d proof. The augmented degree-2d moment SDP can be optimized in polynomial time with exact rational feasibility and a comparison to the original degree-(8d+4) relaxation. Boolean complementation gives the same results for Max-closed systems, including generalized packing and covering. Min-closed systems thus provide a concrete application of the general criteria. All complexity bounds are in the Turing model.

Byzantine Causal Reliable Broadcast with Constant Metadata Overhead

from arXiv: Data Structures and Algorithms

Authors: Purv Patel, Ajay D. Kshemkalyani

Asynchronous Byzantine Reliable Broadcast (BRB) is a fundamental primitive that guarantees agreement and validity in distributed systems subject to Byzantine faults, but it lacks ordering guarantees. Causal message ordering is important for many applications such as blockchain and social networking. Existing solutions for Byzantine Causal Reliable Broadcast (BCRB) have several drawbacks. Such protocols typically append vector clocks or dependency barriers to application messages, resulting in a metadata overhead that scales linearly with $n$, the number of processes in the system. In this paper, we propose the first optimal message overhead BCRB algorithm that guarantees safety. We do this by re-engineering Bracha's BRB algorithm with relatively small but critical modifications, and prove that our algorithm solves BCRB with optimal message metadata. The algorithm achieves constant-size $\mathcal{O}(1)$ message metadata overhead and $\mathcal{O}(n^2)$ messages, resulting in $\mathcal{O}(n^2)$ communication word complexity. This is as against $\mathcal{O}(n^3)$ communication word complexity of existing protocols. The algorithm tolerates $f < n/3$ Byzantine processes, which is the well-known optimal resilience bound, and uses four phases.

Authors: Purv Patel, Ajay D. Kshemkalyani

Asynchronous Byzantine Reliable Broadcast (BRB) is a fundamental primitive that guarantees agreement and validity in distributed systems subject to Byzantine faults, but it lacks ordering guarantees. Causal message ordering is important for many applications such as blockchain and social networking. Existing solutions for Byzantine Causal Reliable Broadcast (BCRB) have several drawbacks. Such protocols typically append vector clocks or dependency barriers to application messages, resulting in a metadata overhead that scales linearly with $n$, the number of processes in the system. In this paper, we propose the first optimal message overhead BCRB algorithm that guarantees safety. We do this by re-engineering Bracha's BRB algorithm with relatively small but critical modifications, and prove that our algorithm solves BCRB with optimal message metadata. The algorithm achieves constant-size $\mathcal{O}(1)$ message metadata overhead and $\mathcal{O}(n^2)$ messages, resulting in $\mathcal{O}(n^2)$ communication word complexity. This is as against $\mathcal{O}(n^3)$ communication word complexity of existing protocols. The algorithm tolerates $f < n/3$ Byzantine processes, which is the well-known optimal resilience bound, and uses four phases.

XBDD: A Highly Optimized ROBDD with Per-Edge Variable-Flip Maps

from arXiv: Data Structures and Algorithms

Authors: Yinglong Gan, Jintao Yu, Shenggang Ying, Yusen Li, Xin Hong

The Reduced Ordered Binary Decision Diagram (ROBDD) is a canonical representation of Boolean functions and is widely used in tasks such as equivalence checking and satisfiability checking of combinational circuits. Classical ROBDD packages greatly improve the efficiency of building ROBDDs through a series of optimization techniques, and compress the node scale of the ROBDD through complement edges. However, existing implementations do not take into account the local polarity differences of isomorphic Boolean functions, and still produce a distinct node for each polarity combination, thereby causing an explosion in the number of nodes. This paper proposes XBDD, a highly optimized ROBDD that, on the basis of fully implementing complement edges and their accompanying engineering techniques, introduces a per-edge variable-flip map. XBDD attaches a flip map to each edge to indicate which input variables must be negated when that edge is followed. This allows nodes that differ only in local input polarities to be merged, further reducing the node count. For certain function families, this sharing even yields exponential compression. We also propose methods that use a bitmap and a map pool to substantially reduce the extra overhead brought by the map, and propose normalization and cofactor operators for the map. In addition, XBDD implements several other engineering optimizations to further improve both time and space efficiency. Experiments show that XBDD trades a controllable time cost for a significant space gain, validating the effectiveness of the per-edge variable-flip map.

Authors: Yinglong Gan, Jintao Yu, Shenggang Ying, Yusen Li, Xin Hong

The Reduced Ordered Binary Decision Diagram (ROBDD) is a canonical representation of Boolean functions and is widely used in tasks such as equivalence checking and satisfiability checking of combinational circuits. Classical ROBDD packages greatly improve the efficiency of building ROBDDs through a series of optimization techniques, and compress the node scale of the ROBDD through complement edges. However, existing implementations do not take into account the local polarity differences of isomorphic Boolean functions, and still produce a distinct node for each polarity combination, thereby causing an explosion in the number of nodes. This paper proposes XBDD, a highly optimized ROBDD that, on the basis of fully implementing complement edges and their accompanying engineering techniques, introduces a per-edge variable-flip map. XBDD attaches a flip map to each edge to indicate which input variables must be negated when that edge is followed. This allows nodes that differ only in local input polarities to be merged, further reducing the node count. For certain function families, this sharing even yields exponential compression. We also propose methods that use a bitmap and a map pool to substantially reduce the extra overhead brought by the map, and propose normalization and cofactor operators for the map. In addition, XBDD implements several other engineering optimizations to further improve both time and space efficiency. Experiments show that XBDD trades a controllable time cost for a significant space gain, validating the effectiveness of the per-edge variable-flip map.

The Reach of Abelian Covers in Hypergraphs

from arXiv: Data Structures and Algorithms

Authors: Joshua Brakensiek, Venkatesan Guruswami, Aaron Putterman

Covers in hypergraphs are frequently studied to capture various forms of dependence between hyperedges. For example, even covers--which check if each vertex appears in an even number of hyperedges--have found much success recently in the study of locally decodable codes. Inspired by a recently-emerging line of work on the non-redundancy of constraint satisfaction problems (CSPs), we introduce and study two novel families of covers of hypergraphs which are stricter than even covers: \emph{Abelian} covers and Catalan covers. Abelian covers are similar to even covers, except that arithmetic is now done over the integers rather than modulo 2, allowing us to capture dependences over arbitrary Abelian groups. Catalan covers capture the behavior of non-Abelian groups by only allowing local cancellations in a sequence of hyperedges. We prove three main results about Abelian and Catalan covers. First, using tools from lattice theory, we show that any $r$-uniform hypergraph with $n$ vertices and $n \log(r)$ hyperedges has an Abelian cover. Second, using tools from algebraic topology, we show that in any $3$-uniform hypergraph, Abelian covers and Catalan covers are equivalent; thereby showing that Catalan covers emerge after $O(n)$ hyperedges in $3$-uniform hypergraphs. Finally, using the theory of nilpotent groups, we show that there exists a $4$-uniform hypergraph which has an Abelian cover but not a Catalan cover. Collectively, these results exactly characterize the reach that Abelian covers have in deducing dependences in hypergraphs. As our primary application, we show that any arity-$3$ CSP with an infinite-domain Mal'tsev extension has linear non-redundancy. This implies near optimal streaming, sparsification, and kernelization algorithms for this family of CSPs. Previously, such a result was only known for the much simpler case of arity-$2$ CSPs.

Authors: Joshua Brakensiek, Venkatesan Guruswami, Aaron Putterman

Covers in hypergraphs are frequently studied to capture various forms of dependence between hyperedges. For example, even covers--which check if each vertex appears in an even number of hyperedges--have found much success recently in the study of locally decodable codes. Inspired by a recently-emerging line of work on the non-redundancy of constraint satisfaction problems (CSPs), we introduce and study two novel families of covers of hypergraphs which are stricter than even covers: \emph{Abelian} covers and Catalan covers. Abelian covers are similar to even covers, except that arithmetic is now done over the integers rather than modulo 2, allowing us to capture dependences over arbitrary Abelian groups. Catalan covers capture the behavior of non-Abelian groups by only allowing local cancellations in a sequence of hyperedges. We prove three main results about Abelian and Catalan covers. First, using tools from lattice theory, we show that any $r$-uniform hypergraph with $n$ vertices and $n \log(r)$ hyperedges has an Abelian cover. Second, using tools from algebraic topology, we show that in any $3$-uniform hypergraph, Abelian covers and Catalan covers are equivalent; thereby showing that Catalan covers emerge after $O(n)$ hyperedges in $3$-uniform hypergraphs. Finally, using the theory of nilpotent groups, we show that there exists a $4$-uniform hypergraph which has an Abelian cover but not a Catalan cover. Collectively, these results exactly characterize the reach that Abelian covers have in deducing dependences in hypergraphs. As our primary application, we show that any arity-$3$ CSP with an infinite-domain Mal'tsev extension has linear non-redundancy. This implies near optimal streaming, sparsification, and kernelization algorithms for this family of CSPs. Previously, such a result was only known for the much simpler case of arity-$2$ CSPs.

On Extensions of the Unanimous Vote Problem

from arXiv: Data Structures and Algorithms

Authors: Evan J. R. Brody, Haya Diwan, Lisa Hellerstein, Thomas Lidbetter

The Unanimous Vote problem is to determine a fixed order in which to flip each of $n$ biased coins, where each coin can be flipped only once, such that the expected number of flips until seeing both a head and a tail (or flipping all coins) is minimized. Duman Keles et al. (arXiv:2510.16678 [cs.DS]) gave an $\mathcal{O}(n \log n)$-time algorithm for this problem. Extensions of the Unanimous Vote problem are a rich source of stochastic optimization problems. We focus on three: (1) a variant in which each coin can be flipped arbitrarily many times (a solution is thus an infinite sequence of coin choices), (2) a generalization with $d$-sided dice, that can each be rolled once, where dice must be rolled until two different outcomes are observed (or all dice have been rolled), and (3) a different generalization with $d$-sided dice, where dice must be rolled until all $d$ outcomes have been observed. For (1), we show that there is an optimal sequence which follows a simple greedy rule; the same rule only gives a 1-additive approximation for the original problem (arXiv:2510.16678 [cs.DS]). The rule also yields a correspondence between a particular optimal sequence and a related mechanical word, which we exploit to characterize the conditions under which this optimal sequence is periodic. We establish tight multiplicative and additive adaptivity gaps for this variant. For (2), we show that two different generalizations of the greedy rule from (arXiv:2510.16678 [cs.DS]) can be combined to obtain a PTAS. For (3), we give an $\mathcal{O}(\log d)$-approximation algorithm by reducing the problem to Submodular Ranking (arXiv:1007.2503 [cs.DS]); the same reduction technique can be used to yield approximation algorithms for other stochastic probing problems. Finally, we pose a number of related open questions.

Authors: Evan J. R. Brody, Haya Diwan, Lisa Hellerstein, Thomas Lidbetter

The Unanimous Vote problem is to determine a fixed order in which to flip each of $n$ biased coins, where each coin can be flipped only once, such that the expected number of flips until seeing both a head and a tail (or flipping all coins) is minimized. Duman Keles et al. (arXiv:2510.16678 [cs.DS]) gave an $\mathcal{O}(n \log n)$-time algorithm for this problem. Extensions of the Unanimous Vote problem are a rich source of stochastic optimization problems. We focus on three: (1) a variant in which each coin can be flipped arbitrarily many times (a solution is thus an infinite sequence of coin choices), (2) a generalization with $d$-sided dice, that can each be rolled once, where dice must be rolled until two different outcomes are observed (or all dice have been rolled), and (3) a different generalization with $d$-sided dice, where dice must be rolled until all $d$ outcomes have been observed. For (1), we show that there is an optimal sequence which follows a simple greedy rule; the same rule only gives a 1-additive approximation for the original problem (arXiv:2510.16678 [cs.DS]). The rule also yields a correspondence between a particular optimal sequence and a related mechanical word, which we exploit to characterize the conditions under which this optimal sequence is periodic. We establish tight multiplicative and additive adaptivity gaps for this variant. For (2), we show that two different generalizations of the greedy rule from (arXiv:2510.16678 [cs.DS]) can be combined to obtain a PTAS. For (3), we give an $\mathcal{O}(\log d)$-approximation algorithm by reducing the problem to Submodular Ranking (arXiv:1007.2503 [cs.DS]); the same reduction technique can be used to yield approximation algorithms for other stochastic probing problems. Finally, we pose a number of related open questions.

Distance flexibility in spatial matching: the value of concentration

from arXiv: Data Structures and Algorithms

Authors: Taha Ameen, Sophie H. Yu

In spatial matching markets, a supply unit's flexibility is measured by its service radius, the maximum distance at which it can serve demand. In dimensions $k \geq 2$, we study how a platform should allocate service radii among the supply nodes subject to a budget on their sum. The platform makes this choice before observing supply and demand locations, with the objective of maximizing the expected fulfilled demand. We show that the shape of a preferred allocation depends on the total budget: under suitable conditions, large budgets favor allocations that are more uniform in the sense of majorization, while small budgets favor concentration. We also characterize a non-uniform allocation that is asymptotically optimal for a very-sparse regime, and show that the uniform allocation is suboptimal in this regime. Our results provide theoretical explanations for the radius allocation questions raised by the numerical experiments in [ASY26b].

Authors: Taha Ameen, Sophie H. Yu

In spatial matching markets, a supply unit's flexibility is measured by its service radius, the maximum distance at which it can serve demand. In dimensions $k \geq 2$, we study how a platform should allocate service radii among the supply nodes subject to a budget on their sum. The platform makes this choice before observing supply and demand locations, with the objective of maximizing the expected fulfilled demand. We show that the shape of a preferred allocation depends on the total budget: under suitable conditions, large budgets favor allocations that are more uniform in the sense of majorization, while small budgets favor concentration. We also characterize a non-uniform allocation that is asymptotically optimal for a very-sparse regime, and show that the uniform allocation is suboptimal in this regime. Our results provide theoretical explanations for the radius allocation questions raised by the numerical experiments in [ASY26b].

Arrival-Time Incentive Compatibility in Random Order Online Bipartite Matching

from arXiv: Data Structures and Algorithms

Authors: Arghya Chakraborty, Varun Gupta

In this work we initiate the study of competitive algorithms with arrival-time incentive compatibility for random-order online bipartite matching in settings where the users care only about receiving service (matched vs. unmatched) and not which offline resource serves them, while the platform's objective is to maximize total matching reward. This captures applications such as ride-sharing where the users primarily care about being matched to a ride while the platform internalizes the cost of dispatching a distant driver; dispatching homogeneous service requests to heterogeneous servers (cloud/edge routing); and assigning customer requests to a pool of providers with different flexibility (e.g., English-only vs. bilingual agents). Our main question is: \textit{Is constant-competitive matching possible for incentive-compatible, random-order edge-weighted matching on complete bipartite graphs?} Motivated by the LP-based treatment of incentive compatibility in the classical secretary problem by Buchbinder et al., we impose a constraint that the ex ante probability of selection is equalized across all arrival positions. We answer our main question in the affirmative and propose the first constant-competitive algorithm for incentive compatible edge-weighted random-order online matching on complete bipartite graphs. The competitive ratio of our algorithm is parameterized by the imbalance factor $k := n/m$ -- where $n$ and $m$ are the numbers of online and offline nodes, respectively, and $k$ is a positive integer. In particular, we obtain a competitive guarantee of the form $c_k - O(1/\sqrt{m})$ where $c_1 \approx 0.162$ and $c_k \to 0.02308\ldots$ as $k \to \infty$. We also present algorithms with strictly improved competitive ratio of $\approx 0.07 + O(1/m)$ for the binary-weighted case ($0$-$1$ rewards).

Authors: Arghya Chakraborty, Varun Gupta

In this work we initiate the study of competitive algorithms with arrival-time incentive compatibility for random-order online bipartite matching in settings where the users care only about receiving service (matched vs. unmatched) and not which offline resource serves them, while the platform's objective is to maximize total matching reward. This captures applications such as ride-sharing where the users primarily care about being matched to a ride while the platform internalizes the cost of dispatching a distant driver; dispatching homogeneous service requests to heterogeneous servers (cloud/edge routing); and assigning customer requests to a pool of providers with different flexibility (e.g., English-only vs. bilingual agents). Our main question is: \textit{Is constant-competitive matching possible for incentive-compatible, random-order edge-weighted matching on complete bipartite graphs?} Motivated by the LP-based treatment of incentive compatibility in the classical secretary problem by Buchbinder et al., we impose a constraint that the ex ante probability of selection is equalized across all arrival positions. We answer our main question in the affirmative and propose the first constant-competitive algorithm for incentive compatible edge-weighted random-order online matching on complete bipartite graphs. The competitive ratio of our algorithm is parameterized by the imbalance factor $k := n/m$ -- where $n$ and $m$ are the numbers of online and offline nodes, respectively, and $k$ is a positive integer. In particular, we obtain a competitive guarantee of the form $c_k - O(1/\sqrt{m})$ where $c_1 \approx 0.162$ and $c_k \to 0.02308\ldots$ as $k \to \infty$. We also present algorithms with strictly improved competitive ratio of $\approx 0.07 + O(1/m)$ for the binary-weighted case ($0$-$1$ rewards).

Maximizing Social Influence in Almost Linear Time

from arXiv: Data Structures and Algorithms

Authors: Saeed Seddighin

Influence maximization is a central algorithmic challenge in network analysis, aiming to identify a set of $k$ seed nodes in a graph with $n$ nodes and $m$ edges that maximizes the expected cascade of information under standard diffusion models. The seminal work of Borgs, Brautbar, Chayes, and Lucier (SODA'14) yielded a fundamental breakthrough\footnote{The conference version of their paper originally claimed a runtime of $\tilde O_ε(n+m)$, but this was subsequently corrected to a runtime of $\tilde O_ε((n+m)k)$ in an updated version of the paper that is available online. We validate the necessity of this additional factor $k$ in Section~\ref{sec:lowerbound} by demonstrating that if their algorithm is restricted to a runtime budget of $\tilde{O}_ε(n+m)$, the approximation ratio deteriorates to $O(k^{-1/4})$.} for this problem by achieving an $\tilde O_ε((n+m)k)$ time algorithm for approximating the solution within a factor of $1-1/e-ε$. In the years since, numerous efforts have attempted to improve the runtime of this algorithm; however, these works have been successful in only shaving logarithmic factors or improving the dependence on $ε$, leaving the existence of an almost linear-time algorithm as an open question. In this work, we resolve this long-standing open question. We present a novel algorithm that approximates the influence maximization problem within a factor of $1-1/e-ε$ in time $\tilde{O}_ε(n+m)$, effectively removing the multiplicative dependence on $k$ from the time complexity.

Authors: Saeed Seddighin

Influence maximization is a central algorithmic challenge in network analysis, aiming to identify a set of $k$ seed nodes in a graph with $n$ nodes and $m$ edges that maximizes the expected cascade of information under standard diffusion models. The seminal work of Borgs, Brautbar, Chayes, and Lucier (SODA'14) yielded a fundamental breakthrough\footnote{The conference version of their paper originally claimed a runtime of $\tilde O_ε(n+m)$, but this was subsequently corrected to a runtime of $\tilde O_ε((n+m)k)$ in an updated version of the paper that is available online. We validate the necessity of this additional factor $k$ in Section~\ref{sec:lowerbound} by demonstrating that if their algorithm is restricted to a runtime budget of $\tilde{O}_ε(n+m)$, the approximation ratio deteriorates to $O(k^{-1/4})$.} for this problem by achieving an $\tilde O_ε((n+m)k)$ time algorithm for approximating the solution within a factor of $1-1/e-ε$. In the years since, numerous efforts have attempted to improve the runtime of this algorithm; however, these works have been successful in only shaving logarithmic factors or improving the dependence on $ε$, leaving the existence of an almost linear-time algorithm as an open question. In this work, we resolve this long-standing open question. We present a novel algorithm that approximates the influence maximization problem within a factor of $1-1/e-ε$ in time $\tilde{O}_ε(n+m)$, effectively removing the multiplicative dependence on $k$ from the time complexity.

A Tight Second-Order Lower Bound for Routing Labels in Trees

from arXiv: Data Structures and Algorithms

Authors: Hanqing Li

In the designer-port routing-labeling problem, every vertex of a rooted tree receives a binary label and the child edges receive distinct port numbers. Given only the labels of a source and a destination, a decoder must return the first port on their path. Gawrychowski, Janczewski, and Lopuszanski gave labels of length $\log_2 n+O((\log_2\log_2 n)^2)$, whereas the previous lower bound was $\log_2 n+Ω(\log_2\log_2 n)$. We prove that every scheme for all $n$-vertex trees needs a label of length $\log_2 n+Ω((\log_2\log_2 n)^2)$ for every sufficiently large $n$. The result allows arbitrary port assignments and imposes no computational restriction on either the encoder or the decoder. Thus the second-order term in the optimal worst-case label length is determined up to constant factors.

Authors: Hanqing Li

In the designer-port routing-labeling problem, every vertex of a rooted tree receives a binary label and the child edges receive distinct port numbers. Given only the labels of a source and a destination, a decoder must return the first port on their path. Gawrychowski, Janczewski, and Lopuszanski gave labels of length $\log_2 n+O((\log_2\log_2 n)^2)$, whereas the previous lower bound was $\log_2 n+Ω(\log_2\log_2 n)$. We prove that every scheme for all $n$-vertex trees needs a label of length $\log_2 n+Ω((\log_2\log_2 n)^2)$ for every sufficiently large $n$. The result allows arbitrary port assignments and imposes no computational restriction on either the encoder or the decoder. Thus the second-order term in the optimal worst-case label length is determined up to constant factors.

Tuesday, September 29

Speeding Up the Process of Mourning

from Theory Dish: Stanford Blog

The world of mathematics, including theoretical computer science, is in turmoil. Even the Millennium Prize Problems, and, worse still, our beloved FOCS/STOC open problems, are no longer beyond the reach of LLMs. Watching this unfold inspires genuine awe and excitement, yet it also brings a real sense of loss. In the stages of mourning, the community seems to have moved away from denial. No more “models are nice, but they cannot do real math.” Instead, within our different mathematical communities, we now live in some combination of anger, bargaining, and depression. If you are a junior mathematician, you have every right to take your time processing this shift. You have my deep sympathies, and we must both support you and ensure you have a central voice in shaping the future of our field. But to senior colleagues, myself included, I say: Snap out of it. This moment is not simple, but there is no time to waste. We need to rise to the challenge. The world is changing at an incredible pace, and our response needs to be decisive and continuous. That may not be the traditional forte of academics, but the magnitude of this moment demands it. Among the [...]

The world of mathematics, including theoretical computer science, is in turmoil. Even the Millennium Prize Problems, and, worse still, our beloved FOCS/STOC open problems, are no longer beyond the reach of LLMs. Watching this unfold inspires genuine awe and excitement, yet it also brings a real sense of loss. In the stages of mourning, the community seems to have moved away from denial. No more “models are nice, but they cannot do real math.” Instead, within our different mathematical communities, we now live in some combination of anger, bargaining, and depression.

If you are a junior mathematician, you have every right to take your time processing this shift. You have my deep sympathies, and we must both support you and ensure you have a central voice in shaping the future of our field. But to senior colleagues, myself included, I say: Snap out of it.

This moment is not simple, but there is no time to waste. We need to rise to the challenge. The world is changing at an incredible pace, and our response needs to be decisive and continuous. That may not be the traditional forte of academics, but the magnitude of this moment demands it.

Among the reactions exhibited by senior mathematicians, I find bargaining and depression particularly harmful. Bargaining, a close cousin of denial, is the hope that our work can stay more or less the same with just a little adjustment. If only we could get the frontier labs to pause or stop proving our theorems, or if we slightly adjusted the rules of our publication venues, things wouldn’t be too bad. Sure, pushing back on frontier labs and addressing urgent concerns about the viability of our publication system are important. But we should not mistake these measures for a way to avoid a fundamental transformation of our profession.

Bargaining slows real action. It also prevents us from enjoying the positive aspects of the AI revolution, including progress on mathematical questions that we genuinely care about. We cannot suddenly move the goalposts and pretend that proving theorems was never the point, or that our open problems were merely proxies for building mathematical understanding. Those theorems are still of deep interest, and studying their proofs remains central to how we gain understanding in the first place. As for me, there are quite a few conjectures whose proofs I would absolutely love to understand, regardless of the source.

As for depression: the next time you have the urge to lament, or even celebrate, being “the last generation of human mathematicians,” perhaps keep it to yourself. Contemplating the end of your profession from the relative comfort of an established, tenured career is a privilege, and it comes with responsibilities. Senior academics are not merely individual researchers; we are stewards of our field, and we owe our junior colleagues active leadership rather than abandonment.

So, what do we need to do, and keep doing again and again?

Right now, the immediate, practical questions of how to adapt our institutions are getting the most attention, and we are already seeing thoughtful suggestions and encouraging initial steps. Today, this means increasing the recognition and incentives we provide for communication, understanding, and community building, and reflecting those priorities in our hiring, promotion, funding, and publication practices. For example, many are pointing out that we should no longer accept badly written papers just because we value the theorems. Similarly, we may now value conceptual work, such as new definitions, novel questions, and fresh techniques, more than ever before.

Yet we cannot treat these reforms as a one-time adjustment. We face the daunting task of continuously recreating our institutions as capabilities evolve. Tomorrow, models may surpass us at communicating their results, and eventually at conceptual work as well, which will force us to shift our core operations yet again. The same applies to how we educate future generations of mathematicians. If the human role increasingly centers on judgment, taste, and a broad perspective, how can newcomers reach that point? All of these questions are on everyone’s mind, and I urge us to be brave enough to pursue dramatic, ongoing transformations.

To guide those transformations, however, we need something deeper. Above all, we need to reevaluate our identity. What is it that makes mathematical knowledge and research valuable? What are we offering society, and how much of it survives in a world where models match or exceed humans in some or all relevant mathematical skills? For quite some time, the implicit social contract has been that society pays us to exercise our intellectual curiosity and we, in return, provide useful skills to the next generations and practical knowledge for the world. The current crisis is driven not only by the power of LLMs, but also by how rarely we have had to examine or articulate this contract. Now that the deal needs to be renegotiated, we should approach it with humility rather than entitlement.

It is easy to feel bleak when confronting these questions, so it helps to ask what a positive vision might look like, even in a future with artificial superintelligence (ASI). We are not there yet, as today’s models still make mistakes and flawed proofs can actively harm learning, but suppose we reach a world where models are much more capable. One optimistic possibility I have been toying with is a future that opens the best parts of the academic experience to everyone. Not that everyone would hold an academic job or create knowledge that is new to the world, but everyone could participate in serious intellectual exploration, in mathematics and beyond. Conversations with reliable models could be truly Socratic: helping us ask questions, develop ideas, and discover things for ourselves, rather than simply supplying answers. In this vision, everyone would have access to forms of intellectual creativity that are now reserved for a fortunate few. Within this world, professional academics (in mathematics and elsewhere) would need to find our own distinct role, and I believe we could. Even if this future supports fewer professional mathematicians, it could nevertheless support a much richer mathematical life.

Finally, we must look further than just our own small piece of heaven. Accelerating mathematics can bring tremendous good to the world if it speeds up applied fields, medicine, and other concrete benefits for society. At the same time, the disruptions and dangers extend far beyond academia: a professional driver losing their job is no less important than a professional mathematician whose work has become less enjoyable. We therefore have a duty to take our professional responsibility toward AI alignment seriously.

AI models are mathematical objects, and their development could not have happened without our collective work. Furthermore, mathematicians, especially theoretical computer scientists, have a critical role to play in helping to govern AI models so that they serve individuals and society rather than harm them. Of course, AI alignment is not merely a mathematical problem, but the mathematical perspective is invaluable. Some of us have long been calling for more significant involvement in navigating the interface between computation and society. It is time for many more to heed that call.

Acknowledgments: Thank you to Sílvia Casacuberta, Lee Cohen, Jabari Hastings, and Charlotte Peale for many meaningful conversations and thoughtful comments on earlier drafts, though the views expressed here are entirely my own. I also want to thank a couple of unnamed models that graciously helped me clarify my perspective.

By Omer Reingold

Pricing Commensurability

from Ben Recht

On the origins of cost-benefit analyses in governmental decision making

Hi there, argmin readers! Today’s post is a live blog of Class 8 of my graduate seminar “Forecasting: A Critical Retrospective.” The syllabus and list of past posts are here.

A bizarre central tenet of “rational decision-making” is that all optimal decisions can be made by computing an appropriate cost-benefit analysis. I riff on this in the introduction to The Irrational Decision, and always lead with more absurd examples when I talk about the book. Should you have a surgery? Should you force your kid to take violin lessons? Should you go for it on 4th down? According to the tenets of rational choice theory, you can answer all of these questions by forecasting a dollar value and probability of every outcome.

Thanks for reading arg min! Subscribe for free to receive new posts and support my work.

With these facts in hand, optimal decision making is merely a mechanical chain of sums and multiplications. This is ludicrous if you think about it for two seconds. And yet it’s become a standard social convention that utilitarian calculation is not only possible but the optimal way to live your life, run a business, or govern a nation.

In today’s class, we try to get at the roots of where this came from and how it became institutionalized. My two favorite references on the history are Theodore Porter’s Trust in Numbers and Elizabeth Popp Berman’s Thinking Like an Economist, both of which trace the history in the United States.

Porter starts before the war, looking at how cost-benefit analyses were formalized to justify water projects by the Army Corps of Engineers. He has a nice short article summarizing the book’s in-depth study. Water projects were crucial for preventing flood damage, routing water to farms, and making waterways more navigable. However, they were also classic pork-barrel projects, where elected officials would funnel money back to their districts. The Corps looked for means to “remove the politics” and demonstrate that each project was worth doing. They settled on cost-benefit analysis, establishing a rigorous system to enumerate all of the potential upsides and itemize all of the potential costs.

These calculations were eventually mandated in the 1936 Flood Control Act:

“...the Federal Government should improve or participate in the improvement of navigable waters or their tributaries, including watersheds thereof, for flood-control purposes if the benefits to whomsoever they may accrue are in excess of the estimated costs, and if the lives and social security of people are otherwise adversely affected.” (italics mine)

Cost-benefit analyses would leave them with a simple, clean, unitary number — the ratio between these costs and benefits — that they could present for project approval. All of the complexity could be reduced to two digits. These digits sufficed to make governance decisions. Significant expertise was needed to ensure these calculations held up to adversarial scrutiny. As Porter writes:

“Objectivity, then, meant above all the standardization of quantitative methods and the training up of people capable of performing them. Every failure of clarity, every gap in the reasoning, every loophole that left space for the quantifier to alter the results in a preferred direction, was a potential weakness, which opponents of the agency were certain to exploit, often in hearings before judges and administrators who would probably be ignorant of the fine points of economic quantification.”

Interestingly, no economists were consulted in constructing the estimates. The engineers prided themselves on their ruthless objectivity and ability to decouple their preferences from the cold hard facts. Moreover, the public preferred cost-benefit analyses to opaque expert judgment. Standard, transparent processes feel like they rule out arbitrariness and capriciousness of bureaucrats. Porter casts cost-benefit analysis as “a quantitative decision technology, practiced mainly in public bureaucracies, often in a highly politically-charged context.”

Popp Berman details how this technology spread through the government, with the establishment of various executive-branch offices staffed by experts to oversee complex problems like healthcare and education. It became institutionalized in policy schools, founded in the 1970s to provide graduates to staff said agencies.

Fast forward to the present, and we just take these cost-benefit analyses for granted. They give an institutionalized illusion of objectivity, but of course all of the calculations are subject to institutionalized norms of expert judgment. These norms tell you where you can commit rounding errors, ignore missing data, or disregard the unenumerable. But these are just institutional norms, and they don’t really hold up to scrutiny. As Larry Lohman details, the “objective” methods of institutionalized cost-benefit analysis are riddled with value-laden assumptions, and objectivity rests on absurd ideas of commensurability and the ability to price all preferences.

Moreover, Charles Manski describes the incredible uncertainty inherent to cost-benefit calculations.1 Manski notes that experts all know these uncertainties are present but choose not to report them for political reasons. You’ll often find cost-benefit analyses reported to three or four digits of precision, creating a further illusion of precision. Manski has a long list of critiques:2

  • Conventional certitude: A prediction that is generally accepted as true but is not necessarily true.

  • Dueling certitudes: Contradictory predictions made with alternative assumptions.

  • Conflating science and advocacy: Specifying assumptions to generate a predetermined conclusion.

  • Wishful extrapolation: Using untenable assumptions to extrapolate.

  • Illogical certitude: Drawing an unfounded conclusion based on logical errors.

  • Media overreach: Premature or exaggerated public reporting of policy analysis.

Together, these conventions conspire to communicate certainty where there is none. They justify decisions as rational by sweeping all of the uncertainty under the rug.

Subscribe now

1

We’ll cover uncertainty quantification in later classes.

2

My impression from economist friends is that Manski has a longer list of critiques than what appears in his published works, but he is too prideful to go full Nicholas Polson and have AI air all of his grievances.

By Ben Recht

Postdoctoral Associate at West Virginia University (apply by December 15, 2026)

from CCI: jobs

WVU’s Lane Department invites applications for a 2-year Postdoctoral Fellow in theoretical computer science starting Jan 1, 2027. Funded by NSF (Algorithmic Foundations), research focuses on algorithm design and computational complexity in mathematical programming. Duties include combinatorial optimization research and teaching one course. A PhD in CS or operations research is required. Website: wvu.taleo.net/careersection/faculty/jobdetail.ftl?job=30378&tz=GMT-04%3A00&tzname=America%2FNew_York Email: […]

WVU’s Lane Department invites applications for a 2-year Postdoctoral Fellow in theoretical computer science starting Jan 1, 2027. Funded by NSF (Algorithmic Foundations), research focuses on algorithm design and computational complexity in mathematical programming. Duties include combinatorial optimization research and teaching one course. A PhD in CS or operations research is required.

Website: https://wvu.taleo.net/careersection/faculty/jobdetail.ftl?job=30378&tz=GMT-04%3A00&tzname=America%2FNew_York
Email: k.subramani@mail.wvu.edu

By shacharlovett

Postdoc at Cambridge (apply by December 1, 2026)

from CCI: jobs

Postdoc position available in Tom Gur’s group at Cambridge on topics including (but not limited to) Classical and/or Quantum aspects of: Complexity, Sublinear Algorithms, Coding Theory, Cryptography, Learning Theory, and connections to Harmonic Analysis & Additive Combinatorics. Website: www.cam.ac.uk/jobs/research-assistantassociate-in-theoretical-computer-science-fixed-term-nr51230 Email: tg508@cam.ac.uk

Postdoc position available in Tom Gur’s group at Cambridge on topics including (but not limited to) Classical and/or Quantum aspects of: Complexity, Sublinear Algorithms, Coding Theory, Cryptography, Learning Theory, and connections to Harmonic Analysis & Additive Combinatorics.

Website: https://www.cam.ac.uk/jobs/research-assistantassociate-in-theoretical-computer-science-fixed-term-nr51230
Email: tg508@cam.ac.uk

By shacharlovett

Jeff Fest: Probabilistic Combinatorics at Rutgers

from Gil Kalai

Greetings from Providence! Next week there will be a conference at Rutgers University in honor of Jeff Kahn. Jeff is a great mathematician whose contributions span all areas of combinatorics. He has also been my friend for four decades and … Continue reading →

Greetings from Providence!

Next week there will be a conference at Rutgers University in honor of Jeff Kahn. Jeff is a great mathematician whose contributions span all areas of combinatorics. He has also been my friend for four decades and is my closest collaborator. Here is the conference description:

Probabilistic combinatorics lies at the heart of modern discrete mathematics, with deep connections to probability, theoretical computer science, and statistical physics. This conference will bring together leading researchers and emerging scholars to highlight recent breakthroughs, explore fundamental open problems, and recognize the profound influence of Jeff Kahn on the field.

There is a wonderful lineup of speakers, and the conference promises to be a great event. I look forward to similar events in extremal combinatorics, geometric combinatorics, matroid theory, posets, fractional combinatorics, finite geometries, and more, celebrating other aspects of Jeff’s work 🙂 . In the meeting, I will talk about problems around Borsuk’s conjecture.

I was also kindly invited to visit Brown University and speak at its applied mathematics colloquium. I am very excited to talk here about my work on quantum computers and to meet many friends and colleagues. These two events, Jefffest and the lecture at Brown, are the anchors of a rather intensive, ambitious, and nostalgic tour of Providence, Boston, New Haven, New Brunswick, and Princeton.

Jeff and me, 2006

By Gil Kalai

My “Knowmads” podcast on science and AI

from Scott Aaronson

Or click here if the above doesn’t work. Recorded in-person in my office at UT Austin, with a bulleted list containing “ARC,” “Scalable Oversight,” and “Models” behind me on my blackboard for some reason (I no longer remember who put those there or why). 90 minutes long. Sometimes you see my disembodied arm waving in […]

Or click here if the above doesn’t work.

Recorded in-person in my office at UT Austin, with a bulleted list containing “ARC,” “Scalable Oversight,” and “Models” behind me on my blackboard for some reason (I no longer remember who put those there or why). 90 minutes long. Sometimes you see my disembodied arm waving in midair because of the way the cameras are combined. As always, I strongly recommend 2x speed for the correct experience.

This might actually be one of my best podcasts ever, although I wasn’t planning on that! Thanks so much to Bhavay Tyagi and Prachi Garella for driving all the way from Houston to record it.

Here’s a strict subset of the topics we covered:

  • The story of AI models solving the Navier-Stokes Millennium Problem, insofar as it’s known
  • Can recent AI proofs be called “truly creative”?
  • The history of AI before the LLM revolution
  • What do we mean when we call LLMs “black boxes”?
  • The achievements of the field of interpretability
  • What exactly happened in the OpenAI/HuggingFace incident
  • Must we avoid all “anthropomorphizing language” when discussing the HuggingFace incident? (spoiler alert: no)
  • Examples of major open problems in quantum computing theory that I cared about for decades and that AI models have recently solved
  • Effects of the current AI cataclysm on the math community, especially students
  • What annoys me the most when I listen to AI talks
  • My experiences at OpenAI, why they hired me, and the watermarking work that I did there

Enjoy!

More AI-related content coming soon, as this blog—like much of the rest of the world—continues its transition to “all AI, all the time” (except still 100% written by an aging, deteriorating biological brain)

And for those who just can’t get enough of my rocking back and forth, using too many filler words, as I explain theoretical computer science! Here’s a second podcast, this one mainly on quantum computing, with Seb Agertoft, who I thank for doing it. Enjoy!

By Scott

Quantum Query Complexity Beyond the Worst Case

from arXiv: Computational Complexity

Authors: Srinivasan Arunachalam, Yanlin Chen, Amin Shiraz Gilani

Smoothed analysis is a central framework in classical algorithms for explaining the performance of algorithms beyond the worst case, often explaining why algorithms perform well in practice. We initiate a systematic study of its quantum counterpart and show the following results. $(1)$ We show that there is a total function whose smoothed quantum query complexity is exponentially smaller than its classical query complexity. $(2)$ We give near-tight characterizations of smoothed randomized and quantum query complexities for symmetric Boolean functions, unifying the worst-case complexity results of [Beals et al, FOCS'98] and average-case complexity results of [Ambainis and de Wolf, STACS'00]. $(3)$ We study string problems such as pattern matching and edit distance and, in various regimes, give polynomial to superpolynomial quantum speedups. Our main technical ingredients include a near-tight quantum algorithm for $\varepsilon$-approximating the number of collisions between two non-repetitive strings, improving the result of Le Gall and Ng [QIC'22]. Together, our results show that smoothing can reveal larger quantum speedups than worst-case analysis suggests, opening a path towards quantum advantage on more realistic inputs.

Authors: Srinivasan Arunachalam, Yanlin Chen, Amin Shiraz Gilani

Smoothed analysis is a central framework in classical algorithms for explaining the performance of algorithms beyond the worst case, often explaining why algorithms perform well in practice. We initiate a systematic study of its quantum counterpart and show the following results. $(1)$ We show that there is a total function whose smoothed quantum query complexity is exponentially smaller than its classical query complexity. $(2)$ We give near-tight characterizations of smoothed randomized and quantum query complexities for symmetric Boolean functions, unifying the worst-case complexity results of [Beals et al, FOCS'98] and average-case complexity results of [Ambainis and de Wolf, STACS'00]. $(3)$ We study string problems such as pattern matching and edit distance and, in various regimes, give polynomial to superpolynomial quantum speedups. Our main technical ingredients include a near-tight quantum algorithm for $\varepsilon$-approximating the number of collisions between two non-repetitive strings, improving the result of Le Gall and Ng [QIC'22]. Together, our results show that smoothing can reveal larger quantum speedups than worst-case analysis suggests, opening a path towards quantum advantage on more realistic inputs.

Sublinear Copies Suffice for Fidelity Estimation with Pauli Measurements

from arXiv: Computational Complexity

Authors: Jayadev Acharya, Abhilash Dharmavarapu, Yuhan Liu, Nengkun Yu

We present a protocol that estimates the quantum fidelity, up to precision $\varepsilon$, between a known target state and unknown lab-prepared state with sublinear, $o(d^{0.9908}/\varepsilon^2)$, number of Pauli basis measurements.

Authors: Jayadev Acharya, Abhilash Dharmavarapu, Yuhan Liu, Nengkun Yu

We present a protocol that estimates the quantum fidelity, up to precision $\varepsilon$, between a known target state and unknown lab-prepared state with sublinear, $o(d^{0.9908}/\varepsilon^2)$, number of Pauli basis measurements.

Distributional Variants of the Aaronson-Ambainis Conjecture

from arXiv: Computational Complexity

Authors: Uma Girish, Kunal Mittal, Barak Nehoran, Ran Raz

A longstanding conjecture in quantum complexity theory asserts that, under the uniform input distribution, quantum query algorithms can be polynomially simulated by classical query algorithms. More precisely, the acceptance probability of any quantum query algorithm can be approximated, on average over uniformly random inputs, by a classical query algorithm, with only polynomial query overhead. The conjecture is central to understanding whether exponential quantum advantages for decision problems necessarily rely on additional structure. We study analogues of this conjecture under other natural input distributions and prove that they are all equivalent to the original uniform-distribution conjecture. We first consider the product distribution $μ_p$, where the input bits are independent Bernoulli variables with fixed bias $p$. We show that for every fixed $p \in (0, 1)$, quantum query algorithms under the $μ_p$ distribution admit polynomial-overhead classical simulations if and only if the same holds under the uniform distribution. Second, we consider the distribution $ν_p$ that is uniform over the slice of strings with Hamming weight $\lfloor pn \rfloor$ and prove a similar equivalence for the $ν_p$ distribution and the uniform distribution. The Aaronson-Ambainis conjecture is a stronger statement that implies the above-mentioned conjecture and is formulated in terms of bounded low-degree polynomials on the Boolean hypercube. It asserts that under the uniform distribution, any such polynomial with nonnegligible variance must have an influential variable. We formulate analogues of this conjecture, where the underlying distribution is a biased product distribution or a uniform distribution over a slice, and prove that all these variants are equivalent to the original Aaronson-Ambainis conjecture.

Authors: Uma Girish, Kunal Mittal, Barak Nehoran, Ran Raz

A longstanding conjecture in quantum complexity theory asserts that, under the uniform input distribution, quantum query algorithms can be polynomially simulated by classical query algorithms. More precisely, the acceptance probability of any quantum query algorithm can be approximated, on average over uniformly random inputs, by a classical query algorithm, with only polynomial query overhead. The conjecture is central to understanding whether exponential quantum advantages for decision problems necessarily rely on additional structure. We study analogues of this conjecture under other natural input distributions and prove that they are all equivalent to the original uniform-distribution conjecture. We first consider the product distribution $μ_p$, where the input bits are independent Bernoulli variables with fixed bias $p$. We show that for every fixed $p \in (0, 1)$, quantum query algorithms under the $μ_p$ distribution admit polynomial-overhead classical simulations if and only if the same holds under the uniform distribution. Second, we consider the distribution $ν_p$ that is uniform over the slice of strings with Hamming weight $\lfloor pn \rfloor$ and prove a similar equivalence for the $ν_p$ distribution and the uniform distribution. The Aaronson-Ambainis conjecture is a stronger statement that implies the above-mentioned conjecture and is formulated in terms of bounded low-degree polynomials on the Boolean hypercube. It asserts that under the uniform distribution, any such polynomial with nonnegligible variance must have an influential variable. We formulate analogues of this conjecture, where the underlying distribution is a biased product distribution or a uniform distribution over a slice, and prove that all these variants are equivalent to the original Aaronson-Ambainis conjecture.

$\mathrm{Almost}\text{-}\oplus\mathrm{P} = \mathrm{BP}\cdot\oplus\mathrm{P}$ and a Random-Oracle Proof of Toda's Theorem

from arXiv: Computational Complexity

Authors: Lance Fortnow

Using the recent exponential correlation bounds of Chattopadhyay, Hatami, Lee, Lovett, Tal and Viola between $\mathbb{F}_2$-polynomials and the XOR of majorities, we show that $\mathrm{Almost}\text{-}\oplus\mathrm{P} = \mathrm{BP}\cdot\oplus\mathrm{P}$, where $\mathrm{Almost}\text{-}\oplus\mathrm{P}$ is the class of languages that lie in $\oplus\mathrm{P}^R$ with probability one for a random oracle $R$. This is the parity analogue of Bennett and Gill's $\mathrm{Almost}\text{-}\mathrm{P} = \mathrm{BPP}$ and Nisan and Wigderson's $\mathrm{Almost}\text{-}\mathrm{PH} = \mathrm{PH}$. The key ingredient is a pseudorandom generator with polynomial seed length that fools $\mathbb{F}_2$-polynomials of polynomial degree on exponentially many variables. As an application we complete a random-oracle proof of the first half of Toda's theorem, $\mathrm{PH} \subseteq \mathrm{BP}\cdot\oplus\mathrm{P}$, following an approach of Regan and Royer. Relative to a random oracle, the polynomial hierarchy collapses into $\oplus\mathrm{P}$ by applying Valiant-Vazirani and Papadimitriou-Zachos level by level, with no probabilistic quantifier ever moved through an oracle. Our result then removes the oracle. We compare this argument with the simple proof of Toda's theorem by Fortnow (2009).

Authors: Lance Fortnow

Using the recent exponential correlation bounds of Chattopadhyay, Hatami, Lee, Lovett, Tal and Viola between $\mathbb{F}_2$-polynomials and the XOR of majorities, we show that $\mathrm{Almost}\text{-}\oplus\mathrm{P} = \mathrm{BP}\cdot\oplus\mathrm{P}$, where $\mathrm{Almost}\text{-}\oplus\mathrm{P}$ is the class of languages that lie in $\oplus\mathrm{P}^R$ with probability one for a random oracle $R$. This is the parity analogue of Bennett and Gill's $\mathrm{Almost}\text{-}\mathrm{P} = \mathrm{BPP}$ and Nisan and Wigderson's $\mathrm{Almost}\text{-}\mathrm{PH} = \mathrm{PH}$. The key ingredient is a pseudorandom generator with polynomial seed length that fools $\mathbb{F}_2$-polynomials of polynomial degree on exponentially many variables. As an application we complete a random-oracle proof of the first half of Toda's theorem, $\mathrm{PH} \subseteq \mathrm{BP}\cdot\oplus\mathrm{P}$, following an approach of Regan and Royer. Relative to a random oracle, the polynomial hierarchy collapses into $\oplus\mathrm{P}$ by applying Valiant-Vazirani and Papadimitriou-Zachos level by level, with no probabilistic quantifier ever moved through an oracle. Our result then removes the oracle. We compare this argument with the simple proof of Toda's theorem by Fortnow (2009).

Computational Complexity of Clifford Template Compilation: Are Quantum Computers Useful for Compiling Quantum Circuits?

from arXiv: Computational Complexity

Authors: Keisuke Fujii

A Clifford template is a finite ordered family of repeatable Clifford operations, and an instantiation specifies how many times each operation is applied. The Clifford template compilation problem asks how to choose these repetition numbers so that the template realizes a target transformation of Pauli operators. This problem arises, for example, when searching for logical operations in quantum error correction using only Clifford operations permitted by physical or fault-tolerance constraints. Although forward Clifford dynamics is efficiently classically simulable, this inverse problem has sharp complexity transitions. For commuting templates with unrestricted integer exponents, feasibility lies in $\mathrm{NP}\cap\mathrm{BQP}$ and a constructive quantum algorithm returns a particular solution together with the full exponent-relation lattice; already at $k=1$, recovering the repetition number contains finite-field discrete logarithm over $\mathbb{F}_{2^r}^{\times}$. In general, restricting every exponent to $\{0,1\}$ removes the Abelian-group closure and makes feasibility NP-complete for variable $k$, even for exactly commuting CNOT-only operations and X-type Paulis. For commuting self-inverse Clifford actions, both binary feasibility and recovery of one solution are classically polynomial-time solvable, but imposing a bound on the total repetition count is NP-complete, even for CNOT-only operations. These results reveal a rich complexity landscape within Clifford template compilation, spanning classically tractable cases, problems admitting quantum polynomial-time algorithms, and NP-complete variants.

Authors: Keisuke Fujii

A Clifford template is a finite ordered family of repeatable Clifford operations, and an instantiation specifies how many times each operation is applied. The Clifford template compilation problem asks how to choose these repetition numbers so that the template realizes a target transformation of Pauli operators. This problem arises, for example, when searching for logical operations in quantum error correction using only Clifford operations permitted by physical or fault-tolerance constraints. Although forward Clifford dynamics is efficiently classically simulable, this inverse problem has sharp complexity transitions. For commuting templates with unrestricted integer exponents, feasibility lies in $\mathrm{NP}\cap\mathrm{BQP}$ and a constructive quantum algorithm returns a particular solution together with the full exponent-relation lattice; already at $k=1$, recovering the repetition number contains finite-field discrete logarithm over $\mathbb{F}_{2^r}^{\times}$. In general, restricting every exponent to $\{0,1\}$ removes the Abelian-group closure and makes feasibility NP-complete for variable $k$, even for exactly commuting CNOT-only operations and X-type Paulis. For commuting self-inverse Clifford actions, both binary feasibility and recovery of one solution are classically polynomial-time solvable, but imposing a bound on the total repetition count is NP-complete, even for CNOT-only operations. These results reveal a rich complexity landscape within Clifford template compilation, spanning classically tractable cases, problems admitting quantum polynomial-time algorithms, and NP-complete variants.

The complexity of computing the covering radius of a Euclidean lattice

from arXiv: Computational Complexity

Authors: Frank Vallentin

In this note, we prove that the covering radius problem for Euclidean lattices is complete for the second level of the polynomial hierarchy. The note also documents the author's first experiment with generative AI as a tool for mathematical research.

Authors: Frank Vallentin

In this note, we prove that the covering radius problem for Euclidean lattices is complete for the second level of the polynomial hierarchy. The note also documents the author's first experiment with generative AI as a tool for mathematical research.

Anticoncentration of Complex Gaussian Hafnians

from arXiv: Computational Complexity

Authors: Priyanshu Pant

Let $G_{2n}$ be a complex symmetric random matrix whose entries above the diagonal are independent standard circular complex Gaussians, and let $H_n=\operatorname{haf}(G_{2n})$. We prove the uniform shifted anticoncentration bound $$ \Pr\!\left( \left| \frac{H_n}{\sqrt{(2n-1)!!}}-z \right| \le \varepsilon \right) \le 2\sqrt{\frac nπ}\,\varepsilon^2 $$ for every $z\in\mathbb C$ and $\varepsilon>0$. This establishes a local anticoncentration property that supports hardness arguments for quantum advantage in Gaussian boson sampling.

Authors: Priyanshu Pant

Let $G_{2n}$ be a complex symmetric random matrix whose entries above the diagonal are independent standard circular complex Gaussians, and let $H_n=\operatorname{haf}(G_{2n})$. We prove the uniform shifted anticoncentration bound $$ \Pr\!\left( \left| \frac{H_n}{\sqrt{(2n-1)!!}}-z \right| \le \varepsilon \right) \le 2\sqrt{\frac nπ}\,\varepsilon^2 $$ for every $z\in\mathbb C$ and $\varepsilon>0$. This establishes a local anticoncentration property that supports hardness arguments for quantum advantage in Gaussian boson sampling.

Consequences of Polylogarithmic Membership Comparability for SAT

from arXiv: Computational Complexity

Authors: Sebastian Ben Daniel

We study the consequences of membership comparators that exclude one possible membership vector, deterministically or with a relative advantage over uniform guessing. For every polynomially bounded arity, a randomized polynomial-time comparator of error at most $(1-1/poly(n))2^{-t}$ gives $ NP/ poly\cap coNP/ poly$ recognition with common advice. The proof uses limited independence, polynomial occurrence certificates, and a self-contained positive-relation advice transfer. For SAT at arity $O((\log n)^d)$, both this relative-gap hypothesis and deterministic comparability imply $PH=S^{NP}$, the uniform bound $PH\subseteq BPTIME(2^{O((\log n)^{d^2})})$, and symmetric verification with polynomial-length certificates and an oracle-free deterministic $2^{O((\log n)^d)}$ predicate. Polynomial-advice deterministic decoding has the same exponent $d$. Applying the randomized simulation to an unconditional diagonal language yields, for every fixed $\varepsilon>0$, $\mathrm{BPP}\subsetneq BPTIME(2^{O((\log n)^{d^2+\varepsilon})})$, without advice. The larger clock remains subexponential under every fixed number of self-compositions. A layered oracle satisfies deterministic comparability and $NP^O=coNP^O$ but excludes randomized NP algorithms with smaller logarithmic power, establishing a relativized limit on the SAT exponent $d$. This expanded version also develops the full weak-advantage regime, where saving $2^{-O((log n)^d)}$ gives randomized SAT exponent $d$ and PH exponent $d^k$ at fixed level $k$; the quasipolynomial and exponential hierarchy consequences; binary-comparator advice bounds; and the certificate-length boundary between the randomized regimes. Under deterministic comparability, uniform deterministic promise-unique search additionally gives $UEXP=EXP$. The ordinary second-level collapse $PH=Σ_2^p$ for $d>1$ remains unproved.

Authors: Sebastian Ben Daniel

We study the consequences of membership comparators that exclude one possible membership vector, deterministically or with a relative advantage over uniform guessing. For every polynomially bounded arity, a randomized polynomial-time comparator of error at most $(1-1/poly(n))2^{-t}$ gives $ NP/ poly\cap coNP/ poly$ recognition with common advice. The proof uses limited independence, polynomial occurrence certificates, and a self-contained positive-relation advice transfer. For SAT at arity $O((\log n)^d)$, both this relative-gap hypothesis and deterministic comparability imply $PH=S^{NP}$, the uniform bound $PH\subseteq BPTIME(2^{O((\log n)^{d^2})})$, and symmetric verification with polynomial-length certificates and an oracle-free deterministic $2^{O((\log n)^d)}$ predicate. Polynomial-advice deterministic decoding has the same exponent $d$. Applying the randomized simulation to an unconditional diagonal language yields, for every fixed $\varepsilon>0$, $\mathrm{BPP}\subsetneq BPTIME(2^{O((\log n)^{d^2+\varepsilon})})$, without advice. The larger clock remains subexponential under every fixed number of self-compositions. A layered oracle satisfies deterministic comparability and $NP^O=coNP^O$ but excludes randomized NP algorithms with smaller logarithmic power, establishing a relativized limit on the SAT exponent $d$. This expanded version also develops the full weak-advantage regime, where saving $2^{-O((log n)^d)}$ gives randomized SAT exponent $d$ and PH exponent $d^k$ at fixed level $k$; the quasipolynomial and exponential hierarchy consequences; binary-comparator advice bounds; and the certificate-length boundary between the randomized regimes. Under deterministic comparability, uniform deterministic promise-unique search additionally gives $UEXP=EXP$. The ordinary second-level collapse $PH=Σ_2^p$ for $d>1$ remains unproved.

Depth-Optimal Quantum Compilation

from arXiv: Computational Complexity

Authors: Francisca Vasconcelos

We achieve the first constant-depth circuit for arbitrary single-qubit gate synthesis. Unlike prior approaches, the construction is fully unitary and requires no pre-supplied catalyst. For any constant $δ>0$, it $\varepsilon$-approximates an arbitrary single-qubit gate using $O(\log^{1+δ}(1/\varepsilon))$ clean ancillae, Hadamard and $T$ single-qubit gates, $O(\log(1/\varepsilon))$-width generalized Toffoli gates, and sublogarithmic-width Fan-Out gates. We further eliminate Fan-Out entirely, showing that Hadamard, $T$, and generalized Toffoli gates alone suffice for constant-depth synthesis. When restricted to the standard bounded-width gate model, our construction has depth $O(\log\log(1/\varepsilon))$, and we prove a matching $Ω(\log\log(1/\varepsilon))$-depth lower bound. Overall, we establish that $Θ(\log\log(1/\varepsilon))$-depth is unavoidable with only bounded-width gates, yet allowing even logarithmic-width multi-qubit gates suffices to achieve constant-depth synthesis. These results also reveal new structure in shallow quantum circuit complexity. We give a depth-preserving real simulation of bounded-error decision computation, showing that every depth-$d$ QAC circuit can be simulated in depth $O(d)$ using only Hadamard, $X$, and generalized Toffoli gates. Thus arbitrary single-qubit rotations and complex amplitudes do not increase the bounded-error decision power of QAC, even at constant depth. In particular, this reduces the long-standing conjecture Parity$\notin$QAC$^0$ to proving a Parity lower bound against circuits consisting only of Hadamard, $X$, and generalized Toffoli gates. More generally, this real normal form exposes a direct correspondence between the standard shallow-depth quantum circuit hierarchy and a hierarchy of Forrelation circuits with restricted oracle families.

Authors: Francisca Vasconcelos

We achieve the first constant-depth circuit for arbitrary single-qubit gate synthesis. Unlike prior approaches, the construction is fully unitary and requires no pre-supplied catalyst. For any constant $δ>0$, it $\varepsilon$-approximates an arbitrary single-qubit gate using $O(\log^{1+δ}(1/\varepsilon))$ clean ancillae, Hadamard and $T$ single-qubit gates, $O(\log(1/\varepsilon))$-width generalized Toffoli gates, and sublogarithmic-width Fan-Out gates. We further eliminate Fan-Out entirely, showing that Hadamard, $T$, and generalized Toffoli gates alone suffice for constant-depth synthesis. When restricted to the standard bounded-width gate model, our construction has depth $O(\log\log(1/\varepsilon))$, and we prove a matching $Ω(\log\log(1/\varepsilon))$-depth lower bound. Overall, we establish that $Θ(\log\log(1/\varepsilon))$-depth is unavoidable with only bounded-width gates, yet allowing even logarithmic-width multi-qubit gates suffices to achieve constant-depth synthesis. These results also reveal new structure in shallow quantum circuit complexity. We give a depth-preserving real simulation of bounded-error decision computation, showing that every depth-$d$ QAC circuit can be simulated in depth $O(d)$ using only Hadamard, $X$, and generalized Toffoli gates. Thus arbitrary single-qubit rotations and complex amplitudes do not increase the bounded-error decision power of QAC, even at constant depth. In particular, this reduces the long-standing conjecture Parity$\notin$QAC$^0$ to proving a Parity lower bound against circuits consisting only of Hadamard, $X$, and generalized Toffoli gates. More generally, this real normal form exposes a direct correspondence between the standard shallow-depth quantum circuit hierarchy and a hierarchy of Forrelation circuits with restricted oracle families.

Nash Equilibria in Auctions with Pacing Strategies: Complexity and Inefficiency

from arXiv: Computational Complexity

Authors: Aris Filos-Ratsikas, Charalampos Kokkalis, Mohamad Latifian

We introduce and study Auctions with Pacing Strategies (APS) games, a full-information model in which utility-maximizing bidders compete across many simultaneous first-price auctions, each choosing a single pacing multiplier that uniformly scales their values into bids. We settle three central questions. First, we show that there are instances that admit no approximate pure Nash equilibria. Then, we prove that the problem of deciding whether an APS game admits an (approximate) equilibrium is NP-complete in general, but can be solved in polynomial time if either the number of bidders or the number of items is fixed. Finally, when an equilibrium does exist, we characterize its inefficiency exactly, showing that both the Price of Anarchy and the Price of Stability equal $\frac{e}{e-1}$.

Authors: Aris Filos-Ratsikas, Charalampos Kokkalis, Mohamad Latifian

We introduce and study Auctions with Pacing Strategies (APS) games, a full-information model in which utility-maximizing bidders compete across many simultaneous first-price auctions, each choosing a single pacing multiplier that uniformly scales their values into bids. We settle three central questions. First, we show that there are instances that admit no approximate pure Nash equilibria. Then, we prove that the problem of deciding whether an APS game admits an (approximate) equilibrium is NP-complete in general, but can be solved in polynomial time if either the number of bidders or the number of items is fixed. Finally, when an equilibrium does exist, we characterize its inefficiency exactly, showing that both the Price of Anarchy and the Price of Stability equal $\frac{e}{e-1}$.

A Quadratic Lower Bound on Determinantal Complexity

from arXiv: Computational Complexity

Authors: Mrinal Kumar, Ben Lee Volk

We prove an $Ω(n^2)$ lower bound on the determinantal complexity of the power sum polynomial $\sum_{i=1}^n x_i^n$ over the field of complex numbers. A similar result was claimed in a recent paper of Sheshadri (arXiv:2606.13628), via an AI-assisted and AI-written proof. Assuming its correctness, this was the first super-linear lower bound for this fundamental algebraic problem for any explicit polynomial. However, the authors of this note were unable to follow the details and verify the argument in arXiv:2606.13628, in spite of considerable effort on their part. The proof we provide here is short, (almost) self-contained and seemingly simpler.

Authors: Mrinal Kumar, Ben Lee Volk

We prove an $Ω(n^2)$ lower bound on the determinantal complexity of the power sum polynomial $\sum_{i=1}^n x_i^n$ over the field of complex numbers. A similar result was claimed in a recent paper of Sheshadri (arXiv:2606.13628), via an AI-assisted and AI-written proof. Assuming its correctness, this was the first super-linear lower bound for this fundamental algebraic problem for any explicit polynomial. However, the authors of this note were unable to follow the details and verify the argument in arXiv:2606.13628, in spite of considerable effort on their part. The proof we provide here is short, (almost) self-contained and seemingly simpler.

Hitting Sets for Polynomials with Small Partial Derivative Spaces

from arXiv: Computational Complexity

Authors: Shubham Bhardwaj, Ramprasad Saptharishi

We give an explicit hitting set of size $\text{poly}(n,d,r)$ for the class of $n$-variate degree-$d$ polynomials whose partial derivative space is bounded by $r$, over any field $\mathbb{F}$ of characteristic zero. In particular, this yields a polynomial sized hitting set for the class of depth-$3$ powering circuits. The main technical insight is the construction of a "formal derivation'' and properties of the associated Wronskian with respect to this derivation, which was previously studied by Moura [Moura_2004] in a very different context. The proofs in this paper are elementary and completely self-contained. AI disclosure: The proof of this result was obtained during conversations [astra_proof] with OpenAI GPT-6 Astra. The proof presented in this writeup is a rewriting (in the authors' words) of the proof obtained by the AI model in a form that we believe is understandable to researchers.

Authors: Shubham Bhardwaj, Ramprasad Saptharishi

We give an explicit hitting set of size $\text{poly}(n,d,r)$ for the class of $n$-variate degree-$d$ polynomials whose partial derivative space is bounded by $r$, over any field $\mathbb{F}$ of characteristic zero. In particular, this yields a polynomial sized hitting set for the class of depth-$3$ powering circuits. The main technical insight is the construction of a "formal derivation'' and properties of the associated Wronskian with respect to this derivation, which was previously studied by Moura [Moura_2004] in a very different context. The proofs in this paper are elementary and completely self-contained. AI disclosure: The proof of this result was obtained during conversations [astra_proof] with OpenAI GPT-6 Astra. The proof presented in this writeup is a rewriting (in the authors' words) of the proof obtained by the AI model in a form that we believe is understandable to researchers.

The Fully Depolarizing Noise Conjecture for Entangled Physical States: A Twenty-Year Perspective

from arXiv: Computational Complexity

Authors: Gil Kalai

In this paper I revisit my 2006 conjecture on correlated errors in entangled physical qubits, originally proposed as a potential obstruction to quantum fault tolerance. The conjecture asserts that, in any physical implementation of a quantum computer, the effective noise channel acting on entangled physical qubits contains a joint fully depolarizing component, with a rate comparable to that of two-qubit gate errors. This hypothesized structural constraint goes beyond standard noise models and, if valid, would pose a significant challenge to scalable quantum fault tolerance. The conjecture remains open, but recent advances in experimental quantum computing bring it within reach of empirical testing on current devices. I also discuss two related directions in my critical study of quantum computation: the role of noise sensitivity and computational complexity in noisy intermediate-scale quantum systems, and the statistical analysis of experimental claims of quantum advantage. Finally, since this paper is written for a volume honoring Yuri Gurevich, I include some reflections on the ways in which my scientific and personal trajectory became intertwined with Yuri's.

Authors: Gil Kalai

In this paper I revisit my 2006 conjecture on correlated errors in entangled physical qubits, originally proposed as a potential obstruction to quantum fault tolerance. The conjecture asserts that, in any physical implementation of a quantum computer, the effective noise channel acting on entangled physical qubits contains a joint fully depolarizing component, with a rate comparable to that of two-qubit gate errors. This hypothesized structural constraint goes beyond standard noise models and, if valid, would pose a significant challenge to scalable quantum fault tolerance. The conjecture remains open, but recent advances in experimental quantum computing bring it within reach of empirical testing on current devices. I also discuss two related directions in my critical study of quantum computation: the role of noise sensitivity and computational complexity in noisy intermediate-scale quantum systems, and the statistical analysis of experimental claims of quantum advantage. Finally, since this paper is written for a volume honoring Yuri Gurevich, I include some reflections on the ways in which my scientific and personal trajectory became intertwined with Yuri's.

A Dichotomy for Cubic Bipartite Holant Problems with Complex Algebraic Weights

from arXiv: Computational Complexity

Authors: Yin, Liu

We classify the exact evaluation of $\operatorname{Holant}(f\mid=_3)$ for every fixed complex algebraic symmetric Boolean ternary signature $f$. An input is a cubic bipartite multigraph: every vertex on one side carries $f$, every vertex on the other side carries ternary equality, and no auxiliary signatures are freely available. The tractable signatures are precisely rank-one tensors, generalized equalities, and equality-preserving cube-root diagonal transformations of six affine signatures, together with nonzero scalings and reversal. Every other signature gives a $\#\mathrm{P}$-hard problem under polynomial-time Turing reductions. We also identify the exact real intersection: it consists of the same tractable families as in the rational classification, with real algebraic parameters. The proof preserves degree exactly three on both sides of every oracle instance. A rank-one matrix extracted by interpolation supplies one unary signature only after its unused factor has been absorbed in triples. Over the complex numbers this absorption has three exceptional projective directions. We combine this constraint with projective matrix-group orbits, explicit ternary replacements, and an exhaustive treatment of finite projective orders. The cases of orders three and five include exact polynomial certificates; the certificate identities and a rational-arithmetic verifier are supplied as supplementary material.

Authors: Yin, Liu

We classify the exact evaluation of $\operatorname{Holant}(f\mid=_3)$ for every fixed complex algebraic symmetric Boolean ternary signature $f$. An input is a cubic bipartite multigraph: every vertex on one side carries $f$, every vertex on the other side carries ternary equality, and no auxiliary signatures are freely available. The tractable signatures are precisely rank-one tensors, generalized equalities, and equality-preserving cube-root diagonal transformations of six affine signatures, together with nonzero scalings and reversal. Every other signature gives a $\#\mathrm{P}$-hard problem under polynomial-time Turing reductions. We also identify the exact real intersection: it consists of the same tractable families as in the rational classification, with real algebraic parameters. The proof preserves degree exactly three on both sides of every oracle instance. A rank-one matrix extracted by interpolation supplies one unary signature only after its unused factor has been absorbed in triples. Over the complex numbers this absorption has three exceptional projective directions. We combine this constraint with projective matrix-group orbits, explicit ternary replacements, and an exhaustive treatment of finite projective orders. The cases of orders three and five include exact polynomial certificates; the certificate identities and a rational-arithmetic verifier are supplied as supplementary material.

The Complexity of Nash Equilibrium in Network Congestion and Coordination Games

from arXiv: Computational Complexity

Authors: Ioannis Anagnostides, Ioannis Panageas, Jingming Yan

We show that computing a Nash equilibrium is CLS-complete for linear network congestion and network coordination games. As a result, finding a KKT point of a bilinear polynomial is CLS-complete.

Authors: Ioannis Anagnostides, Ioannis Panageas, Jingming Yan

We show that computing a Nash equilibrium is CLS-complete for linear network congestion and network coordination games. As a result, finding a KKT point of a bilinear polynomial is CLS-complete.

Accepting-Path Counting at the One-Tape $n\log n$ Threshold

from arXiv: Computational Complexity

Authors: Ondřej Kuželka

We observe that the classical $n\log n$ time threshold for one-tape Turing machines is also a threshold for their accepting-path counts. Below it, every nondeterministic one-tape machine running in strong $o(n\log n)$ time has a rational ordinary generating function of accepting-path counts. At strong $O(n\log n)$ time, the situation changes completely: there is a fixed one-tape machine whose accepting-path function is complete for $\#\mathsf P_1$, the tally analogue of $\#\mathsf P$, under parsimonious polynomial-time tally reductions. A second construction within the same time bound gives positive accepting-path counts with a noncomputable exponential growth rate. The rationality result combines the one-tape time gap with the linear-time counting theorem of Tadaki, Yamakami and Lin. The completeness proof adapts the linear-time universal counting machine of Beame et al. to the one-tape setting.

Authors: Ondřej Kuželka

We observe that the classical $n\log n$ time threshold for one-tape Turing machines is also a threshold for their accepting-path counts. Below it, every nondeterministic one-tape machine running in strong $o(n\log n)$ time has a rational ordinary generating function of accepting-path counts. At strong $O(n\log n)$ time, the situation changes completely: there is a fixed one-tape machine whose accepting-path function is complete for $\#\mathsf P_1$, the tally analogue of $\#\mathsf P$, under parsimonious polynomial-time tally reductions. A second construction within the same time bound gives positive accepting-path counts with a noncomputable exponential growth rate. The rationality result combines the one-tape time gap with the linear-time counting theorem of Tadaki, Yamakami and Lin. The completeness proof adapts the linear-time universal counting machine of Beame et al. to the one-tape setting.

Algebraic-Geometric Parvaresh--Vardy Subspace Designs and Rank Condensers

from arXiv: Computational Complexity

Authors: Gil Cohen, Dean Doron, Noam Goldgraber

A subspace design is a collection of subspaces $H_1,\ldots,H_n$ of $\mathbb{F}_q^k$ with the property that no low-dimensional subspace $W$ intersects the collection "too much". Subspace designs and related objects in linear-algebraic pseudorandomness have found a broad range of applications, ranging from list decoding, to derandomizing algorithms. We construct explicit strong subspace designs over every finite field. In the extremal case where the co-dimension $t$ of each $H_i$ is equal to the dimension of $W$, for every constant field size our construction attains $n=Ω(k)$ and matches the probabilistic intersection bound up to a constant factor. All previous constructions required the field size to grow with $t$ (or $k$). Our subspace designs also imply new construction of rank condensers over arbitrary finite fields. This result is the first to achieve an optimal dependence on $k$ while maintaining both a constant output entropy rate and a constant field size. As an application, we construct lossless rank extractors for linear sources of rank $r$, for all $r < q$, with parameters matching those of Guo, Raj, Shangguan and Zhang (FOCS '26), thereby generalizing their result to prime fields and smaller field sizes. Our construction is based on an algebraic-geometric version of the Parvaresh-Vardy codes (Parvaresh-Vardy FOCS '05, Guruswami ECCC '05), extending the framework underlying the condensers of Guruswami, Umans and Vadhan (JACM '09). We view our construction as a linear-algebraic analysis - tailored to affine sources - of the GUV construction, generalized to functions over algebraic curves. More specifically, inspired by Ta-Shma and Umans (CCC 12') we develop a two-level evaluation scheme, where we first evaluate a function on a curve at extension-field points, and then evaluate a corresponding affine-linear polynomial to obtain outputs over the base field.

Authors: Gil Cohen, Dean Doron, Noam Goldgraber

A subspace design is a collection of subspaces $H_1,\ldots,H_n$ of $\mathbb{F}_q^k$ with the property that no low-dimensional subspace $W$ intersects the collection "too much". Subspace designs and related objects in linear-algebraic pseudorandomness have found a broad range of applications, ranging from list decoding, to derandomizing algorithms. We construct explicit strong subspace designs over every finite field. In the extremal case where the co-dimension $t$ of each $H_i$ is equal to the dimension of $W$, for every constant field size our construction attains $n=Ω(k)$ and matches the probabilistic intersection bound up to a constant factor. All previous constructions required the field size to grow with $t$ (or $k$). Our subspace designs also imply new construction of rank condensers over arbitrary finite fields. This result is the first to achieve an optimal dependence on $k$ while maintaining both a constant output entropy rate and a constant field size. As an application, we construct lossless rank extractors for linear sources of rank $r$, for all $r < q$, with parameters matching those of Guo, Raj, Shangguan and Zhang (FOCS '26), thereby generalizing their result to prime fields and smaller field sizes. Our construction is based on an algebraic-geometric version of the Parvaresh-Vardy codes (Parvaresh-Vardy FOCS '05, Guruswami ECCC '05), extending the framework underlying the condensers of Guruswami, Umans and Vadhan (JACM '09). We view our construction as a linear-algebraic analysis - tailored to affine sources - of the GUV construction, generalized to functions over algebraic curves. More specifically, inspired by Ta-Shma and Umans (CCC 12') we develop a two-level evaluation scheme, where we first evaluate a function on a curve at extension-field points, and then evaluate a corresponding affine-linear polynomial to obtain outputs over the base field.

Randomized Lifting for One-Way Number-on-Forehead Communication

from arXiv: Computational Complexity

Authors: Chenyu Wang

We prove a lifting theorem from two-party public-coin one-way communication to multiparty public-coin one-way number-on-forehead (NOF) communication. For every fixed $k\ge2$ and prime $q>2k$, there is a generalized inner product gadget $\GIP_{q,r}^k:(\F_q^r)^k\to\F_q$ with $r=O_k(q/\log q)$ such that, for every partial Boolean function $f:D\to\bits$, where $D\subseteq\F_q\times\F_q$, \[ R_{1/3}^1(f)-O(1) \le R_{1/6}^{1,\NOF}\bigl(f\circ\GIP_{q,r}^k\bigr) \le R_{1/6}^1(f). \] Thus, composition with the gadget preserves one-way randomized communication complexity up to an additive constant and a change in the error parameter. The lower bound holds in the general one-way NOF model, where the last player sees the entire gadget input. This extends the deterministic one-way NOF lifting theorem of Yang and Zhang to randomized protocols, and extends the randomized lifting result of Wang and Wu from the conservative model to the general one-way NOF model. Our proof introduces a one-way cylinder partition bound that lower bounds public-coin one-way NOF communication complexity. We show that, for the lifted function, this bound is at least half the one-way partition bound of the outer function. The main technical step transfers a dual solution between the two bounds, using Möbius inversion and a discrepancy estimate for generalized inner product to control the loss. Combining this transfer with the characterization of two-party one-way randomized communication complexity by the one-way partition bound yields the lifting theorem.

Authors: Chenyu Wang

We prove a lifting theorem from two-party public-coin one-way communication to multiparty public-coin one-way number-on-forehead (NOF) communication. For every fixed $k\ge2$ and prime $q>2k$, there is a generalized inner product gadget $\GIP_{q,r}^k:(\F_q^r)^k\to\F_q$ with $r=O_k(q/\log q)$ such that, for every partial Boolean function $f:D\to\bits$, where $D\subseteq\F_q\times\F_q$, \[ R_{1/3}^1(f)-O(1) \le R_{1/6}^{1,\NOF}\bigl(f\circ\GIP_{q,r}^k\bigr) \le R_{1/6}^1(f). \] Thus, composition with the gadget preserves one-way randomized communication complexity up to an additive constant and a change in the error parameter. The lower bound holds in the general one-way NOF model, where the last player sees the entire gadget input. This extends the deterministic one-way NOF lifting theorem of Yang and Zhang to randomized protocols, and extends the randomized lifting result of Wang and Wu from the conservative model to the general one-way NOF model. Our proof introduces a one-way cylinder partition bound that lower bounds public-coin one-way NOF communication complexity. We show that, for the lifted function, this bound is at least half the one-way partition bound of the outer function. The main technical step transfers a dual solution between the two bounds, using Möbius inversion and a discrepancy estimate for generalized inner product to control the loss. Combining this transfer with the characterization of two-party one-way randomized communication complexity by the one-way partition bound yields the lifting theorem.

Interactive Proofs of Proximity for Model Evaluation

from arXiv: Computational Complexity

Authors: Geoffroy Couteau, Nikolas Melissaris, Tamara Paris

We study interactive proofs of proximity (IPPs) for model evaluation, where a resource-limited verifier interacts with an untrusted prover, typically the model owner, to certify statistical properties of a model under an unknown input distribution. Our formulation separates sampling the input distribution from querying the model and evaluating its output; distinguishes real audit data (black-box sampling) from generated data (chosen-randomness, or gray-box, access to the sampler); and allows the prover and verifier to use different evaluators. We focus on doubly-sublinear IPPs, where both the verifier and honest prover use sublinear resources, and on (weighted) Hamming weight properties. For ordinary Hamming weight, we give a tolerant doubly-sublinear IPP. For completeness and soundness radii $\varepsilon_c<\varepsilon_f$ and gap $g=\varepsilon_f-\varepsilon_c$, a logarithmic-round instantiation uses $\widetilde{O}(1/g)$ verifier queries and $O(1/g^2)$ honest-prover queries, improving the cubic dependence of Amir, Goldreich, and Rothblum (ITCS 2025). We prove matching query lower bounds up to polylogarithmic factors. For distribution-weighted Hamming weight, black-box sampling requires $Θ(1/g^2)$ verifier samples but only $\widetilde{O}(1/g)$ evaluations; the quadratic sample complexity is necessary in the interior regime. With chosen-randomness access, the problem reduces to ordinary Hamming weight, yielding $\widetilde{O}(1/g)$ calls and evaluations. If the parties' evaluators disagree arbitrarily on a $ρ$-fraction of the distribution and by at most $γ$ elsewhere, our protocols remain doubly sublinear whenever $g>2κ$, where $κ=ρ+(1-ρ)γ$. Applications include auditing accuracy, group fairness, calibration, harmlessness, usefulness, and average-case robustness.

Authors: Geoffroy Couteau, Nikolas Melissaris, Tamara Paris

We study interactive proofs of proximity (IPPs) for model evaluation, where a resource-limited verifier interacts with an untrusted prover, typically the model owner, to certify statistical properties of a model under an unknown input distribution. Our formulation separates sampling the input distribution from querying the model and evaluating its output; distinguishes real audit data (black-box sampling) from generated data (chosen-randomness, or gray-box, access to the sampler); and allows the prover and verifier to use different evaluators. We focus on doubly-sublinear IPPs, where both the verifier and honest prover use sublinear resources, and on (weighted) Hamming weight properties. For ordinary Hamming weight, we give a tolerant doubly-sublinear IPP. For completeness and soundness radii $\varepsilon_c<\varepsilon_f$ and gap $g=\varepsilon_f-\varepsilon_c$, a logarithmic-round instantiation uses $\widetilde{O}(1/g)$ verifier queries and $O(1/g^2)$ honest-prover queries, improving the cubic dependence of Amir, Goldreich, and Rothblum (ITCS 2025). We prove matching query lower bounds up to polylogarithmic factors. For distribution-weighted Hamming weight, black-box sampling requires $Θ(1/g^2)$ verifier samples but only $\widetilde{O}(1/g)$ evaluations; the quadratic sample complexity is necessary in the interior regime. With chosen-randomness access, the problem reduces to ordinary Hamming weight, yielding $\widetilde{O}(1/g)$ calls and evaluations. If the parties' evaluators disagree arbitrarily on a $ρ$-fraction of the distribution and by at most $γ$ elsewhere, our protocols remain doubly sublinear whenever $g>2κ$, where $κ=ρ+(1-ρ)γ$. Applications include auditing accuracy, group fairness, calibration, harmlessness, usefulness, and average-case robustness.

Probing the classical complexity of quantum dynamics experiments

from arXiv: Computational Complexity

Authors: Thomas Schuster, Andreas Elben

A confluence of recent works has shown that many quantum circuits and dynamics are efficiently simulable by classical algorithms that track local information, even when conventional complexity measures such as the entanglement and magic are high. Here, we introduce a novel measure of complexity, the reactivity, to capture this new method of classical attack. Unlike conventional complexity measures, the reactivity does not capture a property of a quantum state or operator in isolation, but rather a quantum experiment as a whole. We provide numerical and rigorous evidence that quantum experiments with low reactivity are simple by a host of measures: they are efficient to classically simulate, learn, and fast-forward. This motivates the search for quantum experiments with high reactivity, which may evade these simplistic features. To this end, we introduce easily implementable experimental protocols---dubbed Pauli path spectroscopy---that allow one to efficiently measure the reactivity of any quantum experiment of interest. Our protocols are applicable even when the experiment itself is beyond the reach of classical simulation.

Authors: Thomas Schuster, Andreas Elben

A confluence of recent works has shown that many quantum circuits and dynamics are efficiently simulable by classical algorithms that track local information, even when conventional complexity measures such as the entanglement and magic are high. Here, we introduce a novel measure of complexity, the reactivity, to capture this new method of classical attack. Unlike conventional complexity measures, the reactivity does not capture a property of a quantum state or operator in isolation, but rather a quantum experiment as a whole. We provide numerical and rigorous evidence that quantum experiments with low reactivity are simple by a host of measures: they are efficient to classically simulate, learn, and fast-forward. This motivates the search for quantum experiments with high reactivity, which may evade these simplistic features. To this end, we introduce easily implementable experimental protocols---dubbed Pauli path spectroscopy---that allow one to efficiently measure the reactivity of any quantum experiment of interest. Our protocols are applicable even when the experiment itself is beyond the reach of classical simulation.

Towards Kinematic Actionable Infeasibility Detection in Motion Planning

from arXiv: Computational Geometry

Authors: Aayush Rath, Lakshya Jindal, Antony Thomas

Motion planning in robotics requires not only computing collision-free paths but also certifying infeasibility when no such path exists. Complete methods are limited to low-dimensional spaces, while sampling-based planners scale efficiently but cannot provide finite-time infeasibility certificates, leaving this problem largely unresolved in high-dimensional spaces. In this letter, we present a geometry-driven framework for certifying infeasibility through an explicit resolution-dependent analysis of configuration space topology. Leveraging signed distance field representations, the proposed method traces separating manifolds induced by obstacle boundaries directly in configuration space, enabling both detection of infeasibility and identification of the specific geometric cause. To address computational challenges, we develop a parallel frontier-expansion algorithm that exploits GPU acceleration for efficient simplicial reconstruction in high-dimensional spaces. We validate the approach on 4-DOF and 5-DOF robot scenarios, certifying infeasibility within seconds for 4-DOF cases and under four minutes for 5-DOF cases. We further discuss avenues for improving scalability to higher-dimensional spaces.

Authors: Aayush Rath, Lakshya Jindal, Antony Thomas

Motion planning in robotics requires not only computing collision-free paths but also certifying infeasibility when no such path exists. Complete methods are limited to low-dimensional spaces, while sampling-based planners scale efficiently but cannot provide finite-time infeasibility certificates, leaving this problem largely unresolved in high-dimensional spaces. In this letter, we present a geometry-driven framework for certifying infeasibility through an explicit resolution-dependent analysis of configuration space topology. Leveraging signed distance field representations, the proposed method traces separating manifolds induced by obstacle boundaries directly in configuration space, enabling both detection of infeasibility and identification of the specific geometric cause. To address computational challenges, we develop a parallel frontier-expansion algorithm that exploits GPU acceleration for efficient simplicial reconstruction in high-dimensional spaces. We validate the approach on 4-DOF and 5-DOF robot scenarios, certifying infeasibility within seconds for 4-DOF cases and under four minutes for 5-DOF cases. We further discuss avenues for improving scalability to higher-dimensional spaces.

Curve Band Depth: A Band-Based Data Depth for Unparameterized Planar Curves

from arXiv: Computational Geometry

Authors: Siyi Wang, Alexandre Leblanc, Paul D. McNicholas

We introduce \emph{curve band depth} (CBD), a band-based data depth for samples of \emph{unparameterized} planar curves. CBD is motivated by band depth and modified band depth for functional data, but targets trajectory data. Unlike the halfspace-based curve depth of \citet{de2021depth} and the curve stabbing depth of \citet{durocher2023csd}, CBD is defined through a geometric band region generated by two curves, and measures the arc-length proportion of a target curve lying inside such bands. We develop a CBD family consisting of an integral version (int-CBD), an infimal version (inf-CBD), and a fast-walk variant (FW-CBD). The fast-walk band is a narrower band construction contained in the global convex-combination band. We establish boundedness, vanishing at infinity, and similarity invariance for these constructions, together with a Borel-measurability result for the induced depth maps under a mild measurability assumption. A length-penalized variant is proposed for samples with heterogeneous curve lengths. We implement the methods via arc-length sampling and polygonal approximations, and evaluate them through classification of overlapping handwriting data and MNIST-derived digit curves, online-signature screening on \texttt{MOBISIG}, and an exploratory clustering task based on decomposed band contributions.

Authors: Siyi Wang, Alexandre Leblanc, Paul D. McNicholas

We introduce \emph{curve band depth} (CBD), a band-based data depth for samples of \emph{unparameterized} planar curves. CBD is motivated by band depth and modified band depth for functional data, but targets trajectory data. Unlike the halfspace-based curve depth of \citet{de2021depth} and the curve stabbing depth of \citet{durocher2023csd}, CBD is defined through a geometric band region generated by two curves, and measures the arc-length proportion of a target curve lying inside such bands. We develop a CBD family consisting of an integral version (int-CBD), an infimal version (inf-CBD), and a fast-walk variant (FW-CBD). The fast-walk band is a narrower band construction contained in the global convex-combination band. We establish boundedness, vanishing at infinity, and similarity invariance for these constructions, together with a Borel-measurability result for the induced depth maps under a mild measurability assumption. A length-penalized variant is proposed for samples with heterogeneous curve lengths. We implement the methods via arc-length sampling and polygonal approximations, and evaluate them through classification of overlapping handwriting data and MNIST-derived digit curves, online-signature screening on \texttt{MOBISIG}, and an exploratory clustering task based on decomposed band contributions.

Updating a Discrete Morse Vector Field for a Lower Star Filtration Vineyard

from arXiv: Computational Geometry

Authors: Kevin Woytowich, Nkechi Nnadi, Elizabeth Munch

In this paper, we provide a construction of an acyclic discrete vector field that is compatible with a total order associated with a lower star filtration on a simplicial complex, called the colex vector field. We show that the colex vector field induces a filtered acyclic vector field, whose resulting Morse complex computes the persistent homology of the lower star filtration on the underlying simplicial complex. We show that the colex vector field can be recomputed quickly when the vertex function that induces the lower star filtration is modified via order-adjacent vertex swaps. We provide a framework for storing and computing the number of paths between cells in the simplicial complex, as well as a method to update these values quickly when the colex vector field changes. Finally, we provide publicly available proof-of-concept code for the ideas shown. When applying it to the Persistent Homology Transform, we show that its runtime is comparable to a more standard matrix reduction approach.

Authors: Kevin Woytowich, Nkechi Nnadi, Elizabeth Munch

In this paper, we provide a construction of an acyclic discrete vector field that is compatible with a total order associated with a lower star filtration on a simplicial complex, called the colex vector field. We show that the colex vector field induces a filtered acyclic vector field, whose resulting Morse complex computes the persistent homology of the lower star filtration on the underlying simplicial complex. We show that the colex vector field can be recomputed quickly when the vertex function that induces the lower star filtration is modified via order-adjacent vertex swaps. We provide a framework for storing and computing the number of paths between cells in the simplicial complex, as well as a method to update these values quickly when the colex vector field changes. Finally, we provide publicly available proof-of-concept code for the ideas shown. When applying it to the Persistent Homology Transform, we show that its runtime is comparable to a more standard matrix reduction approach.

Learned Localized Mesh Refinement

from arXiv: Computational Geometry

Authors: Xiao Zhan, Chrystiano Araújo, Kangle Deng, Maneesh Agrawala, Hsueh-Ti Derek Liu, Mina Konaković Luković

We present a neural method for adaptive triangle mesh refinement, in which an autoregressive model adds geometric detail to selected regions of an input mesh while leaving the rest unchanged, a key capability for efficiently allocating mesh budget. Existing upsampling methods struggle to achieve this. Classical subdivision schemes refine triangulation without semantic awareness of the underlying shape or the ability to recover geometric details missing from a coarse input. Recent neural mesh models generate shapes globally, sacrificing region-specific control. We propose a novel tokenizer that yields combinatorially many valid upsampling trajectories from a single mesh. Trained on such data, our locally-conditioned autoregressive architecture allows for direct manipulation of topology and geometry within target regions of an input mesh. We validate our method against state-of-the-art approaches and demonstrate its ability to perform adaptive upsampling with region-selective control, a capability absent from existing approaches. This unlocks inference-time view-dependent refinement, physics-aware region refinement, and coarse-shape conditioned novel mesh synthesis. We provide code at github.com/seanxzhan/learned-localized-mesh-refinement/.

Authors: Xiao Zhan, Chrystiano Araújo, Kangle Deng, Maneesh Agrawala, Hsueh-Ti Derek Liu, Mina Konaković Luković

We present a neural method for adaptive triangle mesh refinement, in which an autoregressive model adds geometric detail to selected regions of an input mesh while leaving the rest unchanged, a key capability for efficiently allocating mesh budget. Existing upsampling methods struggle to achieve this. Classical subdivision schemes refine triangulation without semantic awareness of the underlying shape or the ability to recover geometric details missing from a coarse input. Recent neural mesh models generate shapes globally, sacrificing region-specific control. We propose a novel tokenizer that yields combinatorially many valid upsampling trajectories from a single mesh. Trained on such data, our locally-conditioned autoregressive architecture allows for direct manipulation of topology and geometry within target regions of an input mesh. We validate our method against state-of-the-art approaches and demonstrate its ability to perform adaptive upsampling with region-selective control, a capability absent from existing approaches. This unlocks inference-time view-dependent refinement, physics-aware region refinement, and coarse-shape conditioned novel mesh synthesis. We provide code at https://github.com/seanxzhan/learned-localized-mesh-refinement/.

Using Persistent Homology to Analyze Access to Heterogeneous-Quality Resources and Heterogeneous-Severity Nuisances

from arXiv: Computational Geometry

Authors: Sarah Tymochko, Gillian Grindstaff, Abigail Hickok, Jiajie Luo, Mason A. Porter

We develop a framework to use multiparameter persistent homology (PH) to examine access to heterogeneous-quality resources and exposure to heterogeneous-severity nuisances in a geographic region. Persistent homology, which is a type of topological data analysis {(TDA)}, has been employed previously to examine resource coverage. Unlike prior approaches, which used one-parameter PH to study resource coverage and nuisance exposure, our method accounts for heterogeneous-quality resources. Our framework, which employs a computationally-efficient approximation of multiparameter PH, allows one to study access to any resource ({or} exposure of any nuisance) using any notion of quality (or severity). Using the city of Chicago as an example region, we employ our framework to detect clusters of poor access to public parks, overexposure to landfills, and both underexposure and overexposure to pubs and bars.

Authors: Sarah Tymochko, Gillian Grindstaff, Abigail Hickok, Jiajie Luo, Mason A. Porter

We develop a framework to use multiparameter persistent homology (PH) to examine access to heterogeneous-quality resources and exposure to heterogeneous-severity nuisances in a geographic region. Persistent homology, which is a type of topological data analysis {(TDA)}, has been employed previously to examine resource coverage. Unlike prior approaches, which used one-parameter PH to study resource coverage and nuisance exposure, our method accounts for heterogeneous-quality resources. Our framework, which employs a computationally-efficient approximation of multiparameter PH, allows one to study access to any resource ({or} exposure of any nuisance) using any notion of quality (or severity). Using the city of Chicago as an example region, we employ our framework to detect clusters of poor access to public parks, overexposure to landfills, and both underexposure and overexposure to pubs and bars.

A data structure for quotient flag complexes

from arXiv: Data Structures and Algorithms

Authors: Konstantin Sorokin, Aleksandr Levin, Maxim Beketov, Anton Ayzenberg

Vietoris-Rips filtrations, which are standard in topological data analysis, are flag complexes, and a simplex tree stores these without any attaching data. In this paper we ask what survives of this economy when a flag complex $K$ is divided by a subcomplex $A$, each connected component of $A$ being crushed to a point. Such a quotient is a CW complex whose cells are the simplices of $K\setminus A$, but their attaching maps are no longer implicit. We show that for flag $K$ the face order of the quotient is strictly graded exactly when $A$ is flag, and that the surviving labelled cells are determined by those of dimension at most 3. For $m$-flag pairs the threshold is $2m+1$, and it drops to $m+2$ when $K$ is flag. The prescribed cells form a regular CW decomposition only when $A$ is full in $K$. These results justify the QF-tree: a cell table that stores, for each surviving simplex, its ordered list of $d+1$ facets with collapsed facets flagged, indexed by a trie of quotient-vertex words. For bounded dimension its size is linear in the number of surviving simplices plus the retained provenance, and we derive and verify a simple formula for the collapsed fraction above which it is smaller than the homotopy-equivalent cone model. Because a collapse changes the attaching data only on the closed star of $A$, the QF-tree can also be applied locally inside a simplex tree. For a ball-shaped $A$ in the sampled Vietoris--Rips regime the closed star is a thin shell, and the median compact budget is below the cone model at every sampled radius. An accompanying library, modelled on Gudhi, implements the QF-tree, its local variant, an editable layer with local quotient updates, gluing, disc attachment, induced maps, cup products, fundamental-group presentations and zigzag persistence, and provided experiments separate the cost of maintaining a quotient from the cost of the algebra computed on it.

Authors: Konstantin Sorokin, Aleksandr Levin, Maxim Beketov, Anton Ayzenberg

Vietoris-Rips filtrations, which are standard in topological data analysis, are flag complexes, and a simplex tree stores these without any attaching data. In this paper we ask what survives of this economy when a flag complex $K$ is divided by a subcomplex $A$, each connected component of $A$ being crushed to a point. Such a quotient is a CW complex whose cells are the simplices of $K\setminus A$, but their attaching maps are no longer implicit. We show that for flag $K$ the face order of the quotient is strictly graded exactly when $A$ is flag, and that the surviving labelled cells are determined by those of dimension at most 3. For $m$-flag pairs the threshold is $2m+1$, and it drops to $m+2$ when $K$ is flag. The prescribed cells form a regular CW decomposition only when $A$ is full in $K$. These results justify the QF-tree: a cell table that stores, for each surviving simplex, its ordered list of $d+1$ facets with collapsed facets flagged, indexed by a trie of quotient-vertex words. For bounded dimension its size is linear in the number of surviving simplices plus the retained provenance, and we derive and verify a simple formula for the collapsed fraction above which it is smaller than the homotopy-equivalent cone model. Because a collapse changes the attaching data only on the closed star of $A$, the QF-tree can also be applied locally inside a simplex tree. For a ball-shaped $A$ in the sampled Vietoris--Rips regime the closed star is a thin shell, and the median compact budget is below the cone model at every sampled radius. An accompanying library, modelled on Gudhi, implements the QF-tree, its local variant, an editable layer with local quotient updates, gluing, disc attachment, induced maps, cup products, fundamental-group presentations and zigzag persistence, and provided experiments separate the cost of maintaining a quotient from the cost of the algebra computed on it.

Structure-Adaptive Tree Field Integrators

from arXiv: Data Structures and Algorithms

Authors: Millend Roy, Soham Samal, Ivan Zelich, Krzysztof Marcin Choromanski

We present a new class of near-linear algorithms for efficiently integrating general tensor fields defined on trees with distance dependent kernels, the Structure-Adaptive Tree Field Integrators (STAD-TFIs). STAD-TFIs exploit the tree's underlying structure through decompositions built around path backbones and single vertex separators, and use two-dimensional fast Fourier transforms to compute interactions jointly. By exploiting this structural information, STAD-TFIs achieve more computationally efficient integration than their regular efficient tree field integrators (TFI) counterparts. We provide a detailed theoretical analysis of our proposed approach and complement it with an exhaustive empirical evaluation, ranging from speed tests on synthetic trees, through accelerated Sinkhorn-based relaxations of the Optimal Transport algorithms on real meshes, to Topological Attention Transformers for vision tasks. To the best of our knowledge, we provide some of the first results showing that efficient to compute and accurate relaxations of the geodesic Sinkhorn-based solutions of the Optimal Transport problem can be derived by applying fast TFI methods.

Authors: Millend Roy, Soham Samal, Ivan Zelich, Krzysztof Marcin Choromanski

We present a new class of near-linear algorithms for efficiently integrating general tensor fields defined on trees with distance dependent kernels, the Structure-Adaptive Tree Field Integrators (STAD-TFIs). STAD-TFIs exploit the tree's underlying structure through decompositions built around path backbones and single vertex separators, and use two-dimensional fast Fourier transforms to compute interactions jointly. By exploiting this structural information, STAD-TFIs achieve more computationally efficient integration than their regular efficient tree field integrators (TFI) counterparts. We provide a detailed theoretical analysis of our proposed approach and complement it with an exhaustive empirical evaluation, ranging from speed tests on synthetic trees, through accelerated Sinkhorn-based relaxations of the Optimal Transport algorithms on real meshes, to Topological Attention Transformers for vision tasks. To the best of our knowledge, we provide some of the first results showing that efficient to compute and accurate relaxations of the geodesic Sinkhorn-based solutions of the Optimal Transport problem can be derived by applying fast TFI methods.

Tight Efficiency Guarantees for Strategyproof Linear Regression

from arXiv: Data Structures and Algorithms

Authors: Yichen Huang, Yuqi Pan, Michael Mitzenmacher, Milind Tambe, Yiling Chen

We study the trade-off between squared-error accuracy and incentive compatibility in linear regression. Agents report private labels associated with publicly known features and prefer predictions close to their true labels. Ordinary least squares (OLS) need not elicit truthful reports. For regression with $d$ parameters, we design a deterministic group-strategyproof mechanism achieving a $(d+1)$-approximation to the least-squares optimum and prove optimality even among universally strategyproof randomized mechanisms, answering an open question of Chen et al. (EC 2018). Relaxing universal strategyproofness to strategyproofness in expectation reveals a sharp separation: squared individual loss retains the factor $d+1$, while absolute individual loss admits the tight ratio $2-1/(\lceil d/2\rceil+1)$.

Authors: Yichen Huang, Yuqi Pan, Michael Mitzenmacher, Milind Tambe, Yiling Chen

We study the trade-off between squared-error accuracy and incentive compatibility in linear regression. Agents report private labels associated with publicly known features and prefer predictions close to their true labels. Ordinary least squares (OLS) need not elicit truthful reports. For regression with $d$ parameters, we design a deterministic group-strategyproof mechanism achieving a $(d+1)$-approximation to the least-squares optimum and prove optimality even among universally strategyproof randomized mechanisms, answering an open question of Chen et al. (EC 2018). Relaxing universal strategyproofness to strategyproofness in expectation reveals a sharp separation: squared individual loss retains the factor $d+1$, while absolute individual loss admits the tight ratio $2-1/(\lceil d/2\rceil+1)$.

Improved SDP Coloring of 3-Colorable Graphs from Recursive Gaussian Certificates

from arXiv: Data Structures and Algorithms

Authors: Ijay Narang, Yukai Tang

We give a randomized polynomial-time algorithm that, for every fixed $\varepsilon > 0$, colors every $3$-colorable $n$-vertex graph using $O\bigl(n^{(13-\sqrt{97})/18+\varepsilon}\bigr) \approx O\bigl(n^{0.17506+\varepsilon}\bigr)$ colors, improving upon the previous best bound of $O(n^{0.19539})$ from Bansal, Huang, and Lee. Our improvement comes from analyzing higher-level neighborhoods through a recursive description of failure in Gaussian SDP rounding. If the rounding returns too small an independent set, it produces local Gaussian certificates at every vertex of a nonempty induced subgraph. We propagate these certificates along walks to higher-level neighborhoods by defining a recursive certificate structure and proving a strengthened cover-composition lemma, which refines the one of Arora, Chlamt{á}{č}, and Charikar. We then construct a bounded potential function that increases by a fixed positive amount at every propagation step, yielding a contradiction. Consequently, the rounding must produce a sufficiently large independent set.

Authors: Ijay Narang, Yukai Tang

We give a randomized polynomial-time algorithm that, for every fixed $\varepsilon > 0$, colors every $3$-colorable $n$-vertex graph using $O\bigl(n^{(13-\sqrt{97})/18+\varepsilon}\bigr) \approx O\bigl(n^{0.17506+\varepsilon}\bigr)$ colors, improving upon the previous best bound of $O(n^{0.19539})$ from Bansal, Huang, and Lee. Our improvement comes from analyzing higher-level neighborhoods through a recursive description of failure in Gaussian SDP rounding. If the rounding returns too small an independent set, it produces local Gaussian certificates at every vertex of a nonempty induced subgraph. We propagate these certificates along walks to higher-level neighborhoods by defining a recursive certificate structure and proving a strengthened cover-composition lemma, which refines the one of Arora, Chlamt{á}{č}, and Charikar. We then construct a bounded potential function that increases by a fixed positive amount at every propagation step, yielding a contradiction. Consequently, the rounding must produce a sufficiently large independent set.

Geometry-Adaptive Mechanisms for Private Synthetic Data

from arXiv: Data Structures and Algorithms

Authors: Raoof Zare Moayedi, Amir R. Asadi, Mohammad Hossein Yassaee, Gholamali Aminian

Generating differentially private synthetic data with meaningful Wasserstein utility guarantees is challenging in high dimensions. For datasets of size \(n\) on $[0,1]^d$ with $d\ge2$, existing pure \(\varepsilon\)-differentially private mechanisms achieve expected $1$-Wasserstein error of order $(\varepsilon n)^{-1/d}$, reflecting the curse of dimensionality. While this rate is optimal in the worst case, it can be overly pessimistic when the data are supported on a lower-dimensional set. We formalize this through a multiscale packing-growth dimension $k$, which captures the geometric complexity of the support via the growth of packing numbers across scales. We propose \emph{Adaptive Pruned-PMM}, a pure $\varepsilon$-differentially private mechanism that combines private depth selection with our pruned variant of the Private Measure Mechanism (PMM) of He et al.\ (2023). The mechanism supports deeper, geometry-adapted hierarchies with expected running time $O\!\left(d(n+d)\log(\varepsilon n)\right)$, which is near-linear in $n$ for fixed dimension and privacy budget. Under an external multiscale packing-growth condition with dimension $k$, we show that, for fixed positive privacy budgets and fixed geometry, the expected $1$-Wasserstein error is of order $(\varepsilon n)^{-1/k}$ for $k>1$ as $n$ grows. We also prove a lower bound under a corresponding internal packing-growth condition, showing that the exponent $1/k$ is sharp within this framework.

Authors: Raoof Zare Moayedi, Amir R. Asadi, Mohammad Hossein Yassaee, Gholamali Aminian

Generating differentially private synthetic data with meaningful Wasserstein utility guarantees is challenging in high dimensions. For datasets of size \(n\) on $[0,1]^d$ with $d\ge2$, existing pure \(\varepsilon\)-differentially private mechanisms achieve expected $1$-Wasserstein error of order $(\varepsilon n)^{-1/d}$, reflecting the curse of dimensionality. While this rate is optimal in the worst case, it can be overly pessimistic when the data are supported on a lower-dimensional set. We formalize this through a multiscale packing-growth dimension $k$, which captures the geometric complexity of the support via the growth of packing numbers across scales. We propose \emph{Adaptive Pruned-PMM}, a pure $\varepsilon$-differentially private mechanism that combines private depth selection with our pruned variant of the Private Measure Mechanism (PMM) of He et al.\ (2023). The mechanism supports deeper, geometry-adapted hierarchies with expected running time $O\!\left(d(n+d)\log(\varepsilon n)\right)$, which is near-linear in $n$ for fixed dimension and privacy budget. Under an external multiscale packing-growth condition with dimension $k$, we show that, for fixed positive privacy budgets and fixed geometry, the expected $1$-Wasserstein error is of order $(\varepsilon n)^{-1/k}$ for $k>1$ as $n$ grows. We also prove a lower bound under a corresponding internal packing-growth condition, showing that the exponent $1/k$ is sharp within this framework.

On the Complexity of Forcing and Anti-Forcing Minimum Cuts

from arXiv: Data Structures and Algorithms

Authors: Tatsuya Gima, Yasuaki Kobayashi, Hiraku Morimoto, Yota Otachi

For an instance of a combinatorial optimization problem, a \emph{forcing set} is a set of elements such that there is a unique optimal solution including it. Symmetrically, an \emph{anti-forcing set} is a set of elements such that there is a unique optimal solution excluding it. In this paper, we study the problems of computing smallest forcing and anti-forcing sets for two classical cut problems, \textsc{Global Min Cut} and \textsc{Min $s$--$t$ Cut}. We also consider variants in which the optimal cut to be uniquely determined is given as input. For each of these problems, we either give a polynomial-time algorithm or prove \NP-completeness.

Authors: Tatsuya Gima, Yasuaki Kobayashi, Hiraku Morimoto, Yota Otachi

For an instance of a combinatorial optimization problem, a \emph{forcing set} is a set of elements such that there is a unique optimal solution including it. Symmetrically, an \emph{anti-forcing set} is a set of elements such that there is a unique optimal solution excluding it. In this paper, we study the problems of computing smallest forcing and anti-forcing sets for two classical cut problems, \textsc{Global Min Cut} and \textsc{Min $s$--$t$ Cut}. We also consider variants in which the optimal cut to be uniquely determined is given as input. For each of these problems, we either give a polynomial-time algorithm or prove \NP-completeness.

On the Guo-Fang-Lu Algorithm for Komlos Discrepancy

from arXiv: Data Structures and Algorithms

Authors: Nikhil Bansal

We give an exposition of the recent polynomial time algorithm of Guo, Fang, and Lu for the Komlos problem. We simplify various arguments, and highlight the key new spectral potential idea and how the algorithm follows naturally from it.

Authors: Nikhil Bansal

We give an exposition of the recent polynomial time algorithm of Guo, Fang, and Lu for the Komlos problem. We simplify various arguments, and highlight the key new spectral potential idea and how the algorithm follows naturally from it.

On Diverse Solutions to Max-k-CSP and Bounded Degree k-SAT

from arXiv: Data Structures and Algorithms

Authors: Mayank Goswami, Adarsh Srinivasan

We study the problem of generating diverse solutions to Max-$k$-CSP and bounded-degree $k$-SAT, focusing on two distinct metrics: constraint diversity and variable diversity. For constraint diversity, the goal is to output $s \geq 2$ assignments to the CSP such that each assignment satisfies a $c$-fraction of the constraints, while maximizing the diversity among the $0$-$1$ indicator vectors of satisfied constraints in the Hamming metric. By reducing this to a multi-criteria optimization problem, we design $poly(n,s)$ time approximation algorithms that return s assignments achieving provable bi-criteria guarantees on both the fraction of satisfied constraints and diversity of the constraint vectors. For variable diversity, the objective is to maximize the Hamming distance between the assignments, while also maximizing the number of constraints satisfied. For Max-$k$-CSP instances when the desired number of solutions is $s=2^{O(n)}$, we implicitly represent these diverse approximate solutions by constructing linear codes within the solution space. Finally, we investigate variable diversity for $k$-SAT in the Lovász Local Lemma regime. In this setting, we establish NP-hardness for the exact diversity problem (computing the diameter of the solution space) and provide a polynomial-time approximation algorithm to efficiently generate diverse satisfying assignments.

Authors: Mayank Goswami, Adarsh Srinivasan

We study the problem of generating diverse solutions to Max-$k$-CSP and bounded-degree $k$-SAT, focusing on two distinct metrics: constraint diversity and variable diversity. For constraint diversity, the goal is to output $s \geq 2$ assignments to the CSP such that each assignment satisfies a $c$-fraction of the constraints, while maximizing the diversity among the $0$-$1$ indicator vectors of satisfied constraints in the Hamming metric. By reducing this to a multi-criteria optimization problem, we design $poly(n,s)$ time approximation algorithms that return s assignments achieving provable bi-criteria guarantees on both the fraction of satisfied constraints and diversity of the constraint vectors. For variable diversity, the objective is to maximize the Hamming distance between the assignments, while also maximizing the number of constraints satisfied. For Max-$k$-CSP instances when the desired number of solutions is $s=2^{O(n)}$, we implicitly represent these diverse approximate solutions by constructing linear codes within the solution space. Finally, we investigate variable diversity for $k$-SAT in the Lovász Local Lemma regime. In this setting, we establish NP-hardness for the exact diversity problem (computing the diameter of the solution space) and provide a polynomial-time approximation algorithm to efficiently generate diverse satisfying assignments.

Efficient Dynamic Algorithms for Graph Neural Networks with Non-Linear Propagation

from arXiv: Data Structures and Algorithms

Authors: Kiarash Banihashem, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Silvio Lattanzi, Danny Mittal

Graph Neural Networks (GNNs) are widely used for representation learning on graphs, but most methods assume static topologies, making them inefficient on evolving networks where edges change over time. Existing dynamic approaches either model graph evolution through temporal GNN architectures without focusing on efficient dynamic maintenance, or are restricted to linear propagation models based on Personalized PageRank. In this work, we study how to efficiently maintain node representations for non-linear GNN propagation under edge insertions and deletions. The propagation has no learned parameters, and only a classifier applied afterward is trained. For a broad class of standard activation functions, we develop a residual-based dynamic algorithm that selectively propagates local errors via push operations, maintaining an approximation to the evolving fixed point without full recomputation. We prove that our method achieves amortized $O(1/ε)$ update time per graph change under a degree-normalized error guarantee. Our approach uses a potential-based analysis in a degree-scaled norm and, in contrast to prior work on the linear case, requires no randomness assumptions on either the update sequence or the input vector. For the linear special case, we additionally provide an exact dynamic algorithm via low-rank matrix inverse updates. Experiments on benchmark datasets show that incorporating non-linearity improves accuracy while preserving efficient update performance, yielding a scalable and theoretically grounded method for maintaining this propagation on dynamic graphs.

Authors: Kiarash Banihashem, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Silvio Lattanzi, Danny Mittal

Graph Neural Networks (GNNs) are widely used for representation learning on graphs, but most methods assume static topologies, making them inefficient on evolving networks where edges change over time. Existing dynamic approaches either model graph evolution through temporal GNN architectures without focusing on efficient dynamic maintenance, or are restricted to linear propagation models based on Personalized PageRank. In this work, we study how to efficiently maintain node representations for non-linear GNN propagation under edge insertions and deletions. The propagation has no learned parameters, and only a classifier applied afterward is trained. For a broad class of standard activation functions, we develop a residual-based dynamic algorithm that selectively propagates local errors via push operations, maintaining an approximation to the evolving fixed point without full recomputation. We prove that our method achieves amortized $O(1/ε)$ update time per graph change under a degree-normalized error guarantee. Our approach uses a potential-based analysis in a degree-scaled norm and, in contrast to prior work on the linear case, requires no randomness assumptions on either the update sequence or the input vector. For the linear special case, we additionally provide an exact dynamic algorithm via low-rank matrix inverse updates. Experiments on benchmark datasets show that incorporating non-linearity improves accuracy while preserving efficient update performance, yielding a scalable and theoretically grounded method for maintaining this propagation on dynamic graphs.

Unlocking Geodesic Gromov-Wasserstein Distances for 3D Modeling

from arXiv: Data Structures and Algorithms

Authors: Krzysztof Marcin Choromanski, Derek Long, Ananya Parashar, Dwaipayan Saha

\textit{Gromov-Wasserstein Distances} (GWDs) provide quantitative ways of comparing probabilistic distributions defined on different metric spaces by applying techniques from the optimal transport theory. As such, GWD can be potentially useful in a large variety of applications ranging from graph matching problems to 3D object detection. However its practical use at scale is significantly limited by cubic time complexity computations involving dense intra-space distance matrices. Even though in the Euclidean metric spaces several techniques (e.g. involving scalable kernel methods) were proposed to address it, to the best of our knowledge, analogous techniques for general geodesic distances on manifolds, or shortest-path distance on graphs in their discretized variants, were not developed. In this paper, we present \textbf{E}fficient \textbf{G}eodesic \textbf{Gro}mov-\textbf{W}asserstein methods (EGGroW), a new class of efficient algorithms designed to calculate geodesic Gromov-Wasserstein distances with entropic Sinkhorn-like approaches, leveraging recently introduced \textit{GenusSink} methods \citep{genussink} and the theory of random features. We provide important downstream applications, namely: 3D pose estimation and 3D template detection. In the latter setting, we formulate a partial 3D template recovery as a staged problem: capacity-constrained scene selection is followed by semi-relaxed recovery of template visibility and correspondence. Our empirical findings show that EGGroW provides accurate solutions when standard Euclidean-based techniques fail and is characterized by light computational footprint, as our theoretical analysis predicts.

Authors: Krzysztof Marcin Choromanski, Derek Long, Ananya Parashar, Dwaipayan Saha

\textit{Gromov-Wasserstein Distances} (GWDs) provide quantitative ways of comparing probabilistic distributions defined on different metric spaces by applying techniques from the optimal transport theory. As such, GWD can be potentially useful in a large variety of applications ranging from graph matching problems to 3D object detection. However its practical use at scale is significantly limited by cubic time complexity computations involving dense intra-space distance matrices. Even though in the Euclidean metric spaces several techniques (e.g. involving scalable kernel methods) were proposed to address it, to the best of our knowledge, analogous techniques for general geodesic distances on manifolds, or shortest-path distance on graphs in their discretized variants, were not developed. In this paper, we present \textbf{E}fficient \textbf{G}eodesic \textbf{Gro}mov-\textbf{W}asserstein methods (EGGroW), a new class of efficient algorithms designed to calculate geodesic Gromov-Wasserstein distances with entropic Sinkhorn-like approaches, leveraging recently introduced \textit{GenusSink} methods \citep{genussink} and the theory of random features. We provide important downstream applications, namely: 3D pose estimation and 3D template detection. In the latter setting, we formulate a partial 3D template recovery as a staged problem: capacity-constrained scene selection is followed by semi-relaxed recovery of template visibility and correspondence. Our empirical findings show that EGGroW provides accurate solutions when standard Euclidean-based techniques fail and is characterized by light computational footprint, as our theoretical analysis predicts.

Parity Tests under Ties: A One-Test Lifting Theorem

from arXiv: Data Structures and Algorithms

Authors: Ron Kupfer

In the unrestricted polynomial decision-tree model, only the number of polynomial sign tests is charged. A parity test asks for the sign of a product of pairwise differences. Such tests underlie low-depth randomized algorithms for maximum finding and top-$k$ selection, but their usual analysis assumes distinct inputs because a tie makes the product vanish. We give a black-box lifting theorem that removes this assumption. After $O(\log n)$ polynomial tests determine the number of nonzero pairwise differences, every subsequent parity test is simulated by one polynomial test, consistently with a fixed lexicographic tie-breaking order. The simulator is an elementary symmetric polynomial in masked first and second powers of all pairwise differences. Thus a depth-$D$ parity-test tree on distinct inputs becomes a polynomial decision tree of depth $D+O(\log n)$ on arbitrary inputs, with no increase in randomized pointwise error for order-selection problems. We obtain maximum finding in depth $O(\log n[\log n+\log(1/δ)])$ with error $δ$, and top-$k$ selection in depth $O(\log^2 n+k\log n)$ with inverse-polynomial error, both without any promise on ties.

Authors: Ron Kupfer

In the unrestricted polynomial decision-tree model, only the number of polynomial sign tests is charged. A parity test asks for the sign of a product of pairwise differences. Such tests underlie low-depth randomized algorithms for maximum finding and top-$k$ selection, but their usual analysis assumes distinct inputs because a tie makes the product vanish. We give a black-box lifting theorem that removes this assumption. After $O(\log n)$ polynomial tests determine the number of nonzero pairwise differences, every subsequent parity test is simulated by one polynomial test, consistently with a fixed lexicographic tie-breaking order. The simulator is an elementary symmetric polynomial in masked first and second powers of all pairwise differences. Thus a depth-$D$ parity-test tree on distinct inputs becomes a polynomial decision tree of depth $D+O(\log n)$ on arbitrary inputs, with no increase in randomized pointwise error for order-selection problems. We obtain maximum finding in depth $O(\log n[\log n+\log(1/δ)])$ with error $δ$, and top-$k$ selection in depth $O(\log^2 n+k\log n)$ with inverse-polynomial error, both without any promise on ties.

Near-Optimal Distributed Domination in Planar Graphs

from arXiv: Data Structures and Algorithms

Authors: Wojciech Wawrzyniak

We give a deterministic $(8+\varepsilon)$-approximation for minimum dominating set on planar graphs in a constant number of rounds of the LOCAL model, for every $\varepsilon>0$. This improves the previous ratio $11+\varepsilon$ obtained by Heydt et al. The ratio is near-optimal in this model: its leading constant is only one above the known lower bound of $7$. Our result closes three quarters of the previous gap, reducing it from $4$ to $1$. Our main contribution is a sharp structural bound. For any dominating set $D$, assigning each vertex outside $D$ to a neighboring center gives disjoint owner blocks. If $k_x$ counts the other blocks containing a neighbor of $x$, then $\sum_{x\notin D}(k_x-2)^+\le(4|D|-12)^+$, where $z^+=\max\{z,0\}$. The bound holds for every such assignment, and equality holds for arbitrarily large minimum dominating sets. We use this bound in their three-phase framework, with new parameters and the same final linear-programming procedure. The algorithm requires neither a planar embedding nor the graph size, and its round bound depends only on $\varepsilon$. The transfer theorem of Bonamy et al. also gives a deterministic $(25+\varepsilon)$-approximation on graphs of bounded Euler genus, with a round bound depending only on $\varepsilon$ and the genus.

Authors: Wojciech Wawrzyniak

We give a deterministic $(8+\varepsilon)$-approximation for minimum dominating set on planar graphs in a constant number of rounds of the LOCAL model, for every $\varepsilon>0$. This improves the previous ratio $11+\varepsilon$ obtained by Heydt et al. The ratio is near-optimal in this model: its leading constant is only one above the known lower bound of $7$. Our result closes three quarters of the previous gap, reducing it from $4$ to $1$. Our main contribution is a sharp structural bound. For any dominating set $D$, assigning each vertex outside $D$ to a neighboring center gives disjoint owner blocks. If $k_x$ counts the other blocks containing a neighbor of $x$, then $\sum_{x\notin D}(k_x-2)^+\le(4|D|-12)^+$, where $z^+=\max\{z,0\}$. The bound holds for every such assignment, and equality holds for arbitrarily large minimum dominating sets. We use this bound in their three-phase framework, with new parameters and the same final linear-programming procedure. The algorithm requires neither a planar embedding nor the graph size, and its round bound depends only on $\varepsilon$. The transfer theorem of Bonamy et al. also gives a deterministic $(25+\varepsilon)$-approximation on graphs of bounded Euler genus, with a round bound depending only on $\varepsilon$ and the genus.

Single-Exponential Algorithms for Directed Feedback Vertex Set on Planar Digraphs

from arXiv: Data Structures and Algorithms

Authors: Daniel Lokshtanov, Saket Saurabh, Jie Xue

We consider Directed Feedback Vertex Set on planar digraphs, parameterized by the solution size $k$. We give a randomized algorithm with one-sided error running in time $(2+\sqrt5)^k n^{O(1)}= 4.24^k n^{O(1)}$, and a deterministic algorithm running in time $8.04^k n^{O(1)}$. Both algorithms use polynomial space. To the best of our knowledge, these are the first single-exponential fixed-parameter algorithms for Directed Feedback Vertex Set on planar digraphs. This contrasts with general digraphs, where the best known algorithms run in time $2^{O(k\log k)}(n+m)$, and whether a $2^{o(k\log k)}n^{O(1)}$-time algorithm exists remains a major open problem. Our main tool is an exact Euler-type counting identity for plane digraphs. It shows that every small solution must carry a large share of the vertices whose in- and out-arcs alternate in the embedding, while solutions avoiding such vertices can be computed by reducing to Directed Feedback Arc Set, which is known to be solvable in polynomial time on planar digraphs via the Lucchesi-Younger theorem.

Authors: Daniel Lokshtanov, Saket Saurabh, Jie Xue

We consider Directed Feedback Vertex Set on planar digraphs, parameterized by the solution size $k$. We give a randomized algorithm with one-sided error running in time $(2+\sqrt5)^k n^{O(1)}= 4.24^k n^{O(1)}$, and a deterministic algorithm running in time $8.04^k n^{O(1)}$. Both algorithms use polynomial space. To the best of our knowledge, these are the first single-exponential fixed-parameter algorithms for Directed Feedback Vertex Set on planar digraphs. This contrasts with general digraphs, where the best known algorithms run in time $2^{O(k\log k)}(n+m)$, and whether a $2^{o(k\log k)}n^{O(1)}$-time algorithm exists remains a major open problem. Our main tool is an exact Euler-type counting identity for plane digraphs. It shows that every small solution must carry a large share of the vertices whose in- and out-arcs alternate in the embedding, while solutions avoiding such vertices can be computed by reducing to Directed Feedback Arc Set, which is known to be solvable in polynomial time on planar digraphs via the Lucchesi-Younger theorem.

Online Covering with Maximum Delay under Subadditive Service Costs

from arXiv: Data Structures and Algorithms

Authors: Tianhang Lu, Runtian Ren, Shengcai Liu

We study online covering in which each instantaneous service pays its purchase cost and one maximum waiting time, with no effect on future requests. For static realizable services, monotone subadditivity suffices for optimal competitive ratios; submodularity is unnecessary. A normalized monotone subadditive lower-bound oracle with realization factor $ρ$ yields ratios $ρ+1$ deterministically and $1/(1-e^{-1/ρ})$ randomly against an oblivious adversary. Exact batch optimization gives the optimal constants $2$ and $e/(e-1)$. The randomized algorithm uses one global threshold on a seed-independent virtual-height trajectory, whose active time is a lower bound on the offline optimum. Weighted vertex cover gives a strict separation from submodularity on a three-edge bipartite path, with polynomial-time batch implementations through min-cut and LP rounding. An offline consecutive-batch normal form also transfers static approximation guarantees to the offline problem.

Authors: Tianhang Lu, Runtian Ren, Shengcai Liu

We study online covering in which each instantaneous service pays its purchase cost and one maximum waiting time, with no effect on future requests. For static realizable services, monotone subadditivity suffices for optimal competitive ratios; submodularity is unnecessary. A normalized monotone subadditive lower-bound oracle with realization factor $ρ$ yields ratios $ρ+1$ deterministically and $1/(1-e^{-1/ρ})$ randomly against an oblivious adversary. Exact batch optimization gives the optimal constants $2$ and $e/(e-1)$. The randomized algorithm uses one global threshold on a seed-independent virtual-height trajectory, whose active time is a lower bound on the offline optimum. Weighted vertex cover gives a strict separation from submodularity on a three-edge bipartite path, with polynomial-time batch implementations through min-cut and LP rounding. An offline consecutive-batch normal form also transfers static approximation guarantees to the offline problem.

Differentially Private Approximation of the John Ellipsoid

from arXiv: Data Structures and Algorithms

Authors: Bar Mahpud, Daniel Omer, Or Sheffet

We study the problem of approximating the John ellipsoid (JE) of a given (centrally symmetric) polytope of $n$ constraints in a Euclidean space under differential privacy (DP). We give the first differentially private algorithm for this problem under the standard model, where neighboring datasets may differ arbitrarily in one a single constraint. Our work also extends to the complimentary problem of Minimum Enclosing Ellipsoid of $n$ points in the Euclidean space. Our approach is based on the recent non-private multiplicative-weights algorithm of~\cite{pmlr-v99-cohen19a}. First we introduce a non-private generalization of the Cohen et al algorithm, yielding a $(1+γ)$-approximation of the JE problem while violating at most $κn$ constraints in $O(\log(1/κ)/γ)$ iterations. This variant works by projecting the intermediate weights assigned to the constraints onto the set of $κ$-dense distributions, similarly to~\cite{bun2020efficientnoisetolerantprivatelearning}. We then design a $ρ$-zCDP variant of this algorithm by adding Gaussian noise to the weighted covariance matrix aggregated in each step of the algorithm. Under a mild goodness assumption on the data we can assert that the resulting noisy matrix is close to the true matrix, thereby achieving essentially the same guarantee as the non-private algorithm provided sufficiently many input points. Thus our method achieves an efficient DP poly-time algorithm under concrete sample complexity bounds.

Authors: Bar Mahpud, Daniel Omer, Or Sheffet

We study the problem of approximating the John ellipsoid (JE) of a given (centrally symmetric) polytope of $n$ constraints in a Euclidean space under differential privacy (DP). We give the first differentially private algorithm for this problem under the standard model, where neighboring datasets may differ arbitrarily in one a single constraint. Our work also extends to the complimentary problem of Minimum Enclosing Ellipsoid of $n$ points in the Euclidean space. Our approach is based on the recent non-private multiplicative-weights algorithm of~\cite{pmlr-v99-cohen19a}. First we introduce a non-private generalization of the Cohen et al algorithm, yielding a $(1+γ)$-approximation of the JE problem while violating at most $κn$ constraints in $O(\log(1/κ)/γ)$ iterations. This variant works by projecting the intermediate weights assigned to the constraints onto the set of $κ$-dense distributions, similarly to~\cite{bun2020efficientnoisetolerantprivatelearning}. We then design a $ρ$-zCDP variant of this algorithm by adding Gaussian noise to the weighted covariance matrix aggregated in each step of the algorithm. Under a mild goodness assumption on the data we can assert that the resulting noisy matrix is close to the true matrix, thereby achieving essentially the same guarantee as the non-private algorithm provided sufficiently many input points. Thus our method achieves an efficient DP poly-time algorithm under concrete sample complexity bounds.

A 51-Addition Alternative-Basis Kernel for Rank-23 $3\times3$ Matrix Multiplication

from arXiv: Data Structures and Algorithms

Authors: Joshua Stapleton, Andrew Perminov

We give a rank-23 algorithm for $3\times3$ matrix multiplication using 51 additions and subtractions in alternative bases. The input and output conversions require five further additions, giving 56 additions in ordinary coordinates. The construction combines linear-program reduction with a search over sparse basis changes, and the resulting programs are distributed as a machine-checkable certificate. We verify correctness by exact coefficient expansion and distinguish the kernel cost from the cost of a complete multiplication.

Authors: Joshua Stapleton, Andrew Perminov

We give a rank-23 algorithm for $3\times3$ matrix multiplication using 51 additions and subtractions in alternative bases. The input and output conversions require five further additions, giving 56 additions in ordinary coordinates. The construction combines linear-program reduction with a search over sparse basis changes, and the resulting programs are distributed as a machine-checkable certificate. We verify correctness by exact coefficient expansion and distinguish the kernel cost from the cost of a complete multiplication.

Stability Dichotomies for Boolean Constraint Satisfaction Problems

from arXiv: Data Structures and Algorithms

Authors: Chunyang Wang, Yuichi Yoshida

We study the stability of Boolean constraint satisfaction problems (CSPs) through the notion of average sensitivity (Varma and Yoshida, SODA 2021; SICOMP 2023). It measures the expected $1$-Wasserstein distance between an algorithm's output distributions before and after the deletion of a uniformly chosen constraint, using the unnormalized Hamming metric. We establish two dichotomies for every finite Boolean constraint language $Γ$, where $n\geq 2$ denotes the number of variables in an instance. For stable solvability, exactly one of the following holds: $\bullet$ either there is an algorithm that solves $\mathrm{CSP}(Γ)$ and has average sensitivity $O_Γ(1)$ for all satisfiable instances; $\bullet$ or every algorithm that solves $\mathrm{CSP}(Γ)$ has average sensitivity $Ω_Γ(n)$ on satisfiable instances of arbitrarily large $n$. The first alternative holds if and only if $Γ$ has finite duality: unsatisfiability can be witnessed on a bounded number of variables. For stable approximability, where a $(1-\varepsilon)$-approximation violates at most an $\varepsilon$-fraction of the constraints in expectation, exactly one of the following holds: $\bullet$ either for every $\varepsilon\in(0,1]$, there is an algorithm that $(1-\varepsilon)$-approximates $\mathrm{CSP}(Γ)$ with average sensitivity $O_Γ(\varepsilon^{-1}\log n)$ for all satisfiable instances; $\bullet$ or there exists $\varepsilon_Γ\in(0,1]$ such that every algorithm that $(1-\varepsilon_Γ)$-approximates $\mathrm{CSP}(Γ)$ has average sensitivity $Ω_Γ(n)$ on satisfiable instances of arbitrarily large $n$. The first alternative holds if and only if $Γ$ has bounded width: local consistency checks on bounded sets of variables detect unsatisfiability.

Authors: Chunyang Wang, Yuichi Yoshida

We study the stability of Boolean constraint satisfaction problems (CSPs) through the notion of average sensitivity (Varma and Yoshida, SODA 2021; SICOMP 2023). It measures the expected $1$-Wasserstein distance between an algorithm's output distributions before and after the deletion of a uniformly chosen constraint, using the unnormalized Hamming metric. We establish two dichotomies for every finite Boolean constraint language $Γ$, where $n\geq 2$ denotes the number of variables in an instance. For stable solvability, exactly one of the following holds: $\bullet$ either there is an algorithm that solves $\mathrm{CSP}(Γ)$ and has average sensitivity $O_Γ(1)$ for all satisfiable instances; $\bullet$ or every algorithm that solves $\mathrm{CSP}(Γ)$ has average sensitivity $Ω_Γ(n)$ on satisfiable instances of arbitrarily large $n$. The first alternative holds if and only if $Γ$ has finite duality: unsatisfiability can be witnessed on a bounded number of variables. For stable approximability, where a $(1-\varepsilon)$-approximation violates at most an $\varepsilon$-fraction of the constraints in expectation, exactly one of the following holds: $\bullet$ either for every $\varepsilon\in(0,1]$, there is an algorithm that $(1-\varepsilon)$-approximates $\mathrm{CSP}(Γ)$ with average sensitivity $O_Γ(\varepsilon^{-1}\log n)$ for all satisfiable instances; $\bullet$ or there exists $\varepsilon_Γ\in(0,1]$ such that every algorithm that $(1-\varepsilon_Γ)$-approximates $\mathrm{CSP}(Γ)$ has average sensitivity $Ω_Γ(n)$ on satisfiable instances of arbitrarily large $n$. The first alternative holds if and only if $Γ$ has bounded width: local consistency checks on bounded sets of variables detect unsatisfiability.

SparseDesign: Scaling Exact Coding-Sequence Design

from arXiv: Data Structures and Algorithms

Authors: Hao Lin, Jingjin Yu

Exact optimization of synonymous coding sequences under a joint folding-energy and codon-usage objective is limited by expensive dynamic-programming splits and large working sets. \textsc{SparseDesign} applies candidate sparsification to the multiloop recurrence of a Turner~2004 dangle-0 solver over a weighted codon automaton. A direct branch is retained only when it strictly improves on every partitionable or endpoint-unpaired realization of the same endpoint states. We prove equivalence to the dense recurrence in real arithmetic, under an explicit scalar branch-interface assumption. With $N$ automaton states, edge set $E$ and $Z$ retained candidates, multiloop work is $O(N^2+N|E|+NZ)$; worst-case time remains cubic for bounded-width automata and total memory remains quadratic. Endpoint ownership permits parallel candidate construction without locks. While synthetic stress families can benefit little from sparsification and exhibit near-quadratic candidate growth, natural proteins show substantial candidate-count reductions. In our 7,600-task campaign, the 2,000-protein human-table panel has median retention of only 3.53\% at $λ=0$ and 2.15\% at $λ=4$, corresponding to approximately 28.3-fold and 46.4-fold reductions relative to all feasible direct intervals. The primary performance experiments use an AMD EPYC 7313 server. For human Dp427c (11,031 nt, $λ=0$), 16-thread packed \textsc{SparseDesign} achieves five-run medians of 236.54 seconds wall-clock time and 14.43 GiB peak RSS. Compared with the single-thread local dense LinearDesign fork on the same server (4,912 seconds, 402.10 GiB RSS), this gives a 20.8-fold wall-clock speedup and a 27.9-fold peak-memory reduction. On a Core i9-14900KF commodity PC with 64 GiB RAM, the same input, layout and thread count achieve 126.42 seconds and 14.43 GiB RSS.

Authors: Hao Lin, Jingjin Yu

Exact optimization of synonymous coding sequences under a joint folding-energy and codon-usage objective is limited by expensive dynamic-programming splits and large working sets. \textsc{SparseDesign} applies candidate sparsification to the multiloop recurrence of a Turner~2004 dangle-0 solver over a weighted codon automaton. A direct branch is retained only when it strictly improves on every partitionable or endpoint-unpaired realization of the same endpoint states. We prove equivalence to the dense recurrence in real arithmetic, under an explicit scalar branch-interface assumption. With $N$ automaton states, edge set $E$ and $Z$ retained candidates, multiloop work is $O(N^2+N|E|+NZ)$; worst-case time remains cubic for bounded-width automata and total memory remains quadratic. Endpoint ownership permits parallel candidate construction without locks. While synthetic stress families can benefit little from sparsification and exhibit near-quadratic candidate growth, natural proteins show substantial candidate-count reductions. In our 7,600-task campaign, the 2,000-protein human-table panel has median retention of only 3.53\% at $λ=0$ and 2.15\% at $λ=4$, corresponding to approximately 28.3-fold and 46.4-fold reductions relative to all feasible direct intervals. The primary performance experiments use an AMD EPYC 7313 server. For human Dp427c (11,031 nt, $λ=0$), 16-thread packed \textsc{SparseDesign} achieves five-run medians of 236.54 seconds wall-clock time and 14.43 GiB peak RSS. Compared with the single-thread local dense LinearDesign fork on the same server (4,912 seconds, 402.10 GiB RSS), this gives a 20.8-fold wall-clock speedup and a 27.9-fold peak-memory reduction. On a Core i9-14900KF commodity PC with 64 GiB RAM, the same input, layout and thread count achieve 126.42 seconds and 14.43 GiB RSS.

Single-or-Sample: Online Fair Allocation for Combinatorial Agents

from arXiv: Data Structures and Algorithms

Authors: Shuchi Chawla, Zhiyi Huang, Pooja Kulkarni, Ruta Mehta, Parnian Shahkar

We study the problem of fairly allocating $m$ indivisible goods among $n$ agents who arrive online, under the notion of maximin share (MMS) fairness. Fair allocation with online arrivals is notoriously challenging: prior work achieves constant-factor MMS guarantees only when agents' preferences belong to a set of valuation functions known in advance, while no guarantees were known without such prior information. We develop a new randomized online algorithm for additive and submodular valuations, that we call Single-or-Sample, and that achieves a constant-factor approximation to MMS simultaneously for all agents, with constant probability. The algorithm requires no prior knowledge about the agents' valuations, and works against adversarial (oblivious) inputs. We further establish a fundamental tradeoff between approximation and success probability. Specifically, for any $c \ge 1$, no online algorithm can guarantee a $1/c$-approximation to MMS with probability exceeding $1 - 1/c^2$, even for binary additive valuations. This rules out constant MMS with probability asymptotically closer to $1$ than a constant. For XOS, the lower bound is much stronger: no algorithm can achieve even $1/\log\log n$-MMS to all agents with a constant probability. We complement this lower bound with an algorithm for the regime $c\in Ω(\log n)$, namely $1/c$-MMS to all agents with probability $(1-O(1/c))$, and show that this tradeoff is tight for XOS. Our constant-factor algorithm introduces several new ideas, combining greedy submodular maximization with randomized allocation and single-item reduction. A key technical ingredient is a new approach for analyzing iterative sampling without replacement. We develop concentration bounds that apply to a broad class of adaptive processes with complex dependencies across elements and rounds, which may be of independent interest.

Authors: Shuchi Chawla, Zhiyi Huang, Pooja Kulkarni, Ruta Mehta, Parnian Shahkar

We study the problem of fairly allocating $m$ indivisible goods among $n$ agents who arrive online, under the notion of maximin share (MMS) fairness. Fair allocation with online arrivals is notoriously challenging: prior work achieves constant-factor MMS guarantees only when agents' preferences belong to a set of valuation functions known in advance, while no guarantees were known without such prior information. We develop a new randomized online algorithm for additive and submodular valuations, that we call Single-or-Sample, and that achieves a constant-factor approximation to MMS simultaneously for all agents, with constant probability. The algorithm requires no prior knowledge about the agents' valuations, and works against adversarial (oblivious) inputs. We further establish a fundamental tradeoff between approximation and success probability. Specifically, for any $c \ge 1$, no online algorithm can guarantee a $1/c$-approximation to MMS with probability exceeding $1 - 1/c^2$, even for binary additive valuations. This rules out constant MMS with probability asymptotically closer to $1$ than a constant. For XOS, the lower bound is much stronger: no algorithm can achieve even $1/\log\log n$-MMS to all agents with a constant probability. We complement this lower bound with an algorithm for the regime $c\in Ω(\log n)$, namely $1/c$-MMS to all agents with probability $(1-O(1/c))$, and show that this tradeoff is tight for XOS. Our constant-factor algorithm introduces several new ideas, combining greedy submodular maximization with randomized allocation and single-item reduction. A key technical ingredient is a new approach for analyzing iterative sampling without replacement. We develop concentration bounds that apply to a broad class of adaptive processes with complex dependencies across elements and rounds, which may be of independent interest.

Tree Search With Distributional Predictions

from arXiv: Data Structures and Algorithms

Authors: Michael Dinitz, Bob Dong

Learning-augmented algorithms use machine-learned predictions to improve classical algorithmic guarantees when the predictions are accurate, while retaining rigorous performance guarantees when they are not. We study this paradigm for search on trees. Given a tree $T$ containing an unknown target vertex $t$, an algorithm may query any vertex $v$ and learn which neighbor of $v$ lies on the unique path from $v$ to $t$. The goal is to find $t$ using as few queries as possible. We consider the distributional setting, in which the target is drawn from an unknown distribution $p$ and the algorithm is given a predicted distribution $\widehat p$ of unknown quality. We give an algorithm with expected query complexity $O\left(H(p)+k\log η\right)$, where $H(p)$ is the Shannon entropy of the true distribution and $η$ is the earth mover's distance between $p$ and $\widehat p$ in the tree metric. We also provide a matching lower bound that shows our algorithm is asymptotically tight. Finally, experiments on real-world and synthetic trees show that our prediction-based algorithm can use substantially fewer queries than a simple baseline that trusts the prediction completely.

Authors: Michael Dinitz, Bob Dong

Learning-augmented algorithms use machine-learned predictions to improve classical algorithmic guarantees when the predictions are accurate, while retaining rigorous performance guarantees when they are not. We study this paradigm for search on trees. Given a tree $T$ containing an unknown target vertex $t$, an algorithm may query any vertex $v$ and learn which neighbor of $v$ lies on the unique path from $v$ to $t$. The goal is to find $t$ using as few queries as possible. We consider the distributional setting, in which the target is drawn from an unknown distribution $p$ and the algorithm is given a predicted distribution $\widehat p$ of unknown quality. We give an algorithm with expected query complexity $O\left(H(p)+k\log η\right)$, where $H(p)$ is the Shannon entropy of the true distribution and $η$ is the earth mover's distance between $p$ and $\widehat p$ in the tree metric. We also provide a matching lower bound that shows our algorithm is asymptotically tight. Finally, experiments on real-world and synthetic trees show that our prediction-based algorithm can use substantially fewer queries than a simple baseline that trusts the prediction completely.

Efficient Support Recovery of Mixtures of Sparse Linear Classifiers with Less Measurements

from arXiv: Data Structures and Algorithms

Authors: Xiaxin Li, Arya Mazumdar

The support recovery problem in mixture of linear classifiers intends to identify which features actually matter when data is generated by a mixture of several linear decision rules. In particular, the aim is to recover the support (nonzero coordinates) of $l$ unknown $k$-sparse vectors from sign measurements. Each measurement is generated by selecting one of the $l$ vectors uniformly at random, and returning the sign of its inner product with a chosen measurement vector. In this paper, we propose adaptive and non-adaptive schemes that significantly improve upon prior results by reducing the number of measurements and achieving sublinear decoding time simultaneously. In particular, our adaptive constructions substantially reduce measurements compared to existing approaches, while also lowering decoding complexity from super-quadratic to sublinear in the ambient dimension. We further provide a non-adaptive scheme that improves previous measurement bounds while maintaining efficient decoding. Overall, our approach yields a more efficient trade-off between sample complexity and decoding time for support recovery in mixture models than previously known methods.

Authors: Xiaxin Li, Arya Mazumdar

The support recovery problem in mixture of linear classifiers intends to identify which features actually matter when data is generated by a mixture of several linear decision rules. In particular, the aim is to recover the support (nonzero coordinates) of $l$ unknown $k$-sparse vectors from sign measurements. Each measurement is generated by selecting one of the $l$ vectors uniformly at random, and returning the sign of its inner product with a chosen measurement vector. In this paper, we propose adaptive and non-adaptive schemes that significantly improve upon prior results by reducing the number of measurements and achieving sublinear decoding time simultaneously. In particular, our adaptive constructions substantially reduce measurements compared to existing approaches, while also lowering decoding complexity from super-quadratic to sublinear in the ambient dimension. We further provide a non-adaptive scheme that improves previous measurement bounds while maintaining efficient decoding. Overall, our approach yields a more efficient trade-off between sample complexity and decoding time for support recovery in mixture models than previously known methods.

The Tight Upper Bound on the Number of Distinct Squares in Circular Words

from arXiv: Data Structures and Algorithms

Authors: Rikuya Hamai

A square is a word $xx$, where $x$ is nonempty. We show that a circular word of length $n$ contains at most $\lfloor 3n/2 \rfloor$ distinct squares of length at most $n$. The proof combines known results relating squares to circuits in Rauzy graphs. The coefficient $3/2$ agrees with the known lower bound.

Authors: Rikuya Hamai

A square is a word $xx$, where $x$ is nonempty. We show that a circular word of length $n$ contains at most $\lfloor 3n/2 \rfloor$ distinct squares of length at most $n$. The proof combines known results relating squares to circuits in Rauzy graphs. The coefficient $3/2$ agrees with the known lower bound.

Tight Convergence Bounds for the Classical Kaczmarz Method

from arXiv: Data Structures and Algorithms

Authors: Runbo Yu, Jelena Diakonikolas

The classical method of Kaczmarz, introduced in 1937, is a textbook iterative method for solving linear systems $A x = b.$ Despite its widespread use, particularly in the context of solving inverse problems where it is commonly included in software packages within Matlab, Python, and Julia, the precise characterization of convergence has long been deemed difficult to obtain. While different bounds on convergence rates have been established, they are largely unsatisfying as they cannot explain the classical, cyclic-update method's efficient convergence in practice. In this work, we obtain a tight characterization of convergence of the classical Kaczmarz method, establishing both linear and sublinear convergence bounds. The worst-case tight (i.e., exactly attained by some instances in the considered family) bounds are expressed in terms of a fixed matrix that depends only on $A,$ but are not fully interpretable in terms of the matrix spectrum and row correlations, which had been observed to have an impact on convergence. We thus provide relaxations of these bounds that are fully expressible in terms of matrix row norms, row correlations, rank, and extremal positive singular values. The provided relaxed bounds explain one-cycle convergence in special cases where the matrix rows are all either parallel or orthogonal to each other. We further argue that the dependence on different parameters appearing in the bounds is necessary and within a small constant factor of the best attainable in the worst case. Finally, our bounds explain why the classical cyclic update is faster than the randomized one when matrix rows are weakly correlated, which is often observed in inverse problems where the cyclic method is used.

Authors: Runbo Yu, Jelena Diakonikolas

The classical method of Kaczmarz, introduced in 1937, is a textbook iterative method for solving linear systems $A x = b.$ Despite its widespread use, particularly in the context of solving inverse problems where it is commonly included in software packages within Matlab, Python, and Julia, the precise characterization of convergence has long been deemed difficult to obtain. While different bounds on convergence rates have been established, they are largely unsatisfying as they cannot explain the classical, cyclic-update method's efficient convergence in practice. In this work, we obtain a tight characterization of convergence of the classical Kaczmarz method, establishing both linear and sublinear convergence bounds. The worst-case tight (i.e., exactly attained by some instances in the considered family) bounds are expressed in terms of a fixed matrix that depends only on $A,$ but are not fully interpretable in terms of the matrix spectrum and row correlations, which had been observed to have an impact on convergence. We thus provide relaxations of these bounds that are fully expressible in terms of matrix row norms, row correlations, rank, and extremal positive singular values. The provided relaxed bounds explain one-cycle convergence in special cases where the matrix rows are all either parallel or orthogonal to each other. We further argue that the dependence on different parameters appearing in the bounds is necessary and within a small constant factor of the best attainable in the worst case. Finally, our bounds explain why the classical cyclic update is faster than the randomized one when matrix rows are weakly correlated, which is often observed in inverse problems where the cyclic method is used.