The death of mathematics

Yesterday will be remembered as the day mathematics died. OpenAI posted the solution on hundreds of mathematical problems on a github repository, all generated by ChatGPT. Some have been formalized, some not, and the pdfs are uniformly unreadable, as except the Riemann one they haven’t been human edited. OpenAI doesn’t even claim all of them are correct, because of course nobody could have checked such a huge amount of material in such a short amount of time.

To add insult to injury their announcement talks about how the release was “informed” by the “recommendations” of a supposedly independent advisory group about “best practices”. Yeah, it was only informed because they didn’t follow the very mild recommendations. I don’t even know how they could have done more damage to mathematics if that was their goal.

It was clear that the days of mathematics were numbered since the announcement of the disproof of the Jacobian conjecture. But one might have thought that they would limit themselves to these high-profile problems that make for good advertisement, and let mathematicians use their models to solve niche problems at their own pace. Which would incidentally allow for checking and a decent write-up of the result. Nope, they decided they wanted to solve everything of interest and scoop absolutely everyone, and dump the resulting turd without the least amount of polishing.

Why? I guess they decided the tiny marketing value of claiming these solutions was higher than whatever money they could make by selling subscriptions to mathematicians. After all there’s no money in pure mathematics, their real customers are big corporations salivating at the thought of creating mass unemployment.

Maybe you think I’m being hyperbolic, and this hasn’t killed mathematics. Even if problem-solving is now outside of human capabilities, surely we can still do the other half of mathematics, theory-building? Indeed LLMs haven’t shown any prowess at theory-building, this is still a purely human endeavour. But the motivation for theory-building is problem-solving. Who is going to start a career in mathematics knowing that they’ll never be able to solve any problem? Maybe they could try to solve problems indirectly, for example by building the theory necessary to tackle P != NP? Some fools will certainly try, but it should be obvious by now that it’s only a matter of time until AI can do theory-building as well.

It’s obviously also just a matter of time until AI becomes capable of doing theoretical physics. Plenty of morons believe it already can, and are clogging the arXiv and Quantum with nonsense. But starting a career in theoretical physics now is suicidal. Maybe experimental physics will last long enough, who know, progress in robotics seems stubbornly slow.

Posted in Uncategorised | Leave a comment

The tragedy of the commons in scientific publishing

The tragedy of the commons is a well-known concept in economics: selfish action by many individuals destroy a common resource, to the detriment of all that used it. Even if each individual is aware of the destruction they have no incentive to restrain themselves, as they would just get less of the resource while the others destroy it anyway. In small communities it’s possible to solve it by appealing to ethics and reputation, but in large communities regulation is necessary.

We’re seeing it happen in scientific publishing now. Here the common resource is the attention of the scientific community, in the form of editors, referees, or readers. Everybody needs it, a paper that doesn’t get attention might as well not exist. And the destruction of the resource is being caused by the submission of too many papers, to the arXiv or to journals.

The arXiv has become unreadable. For example, today’s quant-ph has 157 submissions. Who has the time to read this? Every day? When I was young I used to start my day reading quant-ph, it had one or two dozen abstracts, I could easily skim through everything. Now? Forget it. The way I handle it is to follow authors that I know, and I suspect most people do the same. Which is of course extremely unfair to new authors that might have done something great but are not well-connected. The problem is not only reader time, but moderator time1: the poor volunteers need to go through all this in time for the announcement next morning. The situation is so dire that the arXiv has introduced yesterday a rate-limit of two papers per month per submitter. Much too generous, in my opinion, it should be two per year.

Journals are struggling with this as well. At least Quantum is, as I know it personally: it has a growing backlog of submissions that the editors haven’t even managed to look at, and after the editors manage to process it, it’s becoming ever harder to find referees as they are overloaded as well.

The cause is of course the advent of LLMs. Previously the problem was kept in check by the difficulty of writing a scientific paper. Now, though, any LLM can write something that looks like a scientific paper instantly. And way too many authors believe it is an actual scientific paper and submit it to the arXiv or to a journal. Or even if they don’t believe it, they submit it anyway because there’s no punishment for submitting bad papers, the more the merrier. I saw a particularly egregious example in yesterday’s arXiv: two almost-identical papers, by the same authors: arXiv:2609.31832 and arXiv:2609.31837. They even have slightly different versions of Figure 1, doubtless in order to not get caught by a plagiarism detector.

Such papers completely drown legitimate research. And no, it’s not easy to ignore them: it takes time to read the abstract and realise they’re worthless, it takes time to give them a desk-rejection as an editor.

Since the scientific publishing system is so decentralised, it’s very difficult to implement effective regulation. One of the few centralised institutions we have is the arXiv, so I commend them on the rate-limiting. But this still leaves the journals to fend for themselves. Rate-limiting is not effective at a journal level, because the flood of submissions is dispersed among journal, we don’t get several submissions from the same authors. One thing that might work is to give a penalty for a desk-rejection, e.g. you can’t submit another manuscript for a year if you get desk-rejected. That would give a strong incentive for authors to only submit quality work. But I don’t think that this is plausible at the moment, the punishment is way too harsh for our culture.

What is more plausible is to lower the quality of desk-rejections: instead of reading the paper and discussing it with the editorial board, an editor can simply skim the paper and make the desk-rejection individually. This will of course increase the rate of mistakes, with more high-quality research being rejected and low-quality research being sent to the referees. But what else can we do? Right now everything is just getting ignored.

This won’t solve the problem for the referees, though. On the contrary, they’ll get even more overloaded as we send them submissions as fast as we get them, and with worse curation to boot. For them the only thing that can help is raising the acceptance threshold of the journal. Or people stop using the arXiv as their notebook.

Posted in Uncategorised | 6 Comments

Fuck proof digestion

It happened again. In today’s arXiv Jef Pauwels claims the solution of the I3322 problem: namely that the maximal violation of the I3322 inequality cannot be reached with finite-dimensional strategies, and indeed it is the limit of the Pál-Vértesi strategy. Of course, the result was found by AI. Pauwels claims he simply gave ChatGPT the problem and told it he believed it could do it. And three hours later, a problem that had been open for almost two decades was solved.

Where is the glory in that, though? Anybody could have written that prompt, and indeed one Seth Douglas, who is apparently not even a researcher, claimed the same result a month earlier. Echoing Terence Tao, Pauwels claims the glory now lies in “proof digestion”, making a readable paper out of raw LLM output. Fair enough, unlike the earlier NPA non-convergence result, his paper is actually readable. But then he says he used Claude to do it.

As you might have guessed from the title of the blog post, I disagree.

First of all, I didn’t get into science to polish LLM turds. The reason I accepted low pay, long hours, moving countries, and instability is because I love doing science. Not giving problems to a machine and giving it a pep talk. Even less trying to make sense of LLM drivel. I suspect I’m not alone here. Indeed the students nowadays don’t seem to have motivation to learn. Why bother, if there’s no scientific work left? Almost all of them are just trying to cheat their way to a degree with LLMs. And that’s a drastic change, when I was an undergrad those who paid somebody else to write their thesis were the exception, not the rule.

Secondly, I don’t think there’s anything preventing LLMs from successfully doing digestion (even if it’s not clear how much Claude helped in Pauwels’ case). Trying to find a special role that only humans can do is a fool’s errand. Tao acknowledges that, and tries to reserve for humans the role of being the “bosses” of mathematics, deciding which proofs are relevant and interesting and should be taught. Taught to whom? Relevant to whom? Who the fuck is going to care?

I think this is the death of pure research. When we did the research ourselves, we could tell society that it should pay us because we were learning something, and teaching something, and perhaps even discovering something that might in a century be actually useful. But why should society pay us to type the prompts and “digest” the results? There is no public funding for sommeliers, those who just want to appreciate wine need to pay for it themselves.2

Applied research remains, of course. Society wants a result, and pays people to get it, doesn’t matter how they do it. I just think it will get much more expensive when people no longer do it out of love.

Posted in Uncategorised | 6 Comments

A strange affair in the nonlocality community

This week an amazing result has been announced on the arXiv: a proof that the NPA hierarchy does not converge exactly at a finite level even in the CHSH scenario. This is something that I tried to prove myself but didn’t manage.

However, instead of joy the announcement brought confusion and worry. The strange timeline was

  1. 14/07/2026 – Anubhav Chaturvedi submits the manuscript “No Finite NPA Level Characterizes the Complete Quantum Set in the Simplest Bell Scenario” to PRL.2
  2. 15/07/2026 – Anton Pakhunov submits the manuscript “No finite level of the NPA hierarchy is exact for the doubly-tilted CHSH functional near the critical tilt” to the arXiv.
  3. 16/07/2026 – Anubhav Chaturvedi submits the manuscript “No Finite NPA Level Characterizes the Complete Quantum Set in the Simplest Bell Scenario” to the arXiv.

I don’t believe the timing can be a coincidence, but I also don’t know how could Pakhunov possibly know of Anubhav’s submission to PRL. Moreover, the manuscripts have little in common besides claiming the same result, their claimed proofs are different.

Anubhav is an assistant professor in Gdańsk. He has been working on this problem for years, and I know him personally; we worked together on Self-testing tilted strategies for maximal loophole-free nonlocality, which is the basis for both manuscripts.

Pakhunov is an independent researcher, who doesn’t have any result in nonlocality besides this one. He has four pre-prints in total on the arXiv, starting on April this year, and they are on unrelated subjects.

I would be naturally biased in favour of Anubhav, but I’m not going to jump to conclusions and accuse Pakhunov of malfeasance. I wanted first and foremost understand what the fuck did just happen. I wrote to both of them to hear their sides of the story.

Pakhunov tells me that he studied physical chemistry at the Mendeleev University of Chemical Technology, and has since then worked as a programmer and a devops engineer. His research consists of using AI to search the literature and attempt to solve open problems. It means he can work very fast on disparate subjects; he claims he started working on the NPA finite convergence question on the 10th of July, so merely five days before he put it on the arXiv. But that also means he doesn’t actually know anything about the subject: the very first sentence of his abstract says that we “determined the quantum maximum of the doubly-tilted CHSH functionals by self-testing
methods”, which is false: I found the quantum maximum by directly solving the optimisation problem using Lagrange multipliers and Gröbner bases.

Now Anubhav tells me that he didn’t present his results anywhere or told them to anyone, so there’s no way Pakhunov could have copied them even if he wanted to. The idea that he could have hacked the PRL submission system and produced a different proof in a single day is ridiculous. However, Anubhav tells me that he used the free version of ChatGPT to write up his results, so this is a possible way for them to leak, as anything you tell it is included in the training data.

Therefore a possible explanation for the strange timeline that doesn’t require anyone to be lying or plagiarising is that ChatGPT “learned” Anubhav’s results, and when Pakhunov asked AI to solve the problem it repackaged Anubhav’s solution. That explanation is not very satisfactory because it still leaves a bit of a coincidence in the timing, and also because Pakhunov told me he used Claude instead of ChatGPT.

Of course, I’m not addressing the most important question: are their proofs correct? I tried to read the manuscripts, but both are unreadable. Furthermore, I firmly believe that no human should be condemned to read the output of AI, and that includes me. Unless somebody else steps up, we’re never gonna know.

Posted in Uncategorised | 5 Comments

Science and AI

I’ve just read this essay by the famous mathematician Terence Tao together with an artist about the use of AI in mathematics. It’s well worth a read and very well-written2. It focuses more on familiar territory of the problems caused by the currently existing AIs. Scientifically we have the human-centipedification of knowledge: students using AI to do their homework, professors using AI to grade the students’ homework, researchers using AI to write papers, referees using AI to write reports on such papers, and so on. Socially we have the massive loss of comfortable, safe jobs like call centre operator, software developer, illustrator, musician, writer, etc., further weakening labour and inevitable increasing the capital’s share of income. Environmentally we have the massive construction of data centres running on the dirtiest possible energy, just when climatologies have concluded that the Gulf stream will most likely collapse after all2.

But everybody knows this. What I want to talk about is the future problems they barely mention: what happens when the AI can do scientific research better than us? Note that I’m not talking about “if” but “when”. The idea that there’s something magical about human cognition that can’t be replicated in silicon is plain ridiculous, that I’m only mentioning here because it’s so widespread. It doesn’t mean that current LLMs can do it, though. They have fundamental limitations that prevent them from ever being capable of scientific research: they cannot learn, they don’t have a sense of time, they lack consciousness and a will. Another revolution in AI will be necessary for it to happen. But it will happen.

Here Tao believes that we will be able to live with better-than-human AI like chess players can live with chess engines they can never hope to beat. I think that’s a desperately naïve point of view. The vast majority of chess players do it for fun, not as a job. And chess engines can’t do their actual job. Scientists, on the other hand, do it almost exclusively professionally3 Why would society pay us to do research when it can be done better and cheaper by machines? It obviously won’t.

What will we do then? I can think of three possibilities:

  1. Do a communist revolution and spend our days masturbating, smoking weed, and playing video games.
  2. Do the jobs that AI/robots still can’t handle.
  3. Die.

I don’t think the first alternative is likely; at this point the billionaires will control all the killbots making revolution impossible. And even if some miracle happened it’s not a very appealing future: masturbating, smoking weed and playing video games is fun but gets old really fast. We need more than that in life.

If the first alternative is excluded, the second alternative is almost tautological. We are definitely not going to do the jobs that AI/robots can handle. What are we doing then? If robotics continue to lag behind AI looks like menial jobs, in a rather ironic reversal of the first industrial revolution. And when they do catch up I guess there will always exist billionaires with a fetish for human servants?

Other than that we can of course die, perhaps at the hand of the aforementioned killbots, in order to make space for a continent-wide golf course.

To conclude, I think the future is going to be an unmitigated disaster, and the best we can do is delay the inevitable. But even that seems to much to ask; in reality we have lunatics like Fields medallist Tim Gowers enthusiastically teaching AI to be better mathematicians. Come on, that’s like the Aztecs trying to make Columbus arrive sooner.

UPDATE: a little after I wrote this post, Gowers posted hyping up a minor result in combinatorics found by ChatGPT. Seriously, why would you ever want to use AI for that? That’s pure mathematics, completely devoid of practical applications. The reason people do it is because they find it fun, and the reason society pays people to do it is because it wants to form mathematicians. If AI solves the problem you get neither fun nor mathematicians out of it.

Posted in Uncategorised | 7 Comments

The open access movement has failed

The open access movement has taken over the world. Now pretty much every grant you get comes with an open access mandate, a requirement to publish the resulting papers in an openly accessible manner. It has also captured the public’s imagination, lay people now what open access is, and even those who are never going to read the papers find it suspicious in the rare cases where they are not open access.

We also have plenty of open access journals now. The best is of course Quantum, of which I’m proudly an editor. It is simultaneously of high quality and incredibly cheap (the fee is only 600€, and can be waived). Right now it is struggling with its success: the number of submissions has increased much faster than the number of editors, resulting in long delays. This is an eminently solvable problem, and Quantum is actively recruiting new editors, so this is not the failure I’m talking about.

The failure is that Quantum is pretty much unique. Other journals are usually either high quality and expensive (like PRX Quantum, which charges $3590), or low quality and expensive (like Entropy, which charges CHF 2600). None are cheap. Worst offenders are those that love to publish “high-impact” bullshit like Science Advances, charging $5450, and Nature Communications, charging 7350€. More offensive are those that adopt the “hybrid” open access model, which should be more precisely called “double-dipping” model, as they demand both a subscription and an article processing charge, for instance PRL, charging $4140, and Nature, charging an amazing €10690.

This is a failure because the main point of the open access movement wasn’t getting access to the papers. We have always managed to get access to papers somehow, either through arXiv, sci-hub, or simply asking the authors. I have personally done all three several times, because I was either working somewhere which didn’t pay the subscription I needed, or I couldn’t be arsed to configure the proxy.

The point was making publishing cheap. It’s obscene that the publishing industry has the profit margins of a literal gold mine, while doing almost nothing of the work: they don’t pay for doing the research in the first place, writing the manuscripts, or refereeing them. Sometimes they pay professional editors, sometimes they pay for (terrible) typesetting. Maybe they also print the damn papers somewhere, who knows, I haven’t even seen a printed journal for more than a decade.

The idea was that a closed-access journal had a monopoly on the papers it published, and a reader can’t substitute one paper for another: it really needs that one. This allowed publishers to charge whatever they wanted as subscription fees. With open access, the journals would no longer have any monopoly, but would have to compete with eachother, leading to lower prices.

It didn’t work. Profit margins are still commonly above 30%, and show no sign of decreasing. It amounts to billions of euros per years, and it is a significant drain on scientific research. Why? One might think that a reason is the double-dipping model of open access that I mentioned above, which prevents the “free market” from working. If that were the case fully open access journals would be cheaper, and they are not. It is in any case easy to fix, one can simply tweak the open access mandates to require publishing in fully open access journals.

Another possibility would be to forbid publishing in for-profit journals. That also wouldn’t work, as APS shows that it is possible to make a lot of profit even while being a non-profit. They simply can’t resist the amount of money authors are willing to pay for publishing their papers, and they gladly push their article processing charges to the stratosphere, and spend the money in the non-publishing arms of APS (as they can’t legally make a profit).

The fundamental problem is that scientific research is now a big business, managed throw objective metrics by non-experts. As any such system, it can be gamed, and oh do people game it. You want to get grants, you want to get a permanent position, you’d better publish a lot of papers. Publishing a bullshit paper in a predatory journal is beneficial to pad your CV, so there’s no surprise that companies like MDPI and Hindawi appeared to supply that service. CHF 2600 for a little CV padding? Apparently it’s a fair price, a lot of people pay that. Now publishing a bullshit paper in a high-impact journal? That can even make your career. No surprise that people are willing to pay 7350€ for a Nature Communications. For a permanent position, this is peanuts. People even pay it from their own pockets, in case there’s no funding available. And there usually is, because there’s nothing the universities like more than having their researchers publish in high-impact journals.

The solution is to turn such publications into black marks in one’s CV. Currently they’re not. Take for example the bullshit paper I was complaining about. It has been published in Science Advances. The authors got what they wanted, a high-impact journal in their CVs, their careers will benefit a lot from it. It doesn’t matter that everyone knows the paper is bullshit, because the evaluation is objective, all that matters is the impact factor of the journal. And that won’t be hurt either, the paper will get a lot of citations showing that it is wrong.

How to do that? I think one way is to go away from objective metrics of scientific research, and towards subjective evaluations by experts. Sure, there’s a lot of danger of corruption in the latter system, which is one of the reasons objective metrics are so popular, but I think it’s better to risk hiring somebody’s friends than for sure hiring professional bullshitters. And corruption can be mitigated by requiring the opinion of independent experts.

Another way is to boycott bullshit high-impact journals like Nature Communications or Science Advances. Don’t send them your research, don’t referee for them, don’t cite their bullshit papers. Write serious papers, and send them to Quantum instead. Ironically enough, this would make Quantum win in the impact factor game; currently it has a high impact factor, but it is not as high as those journals, and as such it is not worth as much to one’s career as the bullshit journals. We need it to have a high impact factor, but without turning it into a bullshit high-impact journal.

Another way is to make more of Quantum. Its model is eminently scalable, if a constant fraction of the people submitting papers become editors, it can easily grow to dominate all of quant-ph. But that’s a tiny fraction of scientific research, and the other classes of arXiv need their own journals to save them from greedy publishers.

Posted in Uncategorised | Comments Off on The open access movement has failed

New science blog!

A friend of mine, Miguel Navascués, has just started a new science blog, which I’m happy to advertise here: https://sciencecommunicationexplained.blogspot.com/

His inaugural post, written in his signature incendiary style, is about a subject that has personally pissed him off: another bullshit story in Quanta, this time claiming that his well-known paper on experimentally disproving real quantum mechanics had been overturned. Well, I’m not surprised. After they fell for the wormhole nonsense I don’t expect anything from them anymore. On the contrary, I’m pleasantly surprised they didn’t fall for the nonlocality without entanglement bullshit. Instead, this dubious honour goes to New Scientist.

Still, it’s a clear case of journalistic malpractice. They dutifully interviewed independent experts and some of the original authors for the story. But they didn’t include anything they said about whether the article had in fact been overturned, which is kind of the crucial point. Presumably because they were unanimous that it hadn’t, and this is not the story the journalist (Daniel Garisto) wanted to tell. You see, a story about people proving eachother wrong is much more exciting than reality, which is science being built by results adding to eachother. Actual overturnings are quite rare, as they should, because papers that can be overturned shouldn’t be published in the first place.

The story’s fundamental problem is that it never clarifies what is “real quantum mechanics”. It mentions that Stückelberg developed a real-valued quantum mechanics in 1960; it can reproduce complex quantum mechanics exactly, and hence it is not falsifiable. It never mentions this, of course, because then it would be obvious that this is not the theory that was falsified in 2021. What was falsified was Wotters’ real quantum mechanics, that was introduced in 1990. Bizarrely the story even interviews Wotters, but never mentions this! It introduces him as “a quantum information theorist at Williams College”. Because again if it had mentioned that, then the article could hardly be overturned by a new real quantum mechanics that was introduced in 2025.

All I’m saying is well-known by everybody involved, who doubtless told it to the journalist, who doubtless couldn’t care less.

This is only the superficial, names-and-dates part of the story. To get to the substance, head over to Miguel’s blog, which explains it in detail. He promises to do this not only for this story, but every single time Quanta writes about quantum information or quantum foundations. Let’s hope his stamina lasts!

Posted in Uncategorised | Comments Off on New science blog!

My complex crusade

A couple of years ago I was complaining about the lack of direct support for SDPs with complex numbers, and documenting my efforts to get it working on SeDuMi via YALMIP. That worked, but in the meantime I’ve stopped using MATLAB (and thus YALMIP), and started solving SDPs in Julia via JuMP.

No problem, I did once, I can do it again. Funnily enough, the work was pretty much the complement of what I did for YALMIP: YALMIP already had a complex interface for SeDuMi, it was just disabled. JuMP didn’t have one, I had to write it from scratch. On the other hand, I had to write the dualization support for YALMIP, but not JuMP, because there we get it for free once the interface is working. The codebases couldn’t be more different: not only Julia is a much better programming language than MATLAB, but also YALMIP is a one-man project which has been growing by accretion for decades4. In contrast, JuMP is a much newer project, and has been written by many hands. I has a sane, modular design by necessity.

That was done, but it was pointless, as SeDuMi is just not a competitive solver. To get a real benefit I’d need to add complex support to the best solvers available. My favourite one is Hypatia, which was in a similar situation: the only thing lacking was the JuMP interface. I contacted the devs, and they quickly wrote the interface themselves. Amazing! But that was also the end of the line: to the best of my knowledge no other solver could handle complex numbers natively.

I thought maybe I can hack the solver myself, as the algorithm is really simple. I looked a bit around, and found the perfect target: COSMO, an ADMM solver written in Julia. It was harder than I thought because COSMO fiddled directly with LAPACK, and that requires dealing with a FORTRAN interface from the 70s. Nevertheless I did it, and my PR was accepted. Luckily the interface came for free, as COSMO is written to work with JuMP directly. But it was also a bit pointless, because it turns out COSMO is not really good either, and it’s mostly abandoned.

I decided to go for the best ADMM solver available: SCS. The problem is, it’s written in C2. But I had a strong motivation to do it: I was working on a paper that needed to solve some huge complex SDPs that only SCS could handle. Getting a speedup there was really worth it. And it turned out to be much easier than I thought: the codebase is rather clean and well-organized. The only difficulty was dealing with LAPACK, but I had already learned how to do it from my work with COSMO. It just took ages for my PR to be accepted, so long that my paper was already done by then, and I had no intention of running the calculations again.

With the solver done, I went to work on the JuMP interface, which was just released today. As part of my evil plan to get people to switch to Julia, I didn’t write the YALMIP interface. I mean, writing MATLAB code for something that I’ll never use personally? No way. In any case, the new version of SCS showed a 4x speedup on an artificial problem designed to display the advantage, and a 75% speedup on a real complex SDP I had lying around.

Is that the end of my crusade? I really wanted to add support to a second-order solver, as these are faster and more precise (at the cost of not being able to handle the largest problems). The best one available is MOSEK. But it’s a closed source solver owned by a private, for-profit corporation, so nope. I actually tried to contribute a bug report to them once, and as a result I got personally insulted by MOSEK’s CEO.

An interesting possibility is Clarabel. It is written in Julia, so it should be really easy to hack. And in principle it should be much faster than Hypatia, as it uses the specialized algorithm for symmetric cones, whereas Hypatia uses the generic one that can also handle non-symmetric cones. But my benchmarks showed that it is not competitive with Hypatia, so until that changes there is no point, and I did enough pointless work.

Posted in Uncategorised | 4 Comments

Separable states cannot violate a Bell inequality

An old friend of mine, Jacques Pienaar, wrote to me last Friday asking whether this paper by Wang et al. is bullshit, or is he going crazy. Don’t worry, Jacques, you’re fine. The paper is bullshit.

It claims to experimentally demonstrate the violation of a Bell inequality using unentangled photons. Which is of course impossible. It’s a simple and well-known theorem that separable states cannot violate a Bell inequality. Let me prove it here again for reference: let $\rho_{AB} = \sum_\lambda p_\lambda \rho^A_\lambda \otimes \rho^B_\lambda$ be a separable state shared between Alice and Bob, who measure it using POVMs $\{A^a_x\}_{a,x}$ and $\{B^b_y\}_{b,y}$. Then the conditional probabilities they observe are
\begin{align*}
p(ab|xy) &= \tr[\rho_{AB} (A^a_x \otimes B^b_y)] \\
&= \sum_\lambda p_\lambda \tr(A^a_x \rho^A_\lambda) \tr(B^b_y \rho^B_\lambda) \\
&= \sum_\lambda p_\lambda p(a|x \lambda) p(b|y \lambda),
\end{align*} so we directly obtain a local hidden variable model for them, with hidden variables $\lambda$ distributed according to $p_\lambda$, with response functions $p(a|x\lambda)$ and $p(b|y \lambda)$. Therefore they cannot violate any Bell inequality.

The authors certainly know this. Well, I know Mario Krenn and Anton Zeilinger personally, and I know that they know. The others I can safely presume. This means that the paper is not merely mistaken. It is bullshit.

But what have they done?, you ask. How have they obtained the result they claim? Honestly, I don’t know, and it’s not my problem. When you contradict a well-known theorem you’re the one who has to explain why the theorem is wrong, or why it doesn’t apply to your situation. They don’t explain it. I read the paper, and all they have to say on this subject is the first sentence of the abstract “Violation of local realism via Bell inequality […] is viewed to be intimately linked with quantum entanglement”. They also talk about “the ninety-year endeavor in the violations of local realism with entangled particles.” Apparently it’s not a theorem, just an opinion? Or tradition?

So I’m perfectly justified in washing my hands. I can’t contain my curiosity, though. Is the quantum state they used actually entangled? Or is the violation of a Bell inequality just fake? The state does seem to be separable, so there’s only option left: there is no violation. Their setup is inherently probabilistic, they use four photon sources, which can in total generate either 0, 2, 4, 6, or 8 photons. They postselect on the detection of four photons. Well, well. It’s well-known that you can fake a violation of a Bell inequality via post-selection, if the detection efficiency is $\le 2/3$. What is their efficiency? They don’t reveal. But they do seem to be aware that their experiment is not halal, as they write “We expect that tailored loopholes and local hidden variable to the work reported here can be identified.”

UPDATE: Wharton and Price just uploaded a comment to the arXiv, confirming that the Bell violation is faked through postselection.

UPDATE 2: Now Cieśliński et al. uploaded another comment to the arXiv. They agree with Wharton and Price’s analysis, and additionally show that one can make their experimental setup produce a real Bell violation by adding on/off switches to optionally block the pumping field going to the second downconverters. But then they become part of the measurement device, not state preparation, and the state at this point is entangled.

I am bothered by the language of this comment. Which does not even call itself a comment. Despite refuting everything about the paper, they call it “brilliant” and “outstanding”. They also can’t even bring themselves to directly say the obvious, that a Bell violation with separable states is impossible. They write that “Moreover, we provide an analysis which shows that the statement in their title, about unentangled photons, cannot be upheld.” They also repeatedly talk about the “conjecture” by Wang et al. that they violated a Bell inequality. That’s ridiculous, they didn’t “conjecture” anything, they directly and repeatedly claim a Bell violation. I hope they grow some self-respect and that in the next version of the comment they use proper scientific language. UPDATE 4: I’m glad to report that the updated version of the comment was in fact soberly written.

UPDATE 3: Sabine Hossenfelder made a YouTube video about the paper. I find it mysterious why she should care about it, since she’s a superdeterminist. In any case, the video is embarrassing. Doubly so because it will have orders of magnitude larger audience than the comments. She gives the paper a “0 out of 10 in the bullshit meter”, dismisses the correct explanation of postselection by demonstrating she has no idea what postselection is, and concludes with “of course there’s always the possibility that I just don’t understand it”. Indeed, Dr. Hossenfelder, on this point you are correct.

Posted in Uncategorised | 2 Comments

Recovering solutions from non-commutative polynomial optimization problems

If you have used the NPA hierarchy to bound a Tsirelson bound, and want to recover a state and projectors that reproduce the computed expectation values, life is easy. The authors provide a practical method to do so, just compute a projector onto the span of the appropriate vectors. Now if you’re using its generalization, the PNA hierarchy, and want to recover a state and operators, you’re out of luck. The authors only cared about the case where full convergence had been achieved3, i.e., when the operator constraints they wanted to impose like $A \ge 0$ or $[A,B] = 0$ were respected. They didn’t use a nice little result by Helton, which implies that as long as you’re using full levels of the hierarchy2 you can always recover a solution respecting all the moment constraints, which are stuff like $\mean{A} \ge 0$ or $\mean{A^2 B} = 2$. This in turn implies that if you only have moment constraints, no operator constraints, then the hierarchy always converges at a finite level! This is the only sense in which non-commutative polynomial optimization is simpler than the commutative case, so it is a pity to lose it.

In any case, going from Helton’s proof to a practical method to recover a solution requires shaving yaks. Therefore, I decided to write it up in a blog post, to help those in a similar predicament, which most likely include future me after I forget how to do it.

For concreteness, suppose we have a non-commutative polynomial optimization problem with a single3 Hermitian operator $A$. Suppose we did a complete level 2, constructing the moment matrix associated to the sequences $(\id, A, A^2)$ with whatever constraints you want (they don’t matter), and solved the SDP, obtaining a 3×3 Hermitian matrix4
\[ M = \begin{pmatrix} \mean{\id} & \mean{A} & \mean{A^2} \\
& \mean{A^2} & \mean{A^3} \\
& & \mean{A^4} \end{pmatrix} \]Now we want to recover the solution, i.e., we want to reconstruct a state $\ket{\psi}$ and operator $A$ such that e.g. $\bra{\psi}A^3\ket{\psi} = \mean{A^3}$.

The first step is to construct a matrix $K$ such that $M = K^\dagger K$. This can always be done since $M$ is positive semidefinite, one can for example take the square root of $M$. That’s a terrible idea, though, because $M$ is usually rank-deficient, and using the square root will generate solutions with unnecessarily large dimension. To get the smallest possible solution we compute $M$’s eigendecomposition $M = \sum_i \lambda_i \ketbra{m_i}{m_i}$, and define $K = \sum_{i;\lambda_i > 0} \sqrt{\lambda_i}\ketbra{i}{m_i}$. Of course, numerically speaking $\lambda_i > 0$ is nonsense, you’ll have to choose a threshold for zero that is appropriate for your numerical precision.

If we label each column of $K$ with an operator sequence, i.e., $K = (\ket{\id}\ \ket{A}\ \ket{A^2})$, then their inner products match the elements of the moment matrix, i.e., $\langle \id | A^2 \rangle = \langle A | A \rangle = \mean{A^2}$. This means that we can take $\ket{\psi} = \ket{\id}$, and our task reduces to constructing an operator $A$ such that
\begin{gather*}
A\ket{\id} = \ket{A} \\
A\ket{A} = \ket{A^2} \\
A = A^\dagger
\end{gather*}This is clearly a linear system, but not in a convenient form. The most irritating part is the last line, which doesn’t even look linear. We can get rid of it by substituting it in the first two lines and taking the adjoint, which gives us
\begin{gather*}
A\ket{\id} = \ket{A} \\
A\ket{A} = \ket{A^2} \\
\bra{\id}A = \bra{A} \\
\bra{A}A = \bra{A^2}
\end{gather*}To simplify things further, we define the matrices $S_A = (\ket{\id}\ \ket{A})$ and $L_A = (\ket{A}\ \ket{A^2})$ to get
\begin{gather*}
A S_A = L_A \\
S_A^\dagger A = L_A^\dagger
\end{gather*}which is nicer but not quite solvable. To turn this into a single equation I used my favourite isomorphism, the Choi-Jamiołkowski. Usually it’s used to represent superoperators as matrices, but it can also be used one level down to represent matrices as vectors. If $X$ is a $m \times n$ matrix, its Choi representation is
\[ |X\rangle\rangle = \id_n \otimes X |\id_n\rangle\rangle,\] where
\[ |\id_n\rangle\rangle = \sum_{i=0}^{n-1}|ii\rangle. \] This is the same thing as the vec function in col-major programming languages. We also need the identity
\[ |XY\rangle\rangle = \id \otimes X |Y\rangle\rangle = Y^T \otimes \id |X\rangle\rangle \] with which we can turn our equations into
\[ \begin{pmatrix} S_A^T \otimes \id_m \\
\id_m \otimes S_A^\dagger \end{pmatrix} |A\rangle\rangle = \begin{pmatrix} |L_A\rangle\rangle \\
|L_A^\dagger\rangle\rangle \end{pmatrix} \]where $m$ is the number of rows of $S_A$. Now the linear system has the form $Ax = b$ that any programming language can handle. In Julia for instance you just do A \ b .

It’s important to emphasize that a solution is only guaranteed to exist if the vectors come from a moment matrix coming from a full level of the hierarchy. And indeed we can find a counterexample when this is not the case. If we were dealing instead with
\[ M = \begin{pmatrix} \mean{\id} & \mean{A} & \mean{B} & \mean{BA} \\
& \mean{A^2} & \mean{AB} & \mean{ABA} \\
& & \mean{B^2} & \mean{B^2A} \\
& & & \mean{AB^2A} \end{pmatrix} \] then a possible numerical solution for it is
\[ M = \begin{pmatrix} 1 & 1 & 0 & 0 \\
& 1 & 0 & 0 \\
& & 1 & 0 \\
& & & 1 \end{pmatrix} \] This solution implies that $\ket{\id} = \ket{A}$, but if we apply $B$ to both sides of the equation we get $\ket{B} = \ket{BA}$, which is a contradiction, as the solution also implies that $\ket{B}$ and $\ket{BA}$ are orthogonal.

I also want to show how to do it for the case of variables that are not necessarily Hermitian. In this case a full level of the hierarchy needs all the operators and their conjugates, so even level 2 is annoying to write down. I’ll do level 1 instead:
\[ M = \begin{pmatrix} \mean{\id} & \mean{A} & \mean{A^\dagger} \\
& \mean{A^\dagger A} & \mean{{A^\dagger}^2} \\
& & \mean{AA^\dagger} \end{pmatrix} \] The fact that this is a moment matrix implies that $\mean{A} = \overline{\mean{A^\dagger}}$, which is crucial for a solution to exist. As before we construct $K$ such that $M = K^\dagger K$, and label its columns with the operator sequences $K = (\ket{\id}\ \ket{A}\ \ket{A^\dagger})$. The equations we need to respect are
\begin{gather*}
A\ket{\id} = \ket{A} \\
A^\dagger \ket{\id} = \ket{A^\dagger}
\end{gather*} or more conveniently
\begin{gather*}
A\ket{\id} = \ket{A} \\
\bra{\id} A = \bra{A^\dagger}
\end{gather*} We use again the Choi-Jamiołkowski isomorphism to turn this into a single equation
\[ \begin{pmatrix} \ket{\id}^T \otimes \id_m \\
\id_m \otimes \bra{\id} \end{pmatrix} |A\rangle\rangle = \begin{pmatrix} \ket{A} \\
\overline{\ket{A^\dagger}} \end{pmatrix} \]and we’re done.

Posted in Uncategorised | Comments Off on Recovering solutions from non-commutative polynomial optimization problems