The Era of Software Quality, or the Era of Ostriches?
Humans are bad at writing secure code, and GNOME developers are no exception. GNOME is primarily written using unsafe programming languages where simple mistakes in our code lead to devastating consequences for our users, and we make these mistakes all the time. No matter how much we try, GNOME developers will fail write secure code when using unsafe languages like C, C++, or Vala: it’s just too hard for even experienced developers to do properly.
The above paragraph is taken from the abstracts of my GUADEC 2024 and 2025 talks. At the time, I thought failure was inevitable: we humans were so bad at writing software that we had no chance to do it properly, and I certainly would not have trusted an AI to do better than a human. But the landscape today is completely different than last year. AI has improved considerably, and offers a magic fairy wand solution to this problem: we can now simply ask a language model to look for vulnerabilities in our software. They are quite good at this.
There is zero hope of maintaining quality software in 2026 without AI vulnerability scanning. Any claims to the contrary are unserious and delusional. The tremendous quantity of bugs found in our best-maintained projects, like GLib and fwupd, should speak for itself. Failure to scan our projects is an unfair disservice to our users. If we don’t find the vulnerabilities by scanning projects ourselves, attackers certainly will, because the Linux user base has increased to the point that Linux users are finally numerous enough to be worth targeting. Meanwhile, AI has made it easier than ever to build working exploits, which was previously unheard of.
Already resolved all the detectable vulnerabilities? Then ask the AI to look for non-security bugs as well, to further improve quality. GNOME code is generally much better than it used to be, but there remains considerable room for improvement. For the first time in history, we now have the opportunity to improve software quality to a degree that was never realistic before.
Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the quality of AI-generated vulnerability reports has drastically improved. That is not to say that we no longer have problems with bad vulnerability reports, but in general, nowadays most of them are pretty good. (Daniel Stenberg reports the same pattern for curl.)
AI-generated vulnerability reports have nevertheless introduced many undesirable impacts on GNOME maintainers. They are usually annoyingly verbose and unnecessarily detailed. They often exaggerate the severity of the problem, or make misleading or irrelevant claims. They are occasionally incorrect. Sometimes they include outright fabricated data, such as fake stack traces (which is not the norm, but sadly also not uncommon). A good human reviewer will notice and resolve most of the above problems before creating a bug report on your issue tracker, but often problems are reported by inexperienced humans who do not actually know what they are looking at and simply copy/paste everything blindly. Even when the generated issue report is good and avoids all of the above problems (which is rare), good vulnerability reports in sufficiently high quantity can still overwhelm volunteer maintainers. And even if reporters submit a merge request to resolve the problem so maintainers don’t have to (which is also rare), reviewing those merge requests is itself more unwelcome work for overworked maintainers.
That all is to say: I understand the pain caused by the current wave of AI-generated issue reports. Nevertheless, they are essential and unavoidable. We have to learn to accept and deal with them, not stick our heads in the sand and ignore them.
Some GNOME maintainers have adopted a policy prohibiting AI-generated content in issue reports. Do not do this. Nowadays, the overwhelming majority of vulnerability reports are AI-generated. Projects that choose to ban AI-generated content in issue reports might as well ban all vulnerability reports; the effect will be approximately the same.
I propose the following:
- GNOME maintainers should rewrite their AI contribution policies to permit AI-generated vulnerability reports, as I previously requested four months ago.
- Projects that continue to prohibit AI-generated vulnerability reports are no longer suitable dependencies for GNOME, and should be developed someplace other than GNOME GitLab.
We don’t have to tolerate bad issue reports, but AI use alone should not be disqualifying.
Shouldn’t humans rewrite AI-generated bug reports?
When I complain that maintainers should allow AI-generated vulnerability reports, the most common counterargument is that humans should read the AI’s report, understand it, and rewrite the entire thing to remove all AI-generated content. Some bug reporters actually voluntarily do this, but this is rare.
Vulnerability reporting is a public service, not an obligation. If you ask a reporter to do any amount of extra work, they might be willing to do so, but it’s much more likely that they will either stop looking at your project and move on to something else, or continue looking at your project and publish the vulnerability reports someplace other than your issue tracker.
Rewriting issue reports also does not scale. Let’s say you use AI to find 100 security bugs in a GNOME project, a number consistent with the results of actual scans (read on). Would you really spend months rewriting those bug reports before submitting them to upstream? Validating the AI’s claims, upstreaming the issue reports, and submitting merge requests is already a lot of work. Not many people would be willing to additionally rewrite all the issue reports. That’s more work than everything else combined, and is unrealistic.
Even with just a small number of bugs, I would hesitate to spend much time rewriting an issue report because I have many other tasks I would rather spend my time on. At best, I might prepare a quick summary, but it won’t be as useful as a full report.
The CVE Wave Hits GNOME
The current wave of vulnerability reports is reflected in GNOME’s CVE issuance trends:
| Year | GNOME CVEs | GNOME CVEs Excluding GIMP, Gegl, libxml2, and libxslt |
| 2021 | 21 | 14 |
| 2022 | 14 | 6 |
| 2023 | 13 | 4 |
| 2024 | 37 | 28 |
| 2025 | 97 | 49 |
| 2026 Year-to-date (2026-09-30) | 141 | 74 |
| 2026 Normalized | 188 (141 * 4 / 3) | 99 (74 * 4 / 3) |
The trend here should be pretty clear. Until recently, not many people were reporting vulnerabilities in GNOME. That has changed. We are currently dealing with an order of magnitude more CVEs than just 3 years ago. AI is not the only reason for this; GNOME maintainers have also gotten a little better at flagging issues so that I add them to security tracking. But AI is the primary cause for the increase.
(A few technical notes on this table. CVEs are classified by the year the issue was reported to GNOME, not by the year in the CVE identifier, so e.g. many CVE-2026 issues are counted in 2025. Vulnerabilities reported in 2026 which do not yet have CVEs are not counted, so you can think of the data as being accurate through roughly September 1; multiply the 2026 numbers by 4/3 to make them comparable to the prior years. I count only issues reported to GNOME Security, so any unreported CVEs do not count.)
Although there are still 3 months left in 2026, we will never have data for the rest of the year because I have ended security tracking for new issue reports and nobody else has volunteered to do that work. These CVEs exist only because I request them myself, so I expect the number of CVEs to drastically decrease going forward.
The CVE Wave Hits WebKitGTK
A similar pattern holds for WebKitGTK:
| Year | WebKitGTK CVEs |
| 2015 | 175 |
| 2016 | 57 |
| 2017 | 158 |
| 2018 | 101 |
| 2019 | 99 |
| 2020 | 38 |
| 2021 | 52 |
| 2022 | 50 |
| 2023 | 45 |
| 2024 | 38 |
| 2025 | 66 |
| 2026 Year-to-date (through WSA-2026-0006) | 305 |
CVEs are reported against the year they appeared in a WebKitGTK security advisory, not the year in the CVE ID. The large increase in 2026 is entirely due to AI analysis of Skia and ANGLE. WebKit bundles these libraries because they are not designed to be installed as system libraries, so their vulnerabilities should be counted the same as vulnerabilities in WebKit’s own code. Excluding Skia and ANGLE, there are actually only 21 other WebKitGTK CVEs so far this year, a significant decrease, but excluding CVEs in bundled code would not be fair.
There has actually been a very large increase in WebKit security fixes this year, but this has not resulted in any increase in CVEs. Apple generally creates CVEs for flaws found by external researchers, not often for flaws found by WebKit developers, so the increase in security fixes is not reflected in the total number of CVEs. Only a small fraction of WebKit vulnerabilities receive CVEs.
I had not previously noticed that the count of WebKitGTK CVEs had, until 2026, been decreasing over the past decade. I am not sure why. I also do not know how to explain the low number in 2016.
Announcing the GNOME Bug Bounty Program and Announcing the End of the GNOME Bug Bounty Program
My blog post to-do list says that I need to write a blog post announcing the creation of the GNOME Bug Bounty Program on the YesWeHack platform. Oops, too late. It’s already closed. (Once a task enters my to-do list, it can be a very long time before I get around to doing it.)
The GNOME Bug Bounty Program was generously sponsored by the Sovereign Tech Resilience program of Germany’s Sovereign Tech Agency. I’m not sure precisely when it opened, but the first vulnerability was reported on June 27, 2024, so it would have been sometime shortly before then. We accepted issue reports only for GLib, glib-networking, and libsoup, because GNOME had never operated a bug bounty program before and we did not know what to expect. Starting small had — naively — seemed like a prudent way to avoid a large quantity of issue reports. I had wanted to expand the program to cover all of GNOME, but this failed due to the overwhelming deluge in issues reported against GLib and libsoup.
I requested that the bug bounty program end because I was overwhelmed with incoming AI-generated issue reports. The final issue was reported on February 23, 2026. Here are the results:
| Year | Reports Submitted | Reports Accepted |
| 2024 | 26 | 14 |
| 2025 | 150 | 33 |
| 2026 | 122 | 24 |
| Total | 298 | 71 |
Those numbers for 2026 reflect less than two months’ worth of issue reports, so you can see why it was no longer sustainable.
After the program closed, our work was not done: there was a long backlog of reports to work though. We just last month caught up with accepting the last of the issues reported back in February, and the last bounty was finally awarded earlier today! Even with YesWeHack’s professional triagers analyzing the issue reports before I reviewed them, keeping up with such a large number of vulnerabilities was not easy for me.
At this point, all reports not accepted have been rejected. The program awarded €183,900 in bounties for 71 vulnerabilities: 45 in libsoup, 23 in GLib, and 3 in glib-networking. Award amounts varied from €500 (16 awards) to €7,500 (2 awards). The arithmetic mean award was €2,662.99.
Bug bounty programs are an exception to the rule that most AI-generated vulnerability reports are good. You can see the number of reports accepted is a small fraction of the number of reports submitted. Excluding 30 reports closed as duplicates, that leaves 197 reports rejected. Turns out, people will submit bad reports when financially incentivized to do so. The low percentage of accepted reports even understates the problem, because many of the accepted reports were actually not very good! Many accepted reports did successfully identify valid security problems (in fact, many of the rejected reports successfully identified valid security problems!), but required many rounds of revision and corrections.
Suffice to say, I have reviewed a lot of really bad AI-generated vulnerability reports. But the reports we received via the discontinued bug bounty program are not comparable to the reports received via regular GNOME issue trackers or the security bug report form. We do still occasionally receive bad vulnerability reports, but not often and not many, so it’s not a big problem anymore. When people submit AI-generated reports without hope of a financial award, those reports are generally much better.
Lessons from the Bug Bounty Program
Closing the bug bounty program because it found too many vulnerabilities is not a particularly pleasant result. That said, it was still a partial success in that it uncovered lots of bugs in libsoup and GLib.
I had hypothesized that libsoup was probably not very secure, but I never imagined just how many vulnerabilities would be discovered. To reduce the quantity of incoming issue reports and better reflect actual risk to GNOME users, I eventually removed all denial of service bugs from program scope, and then later removed SoupServer from the scope due to too many request smuggling vulnerabilities, which are HTTP request parsing bugs that pose no threat to GNOME users. Even with those changes, the libsoup vulnerability reports kept coming until I gave up. The silver lining is that libsoup is now relatively much more secure than before. Other bug reporters have been submitting AI-generated bug reports using the normal libsoup issue tracker, so fortunately the improvements to libsoup will continue despite an end to the financial awards.
I had hypothesized that GLib would be much better than libsoup. I’m not sure whether I was correct. Evaluating the severity of GLib flaws is much harder than for libsoup, since GLib vulnerability reports are generally hypothetical in nature: usually some proof of concept program calls a GLib API using valid but improbable values, then something bad happens.
A large portion of the GLib bugs were integer overflow flaws, which generally result in buffer overflow. I am now more scared of integer overflow than anything else. It’s likely that most software projects have many integer overflow problems. Fortunately, we should be able to catch most such problems by adjusting the compiler flags we use. In particular, -Wconversion or -Wint-conversion and -Wsign-compare should help here. Some GNOME projects already use -Wsign-compare, but I suspect most do not. I think few or no GNOME projects use -Wconversion or -Wint-conversion.
Resuming the bug bounty program would only be possible under substantially different conditions. What we were doing was not working well. To resume, we would need to limit the scope to projects that regularly perform their own AI vulnerability scans. We would also most likely want to pay only for functional exploits, rather than for all vulnerabilities. GNOME code is currently not good enough to continue paying for every vulnerability, and it no longer makes sense to pay bounties for issues that can be found by AI scanners.
Red Hat Scans GLib
Red Hat has contracted with AISLE Research to perform AI vulnerability scans of various GNOME projects. We received a large quantity of findings, and are only just now beginning to individually validate and report our findings to upstream. GLib is by far the hardest hit project, which I was not expecting, accounting for more than 40% of our total findings. I’m not certain why, but perhaps this is because GLib provides so many public APIs. Data passed to public APIs is potentially untrusted, so the attack surface is considerable.
Red Hat’s scan of GLib found 118 vulnerabilities. Or at least, it claimed to. However, due to the way we ran the scans, several of these are actually unnecessary duplicates of each other, which we have not fully deduplicated yet, so the number I report is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not very robust. A typelib controls how your program calls libraries; it is effectively calling convention, so it must inherently be fully trusted: a malicious typelib would be able to induce vulnerabilities even without any bugs! I would expect an AI ought to have been able to figure that out, but apparently not. These bugs are still real problems that we ought to fix, but all maintainers agree they are not security vulnerabilities, so let’s count all of them as false positives. That alone creates a 40% false positive rate. Ouch.
I don’t have more stats to share here because we are not yet done working through the issue reports. That said, I am quite pleased with the results thus far. Substantially all of the reports are high-quality. The false positive vulnerability reports are almost all due to one particular misunderstanding and can be treated as good quality non-security bug reports, which are still valuable. Expect many forthcoming CVE assignments for the other findings.
It’s rare for Linux vendors to proactively look for software vulnerabilities, rather than waiting for security researchers to report them. This was a successful experiment in proactively seeking out problems.
Humans Still Useful
In addition to the bug bounty program, the Sovereign Tech Resilience program also sponsored a security audit for GNOME, performed by Codean Labs. This resulted in many findings in various GNOME projects. Most notably, the scope of the audit extended to Flatpak and xdg-desktop-portal, resulting in critical findings.
Most of these issues could have been detected via AI scans, but I am not confident that AIs would have been able to discover the most important findings, like the two Flatpak sandbox escapes that I linked to above. Accordingly, I do not recommend relying on AI alone.
Humanity Still Desired
Although I like AI-generated issue reports, I particularly do not appreciate when I wind up interacting with a robot rather than with a human. It’s pretty obvious when your issue tracker or code review comments are written by an AI. Consider whether outsourcing your writing and your thinking to a language model is truly wise for your public image.
We even have one experienced GNOME developer who is obviously using AI to write all of his posts on GitLab. I am unsure whether he is copy/pasting all of his responses from an AI, or whether he is just a bot now. I especially do not understand the value of this.
Here is a soft proposal, intended only as a starting point for discussion and not as a serious proposal, for what my preferred AI usage policy might look like:
- Newer developers should exercise caution when using AI to write code. Your priority should be learning, and I wonder how much you are really learning when relying on the AI to do work for you.
- Do not use AI to write code comments. Currents AIs are terrible at writing comments. Most comments written by AIs should be deleted. If a comment is truly necessary, then I’d like to see it written in your own words. Presumably AIs will get better at this eventually, but as of 2026, human judgment is still required here.
- Do not use AI to write commit messages. AIs are actually probably better than humans at writing commit messages, but I would still rather hear your own thoughts on the code you are submitting.
- Certainly do not post AI-generated comments on an issue tracker or merge request as if they are your own. You’re not fooling anybody.
Maintain Perspective
Are you scared by the large numbers of recently-discovered vulnerabilities? There is no need to panic. Security bugs are just bugs, and they’re not necessarily more important than other bugs. Occasionally they are emergencies, but far more often they are boring and unexceptional. Security vulnerabilities are not even the biggest digital security threats that users face: those are surely phishing and trojans, with software security bugs a distant third place. No amount of CVE fixing will protect you from those more likely threats.
I don’t want to downplay the severity of security issues either. In fact, evaluating severity is hard. I quite often decide that a bug is not a big deal, only to be proven incorrect. Ideally, we would fix as many security issues as possible, and sooner rather than later. Lifetime issues and out of bounds writes are especially important to fix. Two years ago, I claimed that memory safety vulnerabilities were becoming less threatening, a claim that did not age well: that is surely no longer true due to the drastically increased accessibility of AI exploit generation.
Nonetheless, volunteer maintainers should not feel obligated to fix security issues or treat them as higher-priority than other bug reports. It’s certainly good to fix problems when possible, but my request is only that you do not prohibit issue reports, not that you attempt to personally resolve every security problem yourself. When I add due dates to vulnerability reports, that represents only a disclosure deadline — because issue reports should not stay confidential indefinitely — not an expectation that you fix the issue by that date. Resolving security problems in projects used by big tech companies that depend on your software without contributing back is basically free labor for said companies, and only you can decide whether that’s how you want to spend your volunteer time.
Rust
Yes, even projects written in memory safe languages like Rust still need to allow AI-generated vulnerability reports. Rust will indeed eliminate most memory safety issues (except in unsafe blocks), and you can reasonably expect a Rust project to have an order of magnitude fewer vulnerabilities than a comparable project written in C or C++ or Vala. This is amazing, but not all vulnerabilities are memory safety issues, so this is not an excuse to avoid scanning for flaws.
Although Rust mostly eliminates memory safety risk, any use of Cargo to download dependencies dramatically increases supply chain security risk. The risk of bundling a trojanized dependency arguably — I would even say probably — outweighs the benefit of eliminating memory safety flaws. This problem is inherent to any programming language package manager. Currently the best solution is to not use programming language package managers, but GNOME’s Rust code depends heavily on Cargo. Accordingly, I recommend against using Rust for writing GNOME software.
To Be Continued…
I have exhausted my thoughts on AI vulnerability reports, but there is still much to discuss regarding software quality. Next time, I will discuss additional strategies to improve GNOME quality without significantly relying on AI.


















