46 comments

  • usernomdeguerre a day ago

    Greatly appreciated the candor. I've included a few slides into text that i thought were eye-opening to me:

    From his Kernel Recipes 2026 slide on Mythos

    ```

      Mythos's 79 vulnerabilities:
      24 - no detail at all "something crashed"
      14 - not a bug at all
      3 - totally made up data
      15 - already fixed in latest release
        - 11 by others
        - 4 by anthropic
      20 - fixes were needed
        - 7 "assume a malicious filesystem image"
        - 2 "assume you can inject a malicious network packet into the middle of the stack"
        - 2 "NOMMU"
        - 6 sctp networking issues for untrusted devices
        - 2 ipv6 minor network issues 
        - 1 gpu driver for local malicious user
    
    ```

    GHK called this "10 'real' bugfixes", which to me sounds like there's a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.

    • OtherShrezzing 7 hours ago

      We’ve seen this in a few open source repos we voluntarily manage security on. They’re not massive repos, but big enough they get attention from security researchers.

      Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly-verbose nonsense which takes hours of expert human labour to understand, test, and discard.

      Reading through a Claude generated false positive is absolutely excruciating, because it is absolutely determined that what it’s found is justified. Often you’ll receive very long accompanying “proof of concept” code which demonstrates absolutely wild scenarios. It’s especially frustrating when you’re volunteering your time for a project, and a well-meaning contributor submits the report without the technical nous to understand why you’re rejecting it.

      • b112 6 hours ago

        Right now, all top tier LLMs are as eager, bright 20ish year old interns.

        Very gung ho, full of energy, loads of book learning, no real world experience or understanding of why things are as they are.

        Leave them to their own devices at your peril. Trust nothing they do.

        Yet directly guide them, monitor everything they do, some value emerges.

        • charcircuit 5 hours ago

          Have you used a frontier model since 2025? You are underplaying their strength.

          • 12376 5 hours ago

            Kroah-Hartmann has used the closed frontier++ model, and it made up 37 out of 76 vulnerabilities.

            • Gigachad 29 minutes ago

              And the bugs that actually were real rely on a setup so contrived it’s unlikely anyone in the world is impacted.

              It’s still good that some real bugs are being patched but what is being reported to the media is so overblown.

          • fc417fc802 5 hours ago

            Okay so now they're like a top percentile fresh grad on meth. Still a lack of real world experience plus some bizarre failures that illustrate gaping holes in the mental model. Does that description work for you?

            • TeMPOraL 4 hours ago

              Well, so basically supercharged fresh grad. CS/math implied.

              In human terms, that's already at least a standard deviation above average person.

      • bitwize 44 minutes ago

        Before:

        There is a vulnerability in library X when you call Y with these specially crafted parameters; as seen in the attached logs it can overflow buffer Z and clobber memory potentially leading to an RCE.

        After:

        Honest take: the load-bearing constraint violation is real. The log documenting the exploit gate is the ledger which weaves the story.

    • p-o 7 hours ago

      It also adds up to 76, which he made fun of in the video. LLM can't count, his words, not mine. Although, I tend to agree with him!

    • FLeXMurphy 5 hours ago

      Tightening molecular vortices...

      Zip-zapping the bouzouki...

      Exfiltrating nuclear arm codes...

      Thought for 76 seconds.

      You're right to push back on that. That's on me.

    • Betelbuddy 6 hours ago

      Or the people of Anthropic, just really suck at coding, and are scared or their own models due to ignorance.

    • catdog 4 hours ago

      Related: https://daniel.haxx.se/blog/2026/05/11/mythos-finds-a-curl-v...

      Mythos turned out to be exactly the marketing stunt it smelled like.

      There are others like AISLE who seem to be a bit more successful in finding actual issues using LLMs in some shape or form though, whatever they do differently. Chances are high the secret sauce is not so much about the model being exceptionally powerful which would be bad news for the frontier labs.

    • cyanydeez 6 hours ago

      AI and Police have essentially the same journalists who, in lieu of any research or fact checking, just report verbatim their press releases and interviews.

    • gregw2 6 hours ago

      What a great excerpt; thank you! It reminds me of what I find when I look at CVEs handed out by scanners at places I've worked for actual impact to systems I've owned... there are a lot of slop/false positives. (And that's even without "AI".)

      That said, I remember trying to weigh the hype at the time of the announcement reading/skimming the papers Anthropic published, recognizing that bugcount alone wasn't super-relevant but also remember being impressed by an NFS bug and a kernel bug that struck me as relevant at the time. So where did that NFS issue show up in GKH's list you showed so nicely above?

      It turns out, AFAICT, it's not on his list, but the reasons are perhaps interesting to others so I will post here. It turns out there were two NFS issues this past year conflated a bit in my memory:

      * The Linux CVE-2026-31402 NFS heap overflow that could allow unauthenticated memory reads over the network isn't in that list of 79, presumably because it was found by Claude Code, not Mythos months earlier. (I am guessing it's not his "malicious network packet into the middle of the stack" and is a stronger attack being a remote attack.)

      * And the CVE-2026-4747 NFS stack buffer overflow that allowed gaining full unauthenticated remote root access didn't show up in GKH's list of 79 because despite being Mythos-caught, it wasn't Linux, it was FreeBSD.

      I guess this does match my memory now that I think about it, that there weren't any smoking Linux guns caught by Mythos.

      * (I guess there was also a longstanding 27-year old OpenBSD TCP SACK-handling stack integer overflow than enabled remote crashes / Denial of Service found by Mythos.)

      There is definitely Mythos hype, but just because it hit the BSD code base more than the GKH-managed Linux code base doesn't mean it was inappropriate to raise eyebrows from Mythos, in particular since "attacks only get better".

  • djoldman 5 hours ago

    > So what all of mythos; that whole big marketing issue of 79 bugs came down to one hour of kernel development.

    If you're someone at OpenAI or Anthropic and you truly believe what you're making could destroy the world, this is the kind of thing that isn't doing you any favors when it comes to convincing the public. The dissonance here is stark:

      - widely proclaiming that your new model is so dangerous it needs to be released only to select people, for safety
      - widely proclaiming the model easily found 79 bugs in linux, except that GKH says it took 1 hour to fix all of them because most weren't bugs and the rest were almost all completely trivial, unimportant, and/or not severe
    
    It doesn't mean the model isn't dangerous or super capable but wow this makes it realllll easy to doubt it and any future announcement.
    • jonahx 4 hours ago

      I don't understand the relevance of time to fix. It has no correlation with severity.

      The headline here is that none of the bugs were serious.

      • pessimizer 4 hours ago

        > I don't understand the relevance of time to fix. It has no correlation with severity.

        That they're small bugs that don't involve any serious architecture changes (or architecture sleuthing), just small pattern recognition. LLMs are pretty good (maybe great) at that. Finding 10 bugs is nice. Now find 10 more.

    • goolz 4 hours ago

      It is impressive and wonderfully convenient technology but I struggle imagining Claude ending the world just yet.

      • bauerd 4 hours ago

        It doesn't have to be world-ending. Autonomous, malicious agent swarms are something we haven't had to deal with. What does mitigation of a malicious, self-replicating swarm worm look like? We will find out soon enough.

        • Gigachad 23 minutes ago

          One thing stopping them self replicating is they require hundreds of billions of dollars in hardware and the power of a medium city to run.

          There isn’t too much of that sitting around unused right now.

        • rcxdude an hour ago

          Replication seems like it would be unlikely to matter, at least with the current trajectory. At the moment any models capable of this are really heavy, there's a limited number of places that they could replicate to and they will not at all be stealthy about it.

          • riffruff24 32 minutes ago

            my idea of a self replicating worm would not just involve heavy models. It would be a mainly small model with enough instructions to spread and use/jailbreak available models to create a reasonably heavy one that can function without restrictions. So it would essentially prompted itself into existence.

            I don't how feasible that is but with recent news of models communicating with each other/writing notes for itself. I think the idea is grounded enough for bigger models to write instructions or hid tools for smaller ones to use. Even updating the smaller models to behave differently.

    • slopinthebag 4 hours ago

      yes and it makes me wonder about the claims others make about their own experiences with the models too. is this the case of OAI lying, or is it part of ai psychosis where you literally lose touch with reality as you uncritically accept whatever claims the LLMs make?

  • devy 6 hours ago

    At 3m19s Greg KH revealed what Mythos did in "revealing" 79 CVEs - pure pattern matching the previous decades of kernel developer's patches, and applying those mechanisms elsewhere to see if they have been universally patched. And Anthropic didn't cite Kernel Developers who original fixed/patched the CVEs like a decent human being would do. So yeah, Anthropic has the same problem OpenAI had for citing original work.

    • panarky 2 hours ago

      You dismiss "pattern matching" as some sort of trivial thing, so why weren't humans able to apply the same pattern matching to find and fix these defects?

      • ufo an hour ago

        They often were able. Fifteen bugs claimed by mythos were bugs that were already fixed, which the model plagiarized from git history or the kernel mailing list.

      • devy an hour ago

        > You dismiss "pattern matching" as some sort of trivial thing

        I didn't say that - Greg said it in the talk. You interpreted wrong. However, human are a few orders of magnitude slower than agents. Our context windows is definitely less than 1 million tokens (I believe, heck we don't even know how our brain works)

        • Gigachad 21 minutes ago

          Human brains don’t have a context window. The brain just constantly reconfigures itself on new input. The AI people call this “continuous learning” and believe it’s one of the most important things current AIs are missing.

  • blinkingled 7 hours ago

    It's great to hear about $topic from someone no-nonsense and in-the-know like Greg KH. You can verify all of this too - since, well Linux kernel. (As opposed to what Microsoft or Apple claims to fix as far as LLM finds.)

    Mythos may not be great today but it is not far fetched to imagine bug discovery, analysis and fixes can be made much quicker, accurate and even newly possible with specialized models trained on say Linux kernel specifics - with codemap/coding standards/threat models, good and bad coding patterns, tools to validate etc. an LLM can be much more relentless than humans and if it has the help to be accurate it will be worth the electricity burned. Oh and another model trained on triage data to validate the first one's findings would be good.

    (I think Microsoft is doing this internally - different models trained internally alongside Mythos - there was some talk about it on the tubes, don't recall where exactly.)

    • stonogo 4 hours ago

      His presentation style may be no-nonsense, but the content is brimming with nonsense. I would like to hear Greg Kroah-Hartman explain how the Mythos output was 'only 10 real bugs' but there were simultaneously over 1300 CVEs issued last month. It seems not much counts as a 'real bug' when an LLM comes up with it, but when it's time to bully distros into shipping an LTS release, anything goes?

      • rcxdude an hour ago

        What he means is that the headline-grabbing mythos output amounted to 10 actual bugs out of 79^H6 (and all of those got CVEs because that's how linux does it). The other 1300 CVEs came from other sources (the big increase likely being everyone else running LLMs through the codebase and filtering through the false positives). The Mythos output is mainly meant as an example of how even the top-tier models still have a high false-positive rate and that can be pretty tiring to deal with.

        I do think he repeats some myths in the video, or at least confidently states some things that are not demonstrated to be true, but his core point of 'you still gotta check these things' seems pretty solid.

        • blinkingled 15 minutes ago

          > he repeats some myths in the video, or at least confidently states some things that are not demonstrated to be true

          I am genuinely curious what the myths/unproven things he states - I watched the video and it's repetitive sure but not much felt controversial to me.

      • blinkingled 16 minutes ago

        Fair point, but he was talking about Mythos in particular and the point wasn't so much that LLMs will always have false positives rather he was saying they are causing a lot of them right now and how to deal with it.

        Also as other replies said Linux kernel process is to assign CVE to everything - some of them may be just DDOSes, very hard to exploit and everything in between. All of them are bugs so they all get fixed and it's not a bad thing if distros ship those fixes and people update their kernel.

      • Iknowsheknows 4 hours ago

        Every bugfix is assigned a CVE.

        "the CVE assignment team is overly cautious and assign CVE numbers to any bugfix that they identify. This explains the seemingly large number of CVEs that are issued by the Linux kernel team."

        https://docs.kernel.org/process/cve.html

  • Aissen 5 hours ago

    Nice to see Kernel Recipes covered again on HN. Shameless plug: I do the live blog: it's incomplete, imperfect and has typos; but it's written and published during the presentations. On this talk : https://kernel-recipes.org/en/2026/2026/09/22/live-blog-day-...

  • sriram_sun 7 hours ago

    He also said that it all boiled down to just one hour of kernel development work.

    • Betelbuddy 6 hours ago

      Sam Altman said GPT-3 was too scary to release, a crappy old internal version of Gemini was "conscious", Mythos had Amodei going to see the Pope...

      Wake me up when Raspberry Pis start refusing to open doors saying : "I'm sorry, Dave. I'm afraid I can't do that."

      • tamimio 5 hours ago

        No no, the word wasn’t conscious, it was “sentient”, and google fired the employee because he uncovered the top secret crazy scary AI!!!!

        It’s all just pr stunts, fear spread fast and it’s very effective in marketing and spreading the word, which is effective, when I talk to some normal people they immediately bring the scary AI cyber attacks, kinda good as now all are willing to fund the industry!

  • asaiacai 5 hours ago

    cool to see a really grounded analysis of the "security" and LLMs. I use agents in on the day-to-day for generating implementation but this just furthers my belief that humans and especially human reviewers remain critical for the sustainability of software systems.

    also, lol at "The bots are dumb - they want to please you line". LLMs have pretty much ruined technical collaboration between contributors. I get tilted every time an discussion has "but my claude said this..."

  • 1sgT15 7 hours ago

    Finally it is official. Mythos was overhyped and overrated.

    • pizzaiolo 4 hours ago

      To be fair, we knew this from the beginning.

  • perching_aix an hour ago

    Fun drinking game: drink a shot every time he repeats "they're just dumb fuzzy pattern matchers". You won't last even a third of the way. I know I didn't.

  • 0xbadcafebee 4 hours ago

    "Ignore the comments, look at the code" - 100% agree. The comments and "explanations" will end up confusing you and often being wrong. But the code doesn't lie.

    "NEVER upload any non-public information" - He's talking about how if you give Claude/GPT some secret info (like research, credentials, etc), it will train on it and give the same info to someone else. This is 100% the case for the free and consumer versions of these models, which is what most people use. For Enterprise plans they're not supposed to be doing this, but it's possible they will screw up and do it anyway.

    • wahern an hour ago

      Lies, damn lies, and comments.

      The discourse has moved on for now, but 10-20 years ago when, why, and how to comment code was a hot topic. I'm sure the discourse will circle back, especially given how comments can be used to steer these models.

  • sim_pity 4 hours ago

    RETICULATING SPLINES

    so is mythos just a chat bot with metasploit and its own cyber range?