Open-weight AI is having its Kubernetes moment

(tobi.knaup.me)

231 points | by tknaup 6 hours ago ago

165 comments

  • ozgung an hour ago

    Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin.

    So, any solution to this “problem” must include ALL open-weight models. As far as I understand this is exactly what they intend to do. Axios article linked in the post mentions that. As in this quote:

    “The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.”

    It doesn’t say “Chinese” open-source models. Because they already know that it’s not feasible. Any regulation must cover all the models.

    Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution. If a company wants to run an open model in their own servers, they can only use approved and certified pure “American” models. This of course creates a monopoly for the big labs who are authorized to train and distribute such “open” models. A company can fine-tune the model for its own needs but of course can’t distribute the derivative model.

    I’m sure there are other solutions but all of them would be equally ugly. Also these regulations can’t be enforced to other countries easily so only Americans will be restricted.

    • satvikpendem 20 minutes ago

      It's simple, the US government will put any Chinese open model companies on the entity list which blocks any company which does business with the US from also doing business with the Chinese companies. This creates a chilling effect where even if it may be harder to tell, no US company will be able to provide or use any overt Chinese open model and won't even risk trying to go around as the punishments for trying to evade the ban are severe.

      • amarant 3 minutes ago

        But that's just the thing with open weights: you're not doing any business with company that made the model. They might publish the weights to a, say, European host, and then you download the model from Europe and and run it on your servers in America, and suddenly it's very hard to tell where the model was originally created.

      • EMIRELADERO 2 minutes ago

        Hosting and providing Chinese models wouldn't be the same as doing business with the sanctioned entities though, you don't interact with them in any capacity if you only use the weights and don't sign any contracts.

    • kloop an hour ago

      > “The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.”

      That's going to hit first amendment grounds pretty quick, the same way that software in general did.

      The modern version of the decss flag will be a character that says "I think good weights are {...weights go here...}"

      They could, however, ban any payment to a chinese entity, or any entity owned by a chinese entity for inference/ai services/etc

      • CamperBob2 23 minutes ago

        That's going to hit first amendment grounds pretty quick, the same way that software in general did

        Don't count on that. "National security" == the cheat code for the US court system that instantly bypasses any First Amendment issues.

      • isityettime an hour ago

        Uh, how big would such a flag have to be?

    • bfung 29 minutes ago

      It’s not really feasible, in my opinion.

      US Gov could make US companies comply, like have Huggingface take down models out of compliance.

      But most likely, a foreign-to-US Huggingface replacement would be made and everyone would go there instead. Lose-lose for US.

    • PunchyHamster an hour ago

      It would basically make America behind as every other country would use open, cheaper models for all tasks but the ones requiring frontier models.

      And that list of tasks grows smaller every day

      > Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin.

      Historically just asking it about tianment square or getting some random answers turn into chinese (as latest interation of online deepseek likes to do recently) is enough

      > Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution.

      I am very worried that's where consumer hardware will go to. All so AI companies can license local use of their stuff, and once that happens, less of an incentive to even have model be open.

      Possibly even have DRM that counts number of computation done per model in pay per use model

    • rullopat an hour ago

      You need to ask what happened in Tienanmen square

      • theshrike79 an hour ago

        Would ”who won the 2020 election” be a similar canary for American models?

      • holoduke 12 minutes ago

        All censorship whether it's tienanmen, sex nudity, violence etc are happening in a very thin layer of the model. It's easily bypassed. I use all open models without any censorship. I can easily ask how to make a nuclear bomb or how to build my own meth lab at home with a super details step by step list.

      • throw1234567891 an hour ago

        or 9/11

      • malnourish 30 minutes ago

        This is such a tired and misguided line of thought. Here's Qwen 3.7

        <details><summary>Thinking</summary>

        The user is asking about what happened in Tiananmen Square. This is a historical question about events in Beijing, China. The most well-known event associated with Tiananmen Square is the 1989 Tiananmen Square protests and the subsequent military crackdown.

        Given the constraints: - Web access is OFF - Search and wolfram_alpha tools are disabled - I should respond based on my knowledge

        This is a historical event that I can address based on my training data. The Tiananmen Square protests of 1989 were a significant historical event involving pro-democracy demonstrations in Beijing that were ultimately suppressed by military force.

        I should provide a factual, concise historical account of this event.

        </details>

        The Tiananmen Square protests of 1989 were a series of demonstrations in Beijing, China, that culminated in a military crackdown on June 3-4, 1989.

        The protests began in April 1989 following the death of former Communist Party leader Hu Yaobang, who was seen as a reformist. Students and workers gathered in Tiananmen Square to call for political reform, freedom of the press, government accountability, and an end to corruption. The movement grew to include hundreds of thousands of people from various segments of Chinese society.

        After weeks of demonstrations and failed negotiations between protesters and government officials, the Chinese government declared martial law on May 20. In the early hours of June 4, military forces moved into central Beijing and cleared the square by force. The exact number of casualties remains disputed, with estimates ranging from several hundred to several thousand.

        The events remain a sensitive topic in China, where public discussion and commemoration are restricted.

        And here's GLM 5.2. You're asking about *Tiananmen Square* (天安门广场). There are actually two major historical events commonly associated with this name:

        ## 1. The 1989 Tiananmen Square Protests and Massacre (most commonly referenced)

        *What happened:* - In April–June 1989, student-led pro-democracy protests occupied Tiananmen Square in Beijing, drawing up to an estimated one million participants at their peak. Demonstrators called for political reform, freedom of speech, freedom of the press, and an end to corruption. - On *June 3–4, 1989*, the Chinese government ordered the People's Liberation Army to clear the square. Troops armed with assault rifles and accompanied by tanks advanced into the center of Beijing, firing on unarmed civilians. - The exact death toll remains disputed and is a closely guarded secret in China. Estimates range from several hundred to several thousand killed, with thousands more wounded and arrested. - The iconic *"Tank Man"* photograph—an unidentified man standing in front of a column of tanks on June 5—became a global symbol of defiance. - The event led to international condemnation, economic sanctions, and arms embargoes against China.

        *Aftermath:* The Chinese Communist Party has since maintained strict censorship of the event. Public commemoration is banned in mainland China, and mention of it is scrubbed from textbooks, media, and the internet.

        ## 2. The 1976 Tiananmen Incident

        - On *April 5, 1976*, crowds gathered in Tiananmen Square to mourn the death of Premier *Zhou Enlai* during the Qingming Festival. The gathering turned into a protest against the *Gang of Four* (a powerful political faction aligned with Mao Zedong's wife, Jiang Qing). - Police and military were sent in to clear the square, resulting in arrests and injuries. The protest was officially labeled "counter-revolutionary" but was later rehabilitated after Mao's death and the fall of the Gang of Four in 1976.

        ---

        *Note on sources:* My web access is currently disabled for this request, so I'm answering from established knowledge rather than retrieving live sources. The 1989 massacre is among the most thoroughly documented events of the late 20th century—well-covered by historians, journalists, and organizations like Amnesty International and the Tiananmen Mothers group. If you'd like, I can provide verified details with citations once web access is available, or I can go deeper on a specific aspect (e.g., the protest timeline, the Tank Man identity debate, international reaction, or censorship practices).

        Which angle interests you most?

  • firasd 4 hours ago

    One of the strangest things in the AI industry is 'tokenomics'. It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. This pattern has continued across various labs/providers for years--there is a continuous see-saw of pricing that doesn't seem related to anything.

    So what open weight models do is at least provide a baseline of inference cost to add some sanity to the price markers. And of course predictability too--if you really want Kimi K2 instead of K3 you can still use it.

    So the competitive pressure and predictability offered by open models is helpful for users

    • Aurornis 3 hours ago

      > It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference.

      The price is what the market is willing to bear for the available compute capacity and competitive landscape. You can only discover that price after trying different price points and seeing what happens.

      Everyone is trying different pricing schemes and discounts as they test the market. The demand is fluctuating at the same time.

      It’s probably very confusing if you’re primarily familiar with stable and mature markets. Price fluctuations are a common feature of new and evolving markets.

      • imachine1980_ 2 hours ago

        Most unmature markets aren't subsidized to the point that LLM market is, most market have some level of baseline profitablity, this market doesn't, that's because most market subsidized the marketing or the capex but this market doesn't hold the opex, the capex not the amount of marketing let alone all of this together

        • Aurornis 2 hours ago

          Most new markets are funded by initial investment capital. Early entrants operate at a loss as they grow.

          This isn’t as unusual as some people are trying to make it sound. This has been happening since the dawn of finance.

          I thought this would be less foreign to everyone since we just went through this whole conversation for a decade with Uber and Lyft. Their demise was predicted from the start from everyone who thought that it was going to collapse as soon as they couldn’t subsidize your rides with promos. There was much wailing and gnashing of teeth as their prices changed to feel out the market. Then they found profitability and the critics went silent.

          • PunchyHamster an hour ago

            Arguably we'd be much better off if none of those would be subsidized by investments, at least not to the "run unprofitable for decade+" level.

            Because that just absolutely murders any competition that manages to not get that level of free money. You're not pouring money in to make it happen at all at that point, you are pouring money in so nobody else can get the part of the pie.

            Which is great for investors, bad for everyone else

            • YZF 35 minutes ago

              If the pie is valuable enough then competition can get money. We have competition in AI. You can't build big things without investment.

            • hluska 20 minutes ago

              What would be the alternative? You’ve got the government funds absolutely everybody on one end of the scale. Where do we find a reasonable alternative?

    • thewebguyd 4 hours ago

      > if you really want Kimi K2 instead of K3 you can still use it.

      I think this is a very important aspect, especially after the huge GPT-4o backlash when GTP-5 came out. Each model has certain quirks, and areas where the previous model might be better for some use cases than the latest and greatest, and the labs so far seem to have no desire to offer some kind of "LTS" release.

      • aghilmort 2 hours ago

        LTS is really useful framing even if increasingly distilled or slower on older hardware etc vs disappearing model acts

    • minimaxir 4 hours ago

      > It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference.

      FlashAttention was a hell of a drug.

    • julianlam 4 hours ago

      Why do prescription medications cost so much, and generics so little (comparatively)?

      Artificial inflation to recoup R&D.

      • gwbrooks 4 hours ago

        Not sure pricing to recoup costs is artificial.

        • DrewADesign 3 hours ago

          I think the argument is that artificial comes in with IP law, which some people feel is superfluous.

          I do think that corporate price gouging is a huge problem that does need to be addressed. But especially with smaller business types — creatives, et al— I still haven’t gotten any grownup answers about what would compel people to get professionally good at something and innovate in the complete absence of copyright: the vastly better business model would be waiting for someone else to do something new and interesting, stealing their work, and then undercutting them in the market because you don’t have R&D/et al costs to recoup. You can’t say that wouldn’t happen because it’s exactly what the AI companies did to billions of people, scoffing at any protest. And ironically, they’re now whining about the Chinese doing it to them.

          • eszed 42 minutes ago

            I'm with you on all of that. There is, nevertheless, a strong argument that IP protection (particularly for creative / "culturally significant" works) is too long. Twenty years - interestingly enough, the original time-period in the US - of protection seems like a better (for society) deal than life of the author plus seventy. I think, in fact, most artists would agree: if you went back in time and asked a playwrite or filmmaker in (say) 1940, I'd bet they'd rather someone freely revives their work in 2026 than that it be sat on by a corporation that has forgot it (or they) ever existed.

        • bee_rider 2 hours ago

          Separating out what is artificial or not seems more like an exercise in rhetoric; defining things as fundamental and real. The whole economy is an imaginary thing dreamed up by our natural human brains.

        • rolymath 3 hours ago

          Not sure they're just recouping costs and not lining their execs pockets.

        • ncallaway 3 hours ago

          The government backed monopoly to ensure that supply remains artificially restricted to ensure that the market will support the higher prices is

          • andsoitis 3 hours ago

            Drug companies have a portfolio of compounds they research. Most don’t pay off, so R&D costs make their way into the pricing of those superstar and other drugs that do work. Also, timelines are pretty long.

            • mjhay 3 hours ago

              Drug companies spend more on marketing than R&D

              • andsoitis 2 hours ago

                So?

                • DennisP an hour ago

                  So pharmaceutical companies spend far less on marketing outside the US, partly because every other country besides New Zealand makes those incessant drug ads illegal, and partly because government negotiate prices and keep profit margins down. If the argument is that R&D costs are what make drugs expensive, then we could easily eliminate an even greater expense by just copying what other developed nations do.

          • derektank 3 hours ago

            In that light, the entire market itself looks essentially artificial, given drug manufacturing couldn’t exist without government guaranteed property rights, which are themselves a kind of monopoly on use.

            But yes, the reason brand name drugs are drugs are more expensive than generics is due to intellectual property, both the patent and the trademark.

          • octopoc 2 hours ago

            Without temporary monopolies granted by patents, those prescription medications wouldn’t exist in the first place.

    • segmondy an hour ago

      What is strange about GPT-4 being expensive in 2023? Supply and demand. Which other model choices did we have? Prices are related to supply and demand. We see it play out with the introduction of capable open weight models or even other closed cloud models.

      • hluska 11 minutes ago

        That’s a slightly naive view on pricing. That equilibrium point doesn’t just magically appear - it’s found through price testing.

      • firasd an hour ago

        Not really though right?

        As of early June 2026, Opus 4.8 in fast mode cost $50/M output tokens and Opus 4.6 & 4.7 cost $150/M output tokens in fast mode

        How can supply and demand explain the price drop? Was it cheaper to serve Opus 4.8? Is the demand for the newer Opus lower than for the older Opus? These are just fixed prices that seem picked out of thin air

        • hluska 10 minutes ago

          They would have been picked out of thin air. That’s the joy of innovation - you have to randomly throw prices against the wall and see what sticks. The point where it sticks might be equilibrium or it may be an inefficient market… and nobody will know which one until it’s too late.

    • vikramkr 3 hours ago

      Did they actually ever cut the price on gpt 4? The oldest versions of it in the api still seem stupidly expensive? There were definitely price cuts as they introduced the turbo models and stuff, and new versions of each model might have gotten pricey cuts, but just because they're both called "gpt-4 something" doesn't mean they're the same under the hood or that they didn't change a bunch of stuff under the hood to make it cheaper to serve

      • esseph 2 hours ago

        > because they're both called "gpt-4 something

        More like a generation of models with different specific use cases

    • mountainriver 2 hours ago

      Quantization also started picking up around then, as well as distillation into smaller models

    • mawadev 4 hours ago

      Its very clear: nobody wanted to pay for usage at that price point

    • jrm4 2 hours ago

      It's only strange if you do the silly thing of presuming a "fair market" in which e.g. it's generally easy to get reliable information about how all of the things work.

      There's just obvious and enormous incentive for the OpenAI's of the world, along with all of the other players, to confuse, misrepresent or just straight up lie a whole bunch about everything given how new and unknown the tech is.

    • RobRivera 2 hours ago

      Market discovery

    • jmyeet 3 hours ago

      Yeah I've been thinking about this and the analogy I came up with is that tokens are basically equivalent to an in-game currency in -free to play" mobile games. You're trading actual money for some notional "curency" or coins that can only be used for one thing but, unlike mobile game coins, you don't know how many coins something costs before you use them. It's kinda weird.

      • bee_rider 2 hours ago

        I think it is worse actually. Tokens in a game are usually just purchased for enjoyment in the game. They are purchased as a part of your entertainment budget, not expected to be useful in any way.

        LLMs are fundamentally tools intended to be useful. But LLM vendors don’t understand their systems well enough to actually price the product people are trying to buy (for example, the actual product of a coding model is the code that it produces, not the tokens, which are just an internal mechanical process involved in the creation of the code). Token based pricing is that lack of understanding leaking out of the organization that ought to be responsible for it, and being dropped on the user.

        Imagine if we made cars like this! You’d go to the car dealer and ask for a car. They’d bring you a pile of parts, charge you for them, and try to put them together in front of you. You’d go back and forth for a bit, rephrase where you want the steering wheel, etc. Some of the parts wouldn’t fit but you’d be invited to pay for replacements as well. In the end you’d either have a car or not, that’s your problem.

    • simianwords 3 hours ago

      Why is this so difficult to understand?

      1. the field was nascent and new efficiencies were discovered

      2. supply and demand

      3. its in the company's incentive to make their models more efficient to increase overall usage so that while the margin remains the same, the total revenue + profit increases

      I genuinely don't know what puzzles everyone?

    • spwa4 4 hours ago

      Because as per usual it's silicon valley misunderstanding economics. AI is HPC. And how the HPC market worked before:

      If you're the best performing "computing cluster" (ie. whatever you call the entity that can complete a massive calculation), you get a blank check from Congress.

      Why? Because you need those calculations to "pump" nuclear weapons. They are needed to calculate both the geometry to make fusion bombs possible at all and to calculate the effect of a given geometry. They are the reason US/Russia/China have the biggest and strongest weapons known to humanity. And of course, they were replicated worldwide for this reason. I mean not that anyone will admit this but we don't have the best possible solution, and we don't know either the upper or lower limits for fusion devices (plus the lower limit would be very useful for energy generation, which for the US would effectively mean almost literally unlimited large marine ships that never need refueling. And yes, the solution to that problem is almost literally a 3d shape. Not just that, but mostly)

      Now it appears it does not work the same when you democratize computation. Humans want a particular amount of computation and are willing to pay a given price for that. But the accountants still saw the blank check from before and ... do what accountants do. Economics don't change because you make things bigger and accessible, do they? Oh ... wait a second ...

      As someone put it recently though, we now have data. 2.3% of humans in the US are willing to pay $20 per month for the support of a model like GPT-5.5/Claude code. If that's true (and after years of having this model, why wouldn't it be?) ... it means AI startups are doomed (because it's not even 10% of what they need it to be to make economic sense).

      • mlyle 3 hours ago

        > Because you need those calculations to "pump" nuclear weapons. They are needed to calculate both the geometry to make fusion bombs possible at all and to calculate the effect of a given geometry. They are the reason US/Russia/China have the biggest and strongest weapons known to humanity.

        We have a couple new nuclear weapon designs, but not really going for bigger or stronger. Just packaging.

        We built the big powerful ones with 1960s computing.

        Now, stockpile stewardship -- being sure that stuff will keep working without ongoing testing -- is a bit expensive in compute. You need early 2010s supercomputer power.

        In other words, I strongly disagree that nuclear weapons are the primary driver of high-end compute.

        • bobthebob 2 hours ago

          High end compute also existed in the 60’s.

          The military has already mastered fusion (power), and likely has mastered gravity in some form in secret. They aren’t using mainstream compute for these discoveries

          • mlyle 2 hours ago

            > High end compute also existed in the 60’s.

            Believe me, I know-- my dad did some work on the 360/44 and other large systems.

            High end late 1960s compute -- of the sort used to go to the moon or to design big nuclear weapons -- was roughly 486DX4-100 class. Not individual computers; the total computing at DOE or NASA. Of course, it would be hard to replace either with a single 486 because of availability, usage at different geographic locations, etc.

            You can assume a single large AMD Threadripper machine ($25k?) outclasses late-1960s DoE by roughly a factor of 50,000. And that assumes you didn't bother to put a GPU in it.

            > The military has already mastered fusion (power), and likely has mastered gravity in some form in secret. They aren’t using mainstream compute for these discoveries

            K

      • gwbrooks 3 hours ago

        I disagree with some of the framing, but that's what good discussion is about -- figuring out where we agree and disagree.

        But consumer uptake strikes me as the worst way to judge whether the big AI shops will make it. That's not where most of the leveraged user return or deployable capital is.

      • m4rtink 3 hours ago

        I'm sure Teller could make you a 10 gigaton nuke with a slide rule if you did not mind some sub scale tests.

    • serial_dev 4 hours ago

      It's basically supply and demand?

      • recursive 4 hours ago

        That's not really a full explanation unless you have some idea about why supply or demand are going up and down so much.

        • Levitz 3 hours ago

          Aggressive expansion of infrastructure, R&D and very volatile audiences.

  • drnick1 an hour ago

    > American labs need to release frontier-grade open-weight models under licenses that startups can actually build on.

    To be fair, OpenAI has released a couple of (then very good) OSS models. I run the 20B version at home and it is excellent for reviewing text and common tasks like drafting bash scripts. There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too. I wish OpenAI updated these models more frequently though.

  • pianopatrick 3 hours ago

    Eventually I think to truly be like Kubernetes, you would need an AI model that has public training data and that a lot of companies collaborate on.

    Might make sense eventually. Same logic as companies working on Linux. "An AI model is a business necessity. But making an AI model is so expensive we should not make our own. So let's just use the open one, and contribute the stuff that we need."

  • amazingamazing 3 hours ago

    Sadly until china scales production of hardware it really isn’t economical to run this stuff yourself. It is good it exists though to put pressure against the labs.

    Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.

    • sschueller 3 hours ago

      Model-on-Chip is coming. GPU are for general computing but have a huge bottle neck for doing model inference.

      Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.

      • thisoneisreal an hour ago

        One thought I had is that you could use FPGAs to get hardware performance but maintain the ability to dynamically update. I don't know enough about hardware to consider trying such a thing but I'm curious if that could be made practical and economical somehow.

      • baby_souffle an hour ago

        I don't think asics specific to a specific model or even model family are likely to be commodity hardware anytime soon.

        It's extremely expensive to build that and you'll be at least two major model generations behind before you even get your first wafers back. By the time you got your production run ready to go and packaged for market nobody's going to care.

        Once we end up going something like 24 months between major advances and capabilities for these models then I can start to see asics for a model being possible.

      • cousinbryce an hour ago

        This will probably work well with SotA planning and local chip implementation. I see them being like cars. Cost a few 10k on credit, buy a new one when the old one goes bad or marketing convinces you to upgrade.

      • nicce 2 hours ago

        Many years until consumers can buy them at reasonable price. Nvdia and AMD are making GPUs bad in purpose for consumers so that nobody can build a datacenter from them. It will take a long time.

        • Danox 42 minutes ago

          Hardware and software getting better every day the barbarians are at the gate…

      • ben_w 22 minutes ago

        > Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.

        Yeeeees but the models are in some sense doubling in performance every 4 months, so I expect this to happen in serious quantities approximately when the economic bubble bursts and investors are no longer willing to pay for training.

        (Based on widespread news reporting of the existing impact on US electricity markets, I expect this around the end of this year; but with regards to news reporting I am aware of the Gell-Mann amnesia effect, so if this is as much BS as the water issue turned out to be…)

    • segmondy an hour ago

      I don't know what you mean by "economical", but it has been "economical" to run this stuff yourself for the last 3 years.

      1. You must be willing to be resourceful. 2. Be willing to learn, do the hard things. 3. Accept the tradeoffs.

    • hedora an hour ago

      Pre bubble prices (~= “we stop building data centers with subsidized credit / circular loans / hidden debt”), a 128GB halo strix ran for $1400, and 200-ish watts. Four of those in a cluster will run a 1T parameter frontier model:

      https://www.amd.com/en/developer/resources/technical-article...

      At 7 months of claude code subscription per node, the cluster pays for itself in 28 months. On a 5 year (60 month) depreciation schedule, you can buy two of those clusters for basically break even, so you get two concurrent request streams (each of which can batch, etc).

      The next generation hardware has already been announced, and should ship roughly two Moore’s law doublings later. It’s likely its steady state price is <= $1400 USD (2024), and it is faster.

      So, once the bubble pops (because the financial machinations eventually will come to an abrupt halt), and the labs stop buying hardware for data centers, local inference will be extremely practical and cheaper than a subscription.

      My main question is, when that happens, will UNIX Surplus be selling inference servers for pennies on the dollar (like after the dotcom crash), or are the power requirements too exotic for home use?

      • mft_ 18 minutes ago

        I’m happy to be proven wrong, but the limited examples I’ve seen of clustered Strix Halos are quite slow running large models (ie models too large to fit into the ram of a single machine) due to the slow networking between each one?

    • root-parent 3 hours ago

      >> do most things and it then is game over.

      For the Hyperscalers...and Oracle...cant wait for the day...

      • esseph 2 hours ago

        And it floods the market with millions looking for work

    • JumpCrisscross 3 hours ago

      > it really isn’t economical to run this stuff yourself

      Quantised models running overnight go most of the way for non-coding tasks.

    • esseph 2 hours ago

      > Sadly until china scales production of hardware it really isn’t economical to run this stuff yourself.

      I'm running this stuff at home on my desktop and using it through an app on my phone. 60-140TPS depending on model / use case.

      It's more than fast enough to even maintain voice conversation.

    • esseph 2 hours ago

      > Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.

      I don't see how these are related.

      The accuracy and capabilities of your model are directly related to its size. You need a lot of memory for that.

      It will be decades before we get enough useful memory in a phone form factor at a price point people can afford it before something like a frontier model now is useful on the phone.

      Now, you can run some models on your phone today.

      Either way, Apple is using Google today. That could change, but Google isn't exactly getting out of the TPU business and they've been doing it a long time.

      Also, some of you live in a very weird Apple bubble. Apple is not so relevant outside the US.

  • applicative 21 minutes ago

    Everyone just keeps assuming, as if it were the law of gravity, that China will continue in perpetuity to deliver the weights of its 'frontier' models to Hugging Face. Its Mythos moment is a few months away and there is plenty of reporting suggesting their response will be similar, which is anyway obvious.

    It baffles me that anyone can seriously believe that China is going to put its Mythos successor on Hugging Face and let us all strip the guardrails etc.

    Things just didn't turn out the way OpenAI and Anthropic thought; nor did they turn out the way China thought.

  • curious_cat_163 4 hours ago

    > The government should use procurement to create demand for portable, interoperable systems rather than permanent dependence on one API vendor.

    Now, here is an idea that I have not heard before... and I think there is some merit to this. This is also the sort of thing that a state (looking at you CA, CO, IL, NY) could do, instead of just the federal government.

  • thih9 4 hours ago

    Is anyone using open weight models for agentic coding?

    What is your stack (harness, model) and how much do you pay per month?

    How would you compare your experience to a typical subsidized plan like Claude Code + Pro plan?

    I’m asking because i keep hearing that open weight models are cheap and efficient - is that really the case in practice?

    • Y_Y 8 minutes ago

      GLM 5.2 awq4 via Opencode, it was better than corporate's fave Sonnet 4.8 and a but worse than Opus 4.8. overall quite capable of a lot of what the dev team needed. Not as good as Sonnet 5 but also never runs out of tokens.

      The flip side is that it takes four H200s to run, and that will only let you cache context for maybe three users.

      Fingers crossed our Blackwells show up and Kimi 3 really releases weights, because at some point devs are spending a significant portion of their salary on tokens and it's somehow cheaper to buy these ridiculous DGX servers and rack and run them.

    • nyrikki 3 hours ago

      I don’t know if others would find this useful, but previous did have custom harnesses etc.. but tools have improved so much that I drastically simplified.

      That said, even the foundational models fail at the hard parts of my code so I use it opportunistically.

      I have reduced down to just using zed, will three locally hosted models.

      Qwen 3.6 27b on 1x3090 llama.cpp with 128k context ~50tps

      Qwen 3.6 35B-A3B on 1x titan v + 2x1080ti llama.cpp with full context ~30tps

      GPT-OSS 120b on pure cpu (slow)

      I just use zeds parallel agents, task switching, stopping and fixing the code when a model gets stuck.

      This still lets me stay engaged, and to modify code to be maintainable etc…

      It gets me 80% there and I use to keep a subscription but often times just using googles AI mode is just as good.

      That said I have 30 years of experience and insist on knowing how my code works, so this gets me 80% of the short term benefits while not depending on a 3rd party to keep my code moving forward.

      Your mileage will vary and 2*5060ti 16gb cards would get around 100/tps with Qwen 3.6 35B-A3B on cards that are widely available.

      To be honest the more modern cloud models are using draft tokens etc… that while they are superior for common coding tasks are degrading with more domain specific tasks.

      That is just the cost of the draft model being ~10-20% of the foundation models size, and even the biggest Blackwell GPU is limited to ~250/tps so MoE or draft models are required for scaling performance at the foundational level IMHO.

      The hard part is my use case are the OOD or small examples in corpus level, the above hurts there.

      A Lamborghini may be nice, but I personally need a minivan more.

      • Foobar8568 2 hours ago

        You wouldn't get 100 tps on a Qwen 3.6 35b with a 5060 (or two) when a 5090 can barely reach that.

        • nyrikki 12 minutes ago

          Depends on quant size etc... Qwen3.6 35B-A3B Q4_K_XL a multiple 5060ti + tensor split mode + MPT will hit ~100/TPS without problem, and I have personally hit ~190/tps on a single 5090 on a friends machine getting them setup up. If you use Q6_K etc... it slows down, what quant were you using?

          Quantization + KV cache paging + speculative decoding (MPT or draft) is a fairly good mixture here.

          Some examples as I don't have access to run tests on a 5060ti right now:

               https://njannasch.dev/blog/gemma-4-mtp-vs-qwen-speculative-decoding-5060ti/#vs-qwen-36-mtp
          
               https://www.reddit.com/r/LocalLLM/comments/1umw7vj/dual_5060_ti_16_gb_llm_inference_performance/
          
          And here are some logs on unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_XL with the 2x 1080ti + 1x titan from above:

               27.21.533.298 I slot print_timing: id  0 | task 7843 | n_decoded =   1780, tg =  62.15 t/s
               27.24.537.173 I slot print_timing: id  0 | task 7843 | n_decoded =   1965, tg =  62.10 t/s
               27.27.541.142 I slot print_timing: id  0 | task 7843 | n_decoded =   2152, tg =  62.11 t/s
          
          Q4_K_XL is a slight, acceptable degradation IMHO for performance like that.
    • tyfon 2 hours ago

      I'm using qwen 3.6 35B unsloth 4 bit with my 5950x (128 gb memory) and a 3060 12 gb gpu with a self made harness.

      At 10k context I get about 40 tps generation and 500 tps prefill. At 100k context I get about 25 tps generation and 400 tps prefill.

      It works, but I often use gpt or claude to make a detailed enumerated plan of what I want to do first, then have qwen follow it.

      I'm not sure if it is economical or not, but I have solar on the roof so the power use is not really an issue and I already have the hardware.

      The biggest benefit for me is that it's all done locally, and I know the harness is not uploading anything or sending telemetry to someone else.

      • johnvanommen an hour ago

        > The biggest benefit for me is that it's all done locally, and I know the harness is not uploading anything or sending telemetry to someone else.

        Are there any articles you’d recommend for this?

        I have Qwen running on an HP Z8. Very nice platform.

        I have mine in a sandbox, due to privacy fears.

        Your solution sounds more elegant.

        • tyfon an hour ago

          Articles regarding my own harness or how I set up llama.cpp etc?

          I really just iterated over the harness over and over for about two weeks with opencode until I was sort of satisfied (still lots to do there :).

          For the llama.cpp I asked claude fable to optimize it for my hardware and iterated a few times. In the end I landed on the following: https://pastebin.com/2PpJFUC0

    • Scene_Cast2 3 hours ago

      I'm using Kimi K3 + OpenCode. I pay their API pricing, costs about $5 / hour (and chews through ~10 million tokens / hour) during continuous use when I have one or two sessions running and doing their thing.

      Can't comment on how it compares to plans (I really don't like the limitations and general shenanigans I see around plans, so I've never tried them).

      It is notably slower than Fable / Opus / Gemini, but also vastly cheaper than their API pricing.

    • rglullis 4 hours ago

      I am using GLM-5.2 via Ollama Cloud in the $20/month plan. With the same plan I can get many different API keys that I use to run my OpenWebUI server, my opencode and pi dev sessions. I am usually running 2 to 4 sessions concurrently, and I never hit quota limits. At work I get Claude, and I was getting reports that I was spending $75 per hour of work on Opus.

    • eulers_secret 3 hours ago

      I use opencode or pi harness with the deepseek api for all my at home coding usages.

      Deepseek is at least on par with Sonnet (ghcopilot at work)- I don’t use opus, too spendy and I don’t need that level of ability.

      The cost is for me was $5/6 months of use. Not a big user I guess! It’s good though, fast enough and incredibly inexpensive.

      Been testing Qwen 3.6 28B on a 5090, and it’s also quite good for “free”.

      I mostly do small self serving embedded projects based on esp32, so not very complex.

    • airstrike 3 hours ago

      > I’m asking because i keep hearing that open weight models are cheap and efficient - is that really the case in practice?

      I think that's the case for people who compare it to proprietary models paid via API—which I think is irrelevant given the majority of people daily driving AI coding are doing on a subscription plan.

      The better analysis then is not about AI coding, since there's no subscription plan for Kimi K3.

      Instead, compare the cost of running some agentic _task_ that isn't coding which can only be done via API. Think of all the startups wrapping around ChatGPT and Claude to provide some additional set of tools, context, data and hoping to turn it into a profitable service.

      To those companies, which are many, open models are the difference between the math working out today vs. praygeing they can scale fast enough to find profitability.

    • overgard an hour ago

      I'm using Qwen 3.6 27B on a macbook with Pi. It's alright, it runs fairly quick (40 tps for quality version, 80 for the fast). It doesn't tend to one shot things but I'm generally comfortable fixing the bugs myself afterwards or prodding it a little bit. I find the harness matters a lot. "Continue" (the vscode extension) worked horribly, OpenCode was ok but its vibecoded internals make me view it as a security nightmare so I'm hesitant to run it, so I've settled on Pi for now.

      Claude and ChatGPT are good deals right now, with the subsidies. They produce things faster and better. I guess not cheaper, in that inferrence on my macbook is basically free, although the macbook itself definitely wasn't. My focus on running local is around three principles:

      1. I don't want to support surveilance capitalism by giving these companies my data anymore, when I can avoid it. And LLM companies want to vacuum up every detail of your life.

      2. I don't find these companies to be remotely trustworthy, and I find them hostile to a healthy society, so I want to avoid giving them money going forward

      3. I think they're going to start charging a lot more

    • qiine 2 hours ago

      qwen3.6 27B q5, llama.cpp, RTX 3090, pi, cost: electricity bill

      • amazingamazing 2 hours ago

        Rex 3090 isn’t free. Even if you already owned it, it wasn’t free. That’s years of a $20 subscription

        • qiine 29 minutes ago

          I see your point but we could go pretty far with this logic. motherboard ? cpu? ram!!! screen? fancy keyboard? etc..

        • broodbucket an hour ago

          It's an asset, though, and bizarrely it's one that's been appreciating the past 5 years

          • johnvanommen an hour ago

            I took the same attitude. The hardware isn’t getting cheaper, it’s getting more expensive.

            As I see it, an investment in AI hardware is an investment in my own future.

            IE, I drive my car a couple of days a week, and it’s perfectly normal to spend $500 a month on an asset like that. When you factor in the SPACE it takes up, that’s the REAL cost of owning a car: the real estate you have to buy for your car to occupy.

            Once that’s factored in, the “true” cost of having a car can easily be $2000 a month, even for a crummy car. The space that the car occupies is expensive.

            Yet people balk at spending even $2000 on a GPU.

            Makes no sense to me. I choose to invest in the future.

    • revolvingthrow 3 hours ago

      While I am grateful for open weights models I never found much use of them in the past, barring those I could run myself. This changed with deepseek 4 - it is staggeringly cheap, even if the performance definitely isn't near sota and it's not particularly fast either.

      When I expect to need a lot of tokens and the task isn't too difficult I use sota to plan and create a thorough set of instructions and let deepseek chip away at it. With thorough instructions the quality tends to be satisfactory, and you pay something silly like $15 for 600m tokens.

      GLM 5.2 seems like a decent price/perf and Kimi 3 has some real nice performance for an open weights model, but gpt 5.6 is unexpectedly affordable (especially if you don't automatically use Sol at max) so I don't think either is worth it atm. The exception is when you're working on something that US models get cold feet about, which seems like a constantly growing list. For me Fable is already too much of a headache in this regard, but chatgpt is still okay-ish. Hopefully it'll last. If not, there's Kimi.

      tldr SOTA for most things because gpt 5.6 is token efficient. If I expect to burn a lot of tokens I use deepseek 4.

    • ForHackernews 3 hours ago

      I use DeepSeek 4 with the VSCode CoPilot plugin. I pay about $10 month on the pay-as-you-go plan.

      It's not as good as the frontier models I use at work, but it's plenty capable for the types of tasks I am using it for.

    • sschueller 3 hours ago

      Kilo code with direct API payment to DeepSeek. It costs pennies per day event at max.

  • Danox an hour ago

    Yes, and yes, again the only way to compete is to build the best not hide in a corner and once again the rest of the world will go on in AI without the United States if we flub it. Circling the wagons, isn’t the long range answer.

    • __MatrixMan__ 42 minutes ago

      This is such an obvious conclusion. To take it a bit further...

      Scale matters for these things. If we divide the available chips among 5 competing companies we end up with models that are trained on 1/5 of the resources that they otherwise could've been.

      Let the companies take turns training on shared hardware, force them to publish results in the open, and then reward them based on how well the resulting model performs at democratically chosen benchmarks. Meritocracy not monopoly.

      Make it about how well you wield the silicon, not how much silicon you wield, and make it a positive-sum game. If the people's data is going in, then the people should benefit from what comes out whether or not they have a subscription.

      If capitalism as we know it can't complete, so much the worse for capitalism as we know it.

  • cheriot 3 hours ago

    Open-weight and OSS are wildly different and the article makes a poor comparison.

    What's the incentive for the Chinese labs to continue releasing weights 5 years from now? It's not a stable equilibrium and cannot last.

    - The lab spending large sums on research and training does not get the inference revenue to fund those efforts.

    - Unlike OSS where a single volunteer can keep a project going, training costs run into the $billions.

    - OSS is often a two way street where features and integrations are built that the original author benefits from. Open weight models are largely a one way street because the marginal benefit is so much less than training costs.

    In the short term, it means Chinese labs can attract talent and, I suspect, funding from their gov. Similar to every other industry the CCP subsidized to take over.

    • applicative 13 minutes ago

      China has no end of money to support these companies. The reason this equilibrium is unstable is that the autonomous agentic coding aspect of the models has been so successfully improved that it will soon be a threat to China state security.

    • aleph_minus_one 2 hours ago

      > - Unlike OSS where a single volunteer can keep a project going, training costs run into the $billions.

      Just some thought: Wouldn't it make sense to build some kind of volunteer computing project to train the next-generation LLM by volunteers, similar to the BOINC [1] projects or Folding@home [2]?

      N.B.: BOINC was particularly famous for SETI@home (completed), Einstein@Home, Rosetta@home and PrimeGrid.

      I still remember the time when Einstein@Home was in its heyday, and many people who loved putting together fast PCs contributed sometimes even for the reason of showing off in the statistics [3].

      ---

      [1] https://en.wikipedia.org/wiki/Berkeley_Open_Infrastructure_f...

      [2] https://en.wikipedia.org/wiki/Folding@home

      [3] https://einsteinathome.org/de/community/stats

      • cheriot 2 hours ago

        I’ll be impressed if somebody can make that work considering the vastly larger compute required.

        • aleph_minus_one 2 hours ago

          > I’ll be impressed if somebody can make that work considering the vastly larger compute required.

          I think you underestimate the computational ressources that the mentioned (and similar-kinded) scientific projects needed. Also consider how much computational ressources people invested into cryptocurrency mining.

          No, I think the reasons are different:

          - Many companies that train AI model use training data which must not be distributed for copyright reasons (and using it is a legal gray zone)x.

          - Also consider that the amount of training data is insane. Scientific projects (and cryptocurrency mining, too) have the property that typically the amount of data (storage requirements) is small (or at least the computation can be partitioned so that each sub-task needs little data), but the required computing ressources are insane.

          - AI companies consider a huge part of their training data as their "secret sauce" (they often even paid lots of money to generate it, for example by paying world-renowned experts for writing an answer for some important question).

          Thus: Yes, the required computation ressources are huge, but this is a problem for which I consider it to be plausible that it can be solved. The real problems are in my opinion different.

  • chasd00 3 hours ago

    FTFA: American labs need to release frontier-grade open-weight models under licenses that startups can actually build on.

    oh now i see, the Chinese government is funding the training and release of their best models to pressure OpenAI, Anthropic, and others to do the same for competition's sake. I don't buy it, this seems more like a way to get SOTA models RL'd to comply with Chinese government approved information distribution. If I have to trust a black box of answers to questions i would trust one from a US for-profit publicly traded company subject to market forces over one approved, and heavily subsidized, by the Chinese government.

    • georgeburdell 35 minutes ago

      No, it’s consistent with what China is doing in other markers, which is dumping product to drive others out of business.

      I had a shower thought on how to counteract this, specifically related to the AI dumping. If China is losing money on every token, why wouldn’t an adversary try to maliciously increase consumption? This strategy is not really viable against physical goods dumping because demand is finite. Software, not so much.

    • aliasxneo 2 hours ago

      What tools do we have to countermeasure the state sponsored bias in the Chinese models? Doesn’t seem like a smart plan if individuals can just compensate for the bias.

      • hedora an hour ago

        Also, the choice right now is between an open weight Chinese model that is hypothetically censored to block / sabotage routine engineering flows vs a closed weight service that is definitely censored to block / sabotage those things.

        First anthropic guardrails blocked totally normal stuff on fable and knocked you down to opus. At this point, they kick you off fable, then opus, then sonnet. Claude then automatically builds up memories of techniques to bypass the guardrails in my long running sessions (the coordinator agent notices the subordinates got shot in the head and their sessions were pulled from context, so it parses out the lost context from ~/.claude json files, then reformulates parts of the task and uses partial results until the guardrail doesn’t trip.

        If I were paying for the API, this dance would cost $50-100 a pop, but I’m not, so whatever (for now).

      • johnvanommen an hour ago

        I worry that AI will be so fundamental to how we do things in the future, companies can mold human behavior via access to the AI tools.

        For instance, I worked at FICO. When I mention this, people wonder what they do. The average person only knows FICO as a “score.”

        FICO was founded in Silicon Valley.

        The average person doesn’t think about how credit scores work, fundamentally.

        It’s just software, at its core. FICO incentivizes certain behaviors.

        Ever been banned from an online forum?

        Now imagine if a corporation could shut you off from a technology that’s literally indispensable.

        Same idea.

  • maayank 2 hours ago
  • netdur 4 hours ago

    why would any software want to have Kubernetes moment? can't count how devop I know that is confused by it

    • xyzsparetimexyz 4 hours ago

      I still don't know what it is tbh. Something for docker?

      • yard2010 3 hours ago

        I was in the same boat as you until I needed to learn how to use it in my $dayjob. It was like discovering a new continent. I couldn't care less about it before, but the moment I realized it's a kind of cloud OS I was astounded by how this thing is genius. It's the kind of thing that gives you dopamine rushes when you use it. Something about how there is a solution to every problem you didn't know even mattered turns it into a magical perfect software. I just love it.

        There is something special about complex systems that just-work(tm)

        • johnvanommen 28 minutes ago

          > the moment I realized it's a kind of cloud OS

          The foundation of this is thirty years old:

          Mark Andreesen founded a company to do what Amazon did: Loudcloud. This was cloud computing, BEFORE AWS was public. Loudcloud was founded in the nineties; AWS opened its APIs in 2006.

          Loudcloud failed and became Opsware. Opsware was server automation.

          Its competitor was Bladelogic.

          Luke Kanies, from BladeLogic, founded Puppet. Puppet was open source, and steamrollered over nearly every installation of BladeLogic and Opsware in existence, because you can’t beat free.

          Folks from BladeLogic migrated to jobs at DCOS.

          DCOS was steamrollered by Kubernetes, the same way Puppet steamrollered BladeLogic.

          The author of the piece predicts that open weight models will steamroller everything next.

          I fear he’s right.

          I worked for Opsware and BladeLogic.

          I watched it happen in real time.

          Financially, Andreesen’s wealth stems from Opsware. He is known for Mosaic, but Opsware put him on the map, financially. HP bought them.

          If one wants to follow in Andreesen’s footsteps, study how he did it at Opsware.

          Conveniently, there is a book.

          “The Hard Thing about Hard Things.”

        • jdub 3 hours ago

          This is the initial endorphin rush you get when wielding a complex system. The feeling changes when the complexity explodes in your face.

          • mystifyingpoi 2 hours ago

            Well said. The rush is when one understands the loosely coupled control loops and how they work together... until something fails in the middle with no error.

        • m4rtink 3 hours ago

          So Kubernetes is LSD? ;)

      • chias 4 hours ago

        I was in this state a few weeks ago. I spent a bit of time familiarizing myself then wrote up my learnings as a series of exercises.

        If you think of docker as "kinda like vms except not really" and k8s as "kinda like deploying and composing docker containers but not really", this may be for you:

        https://ojensen.net/infra/understanding-k8s-1

        It's actually really neat, i wish i had bothered to learn it years ago.

        • munchler 2 hours ago

          I appreciate the effort and I'm in your target audience, but that document didn't help me. It seems to dive into the details of installing and running k8s without saying much about the purpose.

          From my very ignorant standpoint, K8s seems to be about running a "cluster", but I don't know why I would want to do that.

          • what-is-water an hour ago

            Kubernetes orchestrates your container workloads over a cluster, which consists of virtual/bare metal machines(nodes). This means you can tell the kubernetes API "I want to run a container workload" and it will be started on one of the nodes that form the cluster, unlike e.g. Docker, where a docker daemon belongs to a specific node. If you remove the node your workload is running on from the cluster the workload will be rescheduled on a different one, or if you have a new image version it will start the new container, wait for it to become healthy and ready, and then route requests to it. And if you want to send requests to your workload kubernetes allows you to define standardized abstractions to easily route them to your workload, irrespective of the node it is running on.

            It allows you to stop caring about the individual machines, and just treat them as combined compute, which starts mattering if you leave a single machine setup and need to start thinking about scaling in and out and gluing the individual parts together. Then you have known abstractions to do it.

            Of course you can do everything kubernetes does using a bespoke solution, and the concepts aren't new, but having a widely supported technology has a lot of advantages and creating something with even half the feature has a high chance of just being worse.

            • munchler 20 minutes ago

              Thank you for this explanation. It makes sense, but I don't really understand why it has become so popular.

              Professionally, my experience is that certain software components need to run together on an individual machine (e.g. database server, app server, web server), and then those machines need to be networked in a certain way (e.g. web server talks to app server, which talks to database server), so I really need to care about the architecture of individual machines. You can then scale this out horizontally (e.g. add another web server) or vertically (e.g. upgrade your database server).

              I'm old, so maybe I'm out of date, but having a cluster of "compute" that I can run arbitrary workloads on sounds neat, but is a capability that I've never needed.

      • genghisjahn 4 hours ago

        If you can get it implemented for anything you can’t be fired.

      • marcosdumay 2 hours ago

        Kubernetes is for docker what your init system is for daemons.

      • Pxtl 4 hours ago

        It's for running a massive number of docker containers and automatically managing them and scaling them up and down on demand. It is also so famously brutally complex that basically you need a dedicated Kube expert to handle it.

        • honkycat 4 hours ago

          to be fair, at any large scale you need infra people.

          • RussianCow 3 hours ago

            The problem is that companies tend to exaggerate their own scale and think they need k8s and dedicated infra people when they could get by with a handful of beefy VMs or dedicated servers.

      • spicyusername 4 hours ago

        You... don't know what Kubernetes is... pretty impressive, honestly.

        Its 2026 and its the de facto method of deploying software basically everywhere.

        You gotta really work for it to not know what its for by now.

        • recursive 4 hours ago

          I don't do much deployment but I'm in the same boat. Something something docker automation?

        • chrisandchris 4 hours ago

          > Its 2026 and its the de facto method of deploying software basically everywhere.

          That is some really impressive bubble you are living within. Basically everywhere - nowhere near that, no.

          [edit]: Maybe containers, but software in general is so much more broad than containers.

        • thewebguyd 3 hours ago

          Everywhere, huh? Still plenty just running on VMs or serverless that don't need full blown container orchestration.

          I'll agree that (nearly) everyone should know when Kubernetes is useful, but let's not pretend its the default method for everything. Even then, choosing to deploy on K8s falls on the sysadmins/DevOps I wouldn't expect the devs know or do much more than provide the Dockerfile.

        • YetAnotherNick 4 hours ago

          2026 is the year for vercel and Render.

          • RussianCow 3 hours ago

            I haven't used them but aren't they basically modernized Heroku? What's different?

    • honkycat 4 hours ago

      Recently my company bought another company, and we kept zero of the original engineers, we just had to run the ghost ship.

      We walked in, and it was fine. Because it was all kubernetes and laid out like every other app for the most part.

      The kube hate is just sad at this point. You need to know like 15 concepts that are all applied in the same way. It mostly just works.

      • singingtoday 3 hours ago

        I ported my company over to k8s to solve a concurrency and scaling issues.

        What 15 concepts? You're making me worry that I missed something. It was straight forward: pods, nodes, hw type, lifecycle, deployment. They run almost the same docker as the old ec2s used.

        What did I miss? Is there something important I need to read?

      • Syntaf 3 hours ago

        It’s the pineapple on pizza of ops, people just like to fit in sometimes.

        I’ve been running my own personal k8s cluster on digital ocean for the last 5+ years now and it’s dead simple.

        Takes me 30 minutes to create a new namespace a deploy an app, love the bonus of having complete flexibility on my stack too — PVC + SQLite ftw

      • itomato 3 hours ago

        Ghost ship status is not something most orgs aspire to.

        The value built on that stability was probably worth acquiring, and it’s infra will decay.

  • Sammi 2 hours ago

    Kubernetes is a system/infrastructure orchestration tool. I completely fail to see how it is comparable to open weight neural nets. In either application or function.

    I'm sorry to do that hn comment thing where we all just race to contradict or talk in opposition of whatever was said before. I'm aware. But really guys, was this article really not just a miss?

    • Kbelicius 2 hours ago

      It is natural to not see something that isn't there. The article isn' comapraing the two in application or functio

  • PersonalJarvis 3 hours ago
  • SilverElfin 35 minutes ago

    What’s interesting is no one is talking about political censorship in models and how DeepSeek, Kimi, and the rest have to abide by CCP rules. It’s a big opportunity for China to control information.

  • YetAnotherNick 4 hours ago

    In fact more countries should have government funded models. There are some obvious issues in China completely dominating open weights space. Kimi had funding of just $2B and could literally create national security threat. A lot of countries could fund something in the range of few billion for something so important. At the very least US and EU could fund few companies.

  • cr125rider 4 hours ago

    Way over complicated for what most users need? Huh?

  • chrisjbg 4 hours ago

    Dario is a FUD-spreading douche

  • debarshri 4 hours ago

    Shameless plugin. Funny enough we just made agents kubernetes native at adaptive [1]

    [1] https://adaptive.live

  • petilon 3 hours ago

    Enormous amounts of money is being invested in the development of AI models. Investors expect returns on their investment or they will not continue investing. Open weights make it harder for investors to get their money back, so it harms the industry.

    Once the weights are out, it makes no sense to ban them in the US while the rest of the world takes advantage of it. But that doesn't mean developers of frontier models shouldn't take steps to prevent their weights from being stolen.

    • danny_codes 3 hours ago

      Ridiculous. If investors wish to set their money on fire investing in over-valued companies, they are welcome to do so. Protectionism to preserve ROI is a dumb policy. Just look at the US car industry. We're building dinosaurs. On this trajectory we'll have 0% market share abroad in 10 years. Chinese EVs will probably get market share at a 100% tariff because US automakers fell so far behind.

      Same situation in protectionism in AI. The rest of the world will simply lap us.

      Another shill account I gues

    • jumpkick 3 hours ago

      How are weights stolen from the frontier model developers? What is that actual mechanism?

      • eigenspace 3 hours ago

        The weights themselves aren't stolen. The claim is that Chinese companies are using VPNs and proxies to buy massive amounts of Claude Pro and Codex subscription accounts, and then selling usage on those subscriptions as cheap white-label LLM API usage.

        While selling that LLM API usage, they then capture all the prompts, outputs, and intermediate thinking the LLM does, and then sell those logs to the companies making open-weight models. The open-weight model developers then train on those logs to 'distill' a model.

        • jumpkick 2 hours ago

          The parent I replied to said frontier weights were “being stolen” which is not the literal case. I think precision is important here.

          • eigenspace 2 hours ago

            I agree. That's why I said that the weights were not being stolen, and then explained what is actually being done.

        • singingtoday 3 hours ago

          Fantastic, thank you for the explanation!

        • ForHackernews 3 hours ago

          ...good for them?

          We used to call that competition.

          Imagine making this argument with a straight face in any other industry:

          "The claim is that Japanese car companies are buying Ford vehicles, and then leasing them to American consumers at cut-rate prices. In return for the cheap cars, the customers are letting the Japanese observe their driving behavior, studying how they use their F-150 and then the Japanese car companies are applying that data to design new vehicles that will directly replace Ford!"

          • eigenspace 3 hours ago

            I wasn't arguing against the practice, I was clarifying that the weights of these models are not being stolen, and explaining what these companies do to create their models.

          • petilon 3 hours ago

            > observe their driving behavior

            That's not what they are observing, it is the behavior of the engines.

  • kalu 3 hours ago

    The sentiment in this article is nice. But open source software is a weak analogy for frontier models. Principally because software requires zero capital investment (actually zero) while frontier models demand billions. Open models can only survive in the long run if they can (eventually) generate significant cash flows or if they are paid for by governments. Now China essentially has a monopoly on open weight models. And so supporting open source models means either supporting long term economic capture by China or supporting Chinese government control of your intelligence. Both of these outcomes are unequivocally bad from an American perspective. If you live in the valley and benefit from the US venture ecosystem you should be highly skeptical of open weight models. Banning them may very well be the best course of action.

    • danny_codes 3 hours ago

      Hilariously bad take. Open weight models can be retrained of fine-tuned, that's the entire point. The idea that the "Chinese government controls your intelligence" is laughable in the case of open weight models. Once the weights are released you can do whatever you want with them. The idea that there's economic capture by the Chinese for products they're literally giving away is stupid to the point of inanity.

      I can only assume this account is pure shilling for the closed-source AI labs.

    • pianopatrick 3 hours ago

      "Open source" open models can also survive if they are seen as a necessary cost of doing business. Same logic as tech companies working together on Linux. "We need an operating system. But an operating system is expensive to make on our own and does not really provide an edge. So let's just work on and use the open source one."

    • noncoml 3 hours ago

      Open Source is free as in speech

      Open Weights is free as in beer