Xiaomi MiMo v2.6

(mimo.xiaomi.com)

415 points | by volf_ 3 hours ago

39 comments

  • rao-v 3 hours ago
    I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

    The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

    If you’re releasing an open model going forward, please consider offering the community more of this transparency!

    • MangoCoffee 1 hour ago
      maybe this is why Dario want to slow down AI development and all the big AI labs in the USA is singing the same song.

      whey they all singing the same tune. it make me question what is their real motives.

      they are afraid of Chinese good enough LLM model killing their margin. we already have story about US companies switch some task to use cheaper Chinese model hosted on Neoclouds.

      • aeyes 1 hour ago
        The reason is money. They want regulation to make it harder for new competitors and competitors from other countries.

        They invested billions into training the models but there is no competitive advantage, we see that within a couple of months everyone catches up. There is no way to profitability unless they get some policies to shields them against competitors that can't comply with the regulatory requirements.

        That is also why there are things like Claude, Codex and Cursor. They are trying hard to build a customer relationship with a higher switching cost that hopefully sticks.

        But the problem is that the AI buildout has become a large percentage of GDP. So obviously the government wants to keep it going because these companies are pumping enormous amounts of money into the economy.

        • jamienk 10 minutes ago
          I’m still not at all sure about the “billions” invested claim. How much of that is cloud running the models? How much is pre and post training (which may or may not be part of what we’d want to include in accounting). Etc. Does anyone have links to good reporting about this: not blind recitations of numbers, but analysis and thought mixes with investigation?
        • lelanthran 25 minutes ago
          > But the problem is that the AI buildout has become a large percentage of GDP. So obviously the government wants to keep it going because these companies are pumping enormous amounts of money into the economy.

          They are pumping enormous amounts of money into each other. Hardly any of that is making its way to people, it's all going to highly automated construction and to energy use.

          Seriously, how many jobs did the $1t in venture capital fund?

          • gunalx 14 minutes ago
            If I pay you 100$ for mowing my lawn, and you me for yours. Technically the GDP increased with 200$.
      • jwolfe 1 hour ago
        Please explain how putting an upper bound on how good the strongest models can be prevents cheaper less strong models from catching up, rather than enabling it. I do not understand this argument at all.
        • rbjorklin 1 hour ago
          The general idea is that Anthropic/OpenAI is pushing this narrative as an attempt at "Regulatory Capture"[1] which would allow them to make it prohibitively expensive for anyone but them to enter the market thus stifling competition.

          * 1: https://en.wikipedia.org/wiki/Regulatory_capture

          • verdverm 1 hour ago
            I heard someone analogize token vendors to car manufacturers, where American companies only want to produce expensive options, the people want cheaper/better alternatives, and we ban BYD because those with enough money are more "persuasive"
            • juiceland 54 minutes ago
              The analogy is a good one, but your explanation is missing one aspect: the country (USA) does have a reasonable interest in having the capacity to build their own models. The “we need to slow down because it’s getting too dangerous” part is probably more related to “we need to slow our public facing development down so the US government can get the best and the American corporations can trickle out what we decide is safe”

              It’s similar with cars. It’s not that American cars are better than Chinese cars on any tangible measurement. But America already shipped most of its manufacturing overseas. Everyone who built those factories is retired. The US should probably hold on to some capacity to make cars, seeing as their entire infrastructure depends on them.

              • verdverm 47 minutes ago
                American Ai/Car manufacturers could build cheaper/open models, some do, the big ones do not. It's not an either or, but a spectrum where they have chosen to build only in a subrange
                • juiceland 36 minutes ago
                  It is the natural result of a country run by lawyers. China is a country run by engineers.
                  • verdverm 29 minutes ago
                    I think it less about lawyer vs engineers and more about money in politics (now unlimited)
          • chanakya 1 hour ago
            How would that slow down the Chinese models, given that the US has no regulatory reach in China?
            • girvo 1 hour ago
              You target the US companies: if they can't use these Chinese models, then they're less of a danger for a now captive audience in the US (and the West generally).

              This is already kind of the case: the big enterprises don't really want to touch the latest Chinese models. It's a real pain, personally, I want to use them at work!

              • pimeys 47 minutes ago
                Show them you can burn tokens in seven sessions day and night with comparable results to Opus with less energy and less than 10 dollars a day, per dev.
                • girvo 34 minutes ago
                  We have. Unfortunately there are political realities that get in the way, and Bedrock for example doesn't have GLM 5.3 (Flash or otherwise) or anything new/useful

                  I do imagine it'll change, but it hasn't yet.

              • verdverm 53 minutes ago
                1. China is a bigger market than the US for Ai, they are on pace to process 100Q tokens this year, roughly the same or more than the US big companies

                2. Enterprise trends are towards open weights, several routers and vendors now have more than half the volume going towards open weights

                • girvo 33 minutes ago
                  Yes, but thats not something a company engaged in regulatory capture for themselves care about: especially if they're worried they'll be outpaced and overtaken by the Chinese labs. Which they will be, IMO.
                  • verdverm 30 minutes ago
                    they care because they know it unlikely open weights will be banned, and thus available to American companies, with regulatory capture (onerous requirements) being a "good enough" "ban" that their big models don't face real competition, regardless of the open weight origin. American companies make open weights too, they are equally threatening to Big Ai financials.
            • RussianCow 1 hour ago
              Because the end goal is to ban non-US AI companies from being able to do business in the US.
            • verdverm 1 hour ago
              it wouldn't slow down China as much as make it impossible for American companies to use non-American options, they care about their margins and don't want to be commoditized
            • the_sleaze_ 1 hour ago
              [flagged]
              • pixl97 1 hour ago
                OK, so how does this help the US?

                If the US slows down this may lead to people that would have went to US labs to go to other countries.

        • lelanthran 23 minutes ago
          > Please explain how putting an upper bound on how good the strongest models can be prevents cheaper less strong models from catching up, rather than enabling it. I do not understand this argument at all.

          They are not proposing to regulate only the strongest models. They are proposing to regulate all models. If they are already on top, regulation may stop them from proceeding further, but it also stops the cheaper alternatives from catching up.

          If they feel they have reached the asymptote of the curve, then regulation doesn't affect them, it affects those who have yet to reach the asymptote.

          • cogman10 2 minutes ago
            Particularly, the route they seem to want to go is "safety".

            My guess is that Anthropic and OpenAI will push for "safety" regulations which require byzantine testing that, shocker, Anthropic and OpenAI can pass but the chinese models cannot. The route they'll take is import bans and potentially even general bans on products producing or using "unsafe" models.

            They'll further likely try and push AI "safety" treaties from the US to other nations to further lock in their lead.

            That's why, IMO, we've been seeing so many "OMG, AI will destroy the world and these AI researchers are so scared" articles.

        • lytedev 1 hour ago
          I don't think "putting an upper bound" was OPs phrasing?
          • jwolfe 1 hour ago
            That's what pacing the frontier is, and is what the labs are pushing for.
        • bellowsgulch 1 hour ago
          That’s not the argument.
          • jwolfe 1 hour ago
            Please elaborate on what the AI labs are specifically requesting and how that results in slowing down Chinese model progress below the frontier.
      • ed_balls 33 minutes ago
        Does anyone know what are the proposed regulations? Controlling software is impossible, so the only option is banning hardware ownership. No more mac studio.
      • Pxtl 1 hour ago
        Dario has always wanted the AI development to slow down and be more careful. Safer AI development was a core reason that Anthropic split off from OpenAI.

        What's different today is that now all the big LLM firms want to slow down AI development. When men like Musk and Altman (both known for habitually shooting their mouths off and saying whatever they need to whoever needs to hear it regardless of truth) suddenly agree with Amodei, that's when things start to smell off.

        • verdverm 1 hour ago
          > What's different today is that now all the big LLM firms

          not all, just a few American ones (~PayPal Mafia + Google), there are other big American LLM developers (notables include Nvidia, Meta, and Palantir) that do not agree

    • earthnail 2 hours ago
      Thanks so much for sharing this. As someone who mostly watches from the sideline, can you share what you can see in this dashboard that someone like me can't see? Is it the metrics themselves that they measure (the metrics tab is absurdly detailed), something in the notices, or something else I missed?
      • rao-v 1 hour ago
        I might turn this into a blogpost if folks are interested, but my god there is so much clever info in that dashboard.

        Here is one really neat bit:

        A cutting edge training idea (for agents, it's been used elsewhere for ages) is on-policy RL, basically, it's not enough to say "here is an end to end agentic sequence (including tool calls etc.) that is perfect" you want to say "here is a sequence you might actually have generated that turns out to be correct".

        Basically, it's more training efficient to improve models with small tweaks to do more of the right thing they are already doing sometimes than from some perfect oracular "this is the way" answer.

        (if you've ever tried to teach humans new skills, you’ve probably noticed this too!)

        When you do that, you care about how far the model you are updating (improving) has deviated from the one being used to generate rollouts (agentic rollouts for hard problems can take hours with lots of tool calls, so you can't keep redeploying every slight improvement).

        Lo and behold, the dashboard literally has:

        partial/avg_staleness (likely the measure of how many micro iterations the "generate answers" model is behind the "improving based on the occasional right answer" model)

        train_infer_diff/new_infer/kl (a more direct KL divergence based way of measuring how differently the two models generate tokens)

        How cool is that?!

        And don't get me started on the clever ideas hiding behind dynsam/avg@n ...

        • pimeys 38 minutes ago
          I would really enjoy that blog post.
        • jeffmcjunkin 1 hour ago
          I'd read the heck out of that blogpost. You have my interest.
        • oceansweep 33 minutes ago
          Please do!
        • dgellow 1 hour ago
          Please do
      • tancop 2 hours ago
        The best thing they did is being open about all the setbacks they had to deal with. They logged every restart with a reason, talked about dropping a cyber dataset after it degraded coding benchmarks. Also published real time training loss, benchmark scores after every checkpoint and running cost estimates.

        Really the only thing missing was dataset descriptions, the dashboard only had random IDs like "dataset-zrso". I guess it's their lawyers fault.

      • verdverm 2 hours ago
        the existence, who else has a live dashboard for the RL late-training?
    • ignoramous 1 hour ago
      > got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores

      Xiaomi MiMo is led by Luo Fuli, a former Alibaba & DeepSeek employee. Perhaps it is due to Luo just how similar Xiaomi's tech & GTM approach is to DeepSeek's.

      - How Luo Fuli Keeps an Earthy Touch as she Soars Through the AI World, https://newsen.pku.edu.cn/news_events/news/people/15385.html (https://archive.vn/I8Pmu).

      - Luo Fuli, the 30-year-old ‘AI genius girl’ behind DeepSeek’s success?, https://e.vnexpress.net/news/tech/personalities/who-is-luo-f... (https://archive.vn/sb3B6).

      • pimeys 36 minutes ago
        Open tech is cool. Speeds up all progress...
    • kingstnap 2 hours ago
      [dead]
  • margorczynski 12 minutes ago
    China will most probably win the AI race in the long run because of one major bottleneck the US has - energy. The electric energy and grid buildout in China has been massive since a long time and there is simply no way for the US to quickly catch up.

    No matter how much cash you throw you can't just materialize a 100 nuclear reactors to power the data centers.

  • lwansbrough 2 hours ago
    Anyone else more excited about Chinese models than American models these days? Big thing for me is affordability.
    • tacomagick 2 hours ago
      Absolutely! Chinese models are both cheaper and more capable in many cases, compared to the American models and their makers continuously fumbling or reducing model capability with each update. Deepseek decreased costs when they released Flash 4.1 you would not see any American company do this, in reverse they would try charge you more.
      • user43928 2 hours ago
        OpenAI decreased prices with the 5.6 model family.

        And later they further cut Sol and Terra pricing by 20% (maybe only in the API) and Luna by 80%.

        In fact Luna still outperformed DeepSeek Flash 4.1 in cost per task on Artificial Analysis when I last checked.

        However, Luna is slightly less intelligent. I have a feeling that it's pretty dumb and prone to hallucination unless running at xhigh or max effort, where it somehow manages to work quite well.

        I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.

        The competition is great, and I hope Chinese models will continue to force leading US labs to offer models at a low price point.

        That said, I don't think the Chinese labs have anything over OpenAI and Anthropic when it comes to capability or efficiency - I have no reason not to believe the US labs have even lower cost to serve the models.

        • tacomagick 2 hours ago
          OpenAI had to cut costs because of Anthropic. I also do not trust the benchmarks when it comes to models anymore. I have tried both Claude and OpenAI models and while it is true that the 5.6 series is smarter than Deepseek (at the time i tested it against 4.0) at that price it is still not worth it and sometimes randomly refuses to do tasks or stops midway etc.

          Do also remember China is this far in the AI race despite all chip restrictions from America. If they were in equal standards I truly think Chinese models would have long surpassed American ones. Also would like to remind how Anthropic CEO is being hostile and blaming Chinese models with distilling meanwhile their own models claimed to be Qwen¹ and their stance against open models is negative² and they still keep blaming China for it.

          1- https://news.ycombinator.com/item?id=48671252

          2-https://www.anthropic.com/news/position-open-weights-models

          • goosejuice 1 hour ago
            > Also would like to remind how Anthropic CEO is being hostile and blaming Chinese models with distilling

            Why wouldn't he? If there really was 25,000 accounts breaking ToS any CEO would at minimum be upset. Evidence of Claude distilling qwen would be damning but that a) makes no sense b) doesn't exist afaik.

          • user43928 1 hour ago
            Not sure about that.

            Given the difference in compute, it seems plausible.

            However, the researchers at the US labs are surely no less talented, and they have better access to hire talent globally.

            They too have to serve their models efficiently at a large scale, and with current capacity constraints this must be a top priority.

        • Implicated 1 hour ago
          > I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.

          So you don't have much perspective on things, it seems. Let me introduce you to the GLM 5.2 and then 5.3/5.3 flash series of... "oh, wow, I should have bought some RTX PRO 6000's while they were 'cheap'" stage of progression.

          As someone carrying multiple max subscriptions to both claude and codex - primary workhorse is glm 5.3 flash running on rented GPUs for less than a latte/hr.

          I also found qwen 3.6 27B nearly useless for my own needs. DS4 flash 0731 and then 4.1 have been nearly as eye opening as glm 5.3 flash, but have their own warts.

          • user43928 20 minutes ago
            Why use GLM 5.3 Flash when you also have access to Astra, Sol, Fable?

            Or I guess the other way around, if GLM 5.3 Flash is so good, why Claude and Codex?

          • CamperBob2 19 minutes ago
            Try DS4.1 Flash. It's another eye-opener. If you run it in Claude Code, it's easy to forget you're not actually talking to a high-end Opus model.
      • goosejuice 2 hours ago
        > Deepseek decreased costs when they released Flash 4.1 you would not see any American company do this, in reverse they would try charge you more.

        OpenAI reduced prices and Anthropic increased weekly usage limits.

    • joshheitzman 1 hour ago
      Absolutely! DeepSeek-V4-Flash-0731 has become my daily driver. It's pretty amazing what it can do for what it costs at deepinfra.com (I don't use deepseek as a provider since they train on your data [at least their honest about it]). GLM-5.1 was my daily driver before that and Kimi K2.5 before that.
      • tristanMatthias 1 hour ago
        How does it compare to 4.1 flash? Curious why folks don’t use the more “modern” one.
        • randbyte 3 minutes ago
          4.1 flash is very fast and capable. Token efficiency is not great so it fill up context window much faster compared to similarly capable models.

          glm 5.3 flash is a tad slower but a bit more capable and way more token efficient.

          Source: self hosted tested on rented GB200 node at 8bit.

        • joshheitzman 34 minutes ago
          I haven't tried 4.1 flash as I'm assuming its a preview. I did not get good results from the preview version of 4.0 flash (i.e. the one that did not include the month and date of release in its name).
          • CamperBob2 18 minutes ago
            4.1 Flash is a horse of a very different color. It cooks. IMHO it's probably a preview of DS5, rather than a true DS4-series model.
      • kingforaday 1 hour ago
        Are you finding DS better then kimi k3 and glm-5.3? Do you mind sharing your primary use case?
        • pimeys 33 minutes ago
          I've used Kimi K3 for a few months as my main model and DeepSeek 4.1 is as fast and about 10x cheaper.

          I just had like four big sessions going today, paid about $8 in tokens. I see no reason to pay more, this is more than I need for intelligence.

        • joshheitzman 58 minutes ago
          My primary use is AI coding agent. Its vastly cheaper than Kimi K3 and I haven't found a scenario where I really need Kimi K3 versus smaller models. GLM-5.3 Flash is good but there is series of bugs in the vllm middleware that prevent GLM models from getting all of their reasoning content returned to them that impairs inference quality. A lot of inference providers use vllm which makes it hard to find a good provider for GLM. I've been using friendli.ai but using GLM-5.3 Flash from them is more expensive then using DS V4 Flash from deepinfra.com simply because deepinfra.com is so cheap. The DS V4 Flash cost at together.ai is similar to the GLM-5.3 Flash from friendli.ai or at least that's what I found in my benchmarks a week ago: https://www.linkedin.com/posts/joshheitzman_i-ran-a-fuller-r...
    • solarkraft 56 minutes ago
      I couldn’t tell you what western model I was last excited about. Probably Glimmer.
      • verdverm 43 minutes ago
        Jev seems to have people excited, I'm more excited for the Kevs
    • SyneRyder 2 hours ago
      Yep, I'm trending in that direction, and I'm someone with Claude stickers all over my laptop. My main app dev work is still going to Claude, but everything else is going to China even at API rates now.

      One simple task: I needed an LLM to go through and clean up a few thousand page descriptions and titles in my personal search engine index, where the human web page authors had put in no effort sigh. I did a shoot out between Claude, Luna, GLM 5.3 Flash and Deepseek. Despite the high cost, Claude's descriptions were terrible, and even Opus warned me that the descriptions coming back from Haiku were "generalized, not accurate". I expected I would choose Luna because of price, and occasionally it did have wonderful descriptions (one captured emotion in a way no other model did). But in the end, the GLM 5.3 Flash descriptions were the easiest to read, they flow well while also being accurate & including necessary keywords, and being highly affordable. So it won out. It's a task that is nowhere near frontier, but a task where somehow China is better than frontier.

      • rapind 1 hour ago
        API rates still aren’t quite competitive with the OpenAI x20 accounts, but they are definitely getting close with deepseek 4.1 flash. I spent a few days with only 4.1 and was very impressed.
    • verdverm 2 hours ago
      I have a contrarian opinion that China passing America in Ai is the Sputnik moment we need to leave the hubris behind and get our mojo back

      debatable if a turn around is possible before '29

      • machomaster 40 minutes ago
        The analogy makes little sense. The USA was not in front of the USSR and Sputnik merely showed that. It is at this point that the Americans woke up, put a lot of effort and finally were able to surpass the Soviets during the Apollo missions.

        China was never ahead of the USA in AI. So perhaps a more proper analogy is the Moon landing. In real history the side that lost the race never got its mojo back...

        • verdverm 35 minutes ago
          I'm looking forward, towards the future, when I use "passing ... we need", need being key here as it implies something we don't yet have

          I expect this to happen within 12-18 months, the differentiation has shrunk, many models are now sufficiently capable for most tasks

          • machomaster 26 minutes ago
            I understood that.

            I was simply saying that when (not if) Chinese AI models will pass Americans, it will probably be game over and Americans will never catch up, let alone become leaders again.

            Check the names of the researchers in the DeepSeek's latest paper. Full of Chinese names. Check the list of names in Google's paper. A very similar view. Anecdotal, but quite thought-provoking...

    • bellowsgulch 1 hour ago
      Yes, an expensive American LLM has zero capabilities as far as I’m concerned because I’m never going to pay for it.
    • swingandamiss 2 hours ago
      No, because I'd rather not support our economic and military rivals.
      • lwansbrough 2 hours ago
        I'm Canadian so this sentiment has little value in 2026 unfortunately.
        • ActionHank 2 hours ago
          Also, frankly, as a fellow Canadian it's pretty clear that the biggest "rival" the US has right now is itself. Just passed out in the corner puking on itself shouting about all the foreigners who won't talk to it.
        • zemvpferreira 1 hour ago
          As much as the US has been easy to hate lately, I don't hesitate to say Xi Jinping as the most powerful man on Earth would be much, much worse.
        • tancop 1 hour ago
          I'm from Europe and I hate America way more than China now. Used to be about equal but then Trump started extorting Ukraine, threatening their own allies and sending billions to Israel to help with a genocide. I think that exposed America for what it really is.
          • boelboel 1 hour ago
            China is enabling russia way more than trump, China doesn't care too much about 'morals' either. Chinese companies have been quite important in the construction sector of the WB settlements. Even though I'm not a great fan of Trump I don't see a reason at all to prefer the chinese.
            • machomaster 38 minutes ago
              China is a somewhat neutral player, supplying both Russians and Ukrainians. Their attitude and action is way less one-sided than Trump's; especially in the first year of his latest presidency.
              • boelboel 29 minutes ago
                With trump his actions being one sided you mean one sided towards ukraine? They still get lots of Intel from Americans and Americans hardly but anything from Russia. But you're right that china supplies both I wouldn't exactly call that neutral as much as just in their self interest.
            • SSLy 40 minutes ago
              buy inference from european companies running open chinese (or that one from google) models
            • peterashford 40 minutes ago
              As a New Zealander, I would agree - no reason to prefer the Chinese. But Trump's America is not an attractive option either and there's no reason to prefer it. And given the choice between two ugly options, the rational choice is the cheaper one, surely.
        • scottyah 2 hours ago
          [flagged]
          • lwansbrough 2 hours ago
            Because at present the pedophile US president is making it his mission to molest my country. China, for all its faults (including espionage, which the US is also guilty of) is mostly focused on conducting trade.
          • verdverm 2 hours ago
            Half of Canada now uses the word 'enemy' when asked for an adjective to describe America or China. We're equivalent in their eyes now because we elected Trump a second time and all that he has said and done in 2.0
            • cwillu 1 hour ago
              It's closer to a cousin you used to be close with despite some moral failings, but who has now has a substance abuse problem and is lashing out at family and friends.

              Not an enemy, just a danger.

        • rayiner 2 hours ago
          Canadians warming up to China makes me think of Germany becoming increasingly reliant on Russia in the 2010s.
          • rapind 1 hour ago
            Murica just has a MAGA problem. We can still be friends if and when you sort that out. Us Canadians like most of you quite a lot.
          • cgio 1 hour ago
            Yes, someone can still blow up a pipe and they look the other way. On the other hand, you can also draw parallels to themselves becoming increasingly reliant on US vs UK in the past.
      • joshheitzman 1 hour ago
        Does it count as supporting a rival if your an American using an American inference provider self-hosting an open weight model from a Chinese lab?
      • Freedom2 2 hours ago
        Agreed, and also because I support freedom of speech!
        • girvo 2 hours ago
          Neither the US nor the Chinese companies are on your side then. They both censor, just different topics.

          But at least I can run Chinese models locally, and strip a lot of that censorship/refusal.

        • peterashford 37 minutes ago
          As long as that speech doesn't come from CNN or criticise Charlie Kirk, Israel or Trump? I'm sceptical about how much the US really values free speech
  • simonw 2 hours ago
    • phainopepla2 2 hours ago
      I think we can say pretty confidently they aren't pelican-bench-maxxing
    • written-beyond 1 hour ago
      can you update this website, I just wish the entire layout wouldn't shift when the page gets loaded and the timestamps in the title look very ugly and take up a lot of space.
    • brcmthrowaway 2 hours ago
      Just me, or do these look bad?

      Qwen3.8-27b pelican was amazing on Mac.

      https://www.nudgehost.com/dpjn3uwe

      • knicholes 47 minutes ago
        Two legs on one side is a little sus.
      • idiotsecant 2 hours ago
        Looking terrible isn't nessesarily a bad thing. The pelican is heavily pre trained now. Having a crappy pelican means you didn't try to juke the stats.
        • broodbucket 2 hours ago
          Apologies for not taking the time to find it, but there was a post that tried to determine if the pelican was benchmaxxed across a bunch of models by comparing it to other SVGs, and found that it wasn't at all.
    • handfuloflight 2 hours ago
      How does this translate to coding performance, which is what most of HN cares about (...I assume)?
      • simonw 2 hours ago
        It means they're good at writing SVGs, in particular SVGs of animals riding modes of transport!
      • lanyard-textile 2 hours ago
        I only visit HN for the pelicans, personally.
    • Imanari 2 hours ago
      ish… at least we can be sure they don’t benchmaxx the pelicans lol
  • stymaar 3 hours ago
    Flash[1]: 309B total / 15B activated parameters

    Pro [2]:, 1.02T total / 42B activated parameters

    [1]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL

    [2]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL

    • verdverm 2 hours ago
      • gandreani 2 hours ago
        Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model?
        • mydreamof 2 hours ago
          It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data
        • simonedepertis 2 hours ago
          [flagged]
    • verdverm 2 hours ago
      curious why the HF pill (on the right) always has inaccurate values
      • bopbop9876 44 minutes ago
        I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants.
      • stymaar 2 hours ago
        I noticed the same, and I wonder as well.
        • verdverm 2 hours ago
          I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
    • segmondy 2 hours ago
      more like 500B in FP8
  • nemothekid 3 hours ago
    Looking at the frontend design examples; why do these models seem to love the "01 - UPPERCASE TEXT" motif. It's everywhere now (see https://try.cloudflare.com/, which has '01 · QUICK TUNNELS', but no "02" anywhere).
    • danvayn 2 hours ago
      My guess is that by function they break down frontend sections or components into pieces and I believe document things for themselves on some level, or purposely are verbose in this way. It is probably also shaped by users and existing web patterns. They probably get reinforced by models the more common they become.
    • pphysch 2 hours ago
      The extraneous small-caps labels are one of the main idiosyncrasies of AI generated markup. I wonder how much of this is a "scaffolding" technique to help the model build stable designs. But was it reinforced in RLHF or an emergent behavior of the models?
    • sandblast 3 hours ago
      Nice catch!
  • user43928 2 hours ago
    I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.

    Maybe Terminal Bench 4.0 and ExploitGym are reasonable.

    Terminal Bench 4.0

      GPT 6 Astra             59.6
      Claude Fable 5.1        55.1
      Claude Opus 5           49.0
      MiMo-V2.6-Pro           34.9
      MiMo-V2.6-Flash         28.8
      DeepSeek V4.1 Flash     26.8
      MiMo-V2.5-Pro            1.5
    
    ExploitGym

      GPT 6 Astra             42.4
      Claude Fable 5.1        30.4
      Claude Opus 5           22.1
      MiMo-V2.6-Pro           17.8
      MiMo-V2.6-Flash          6.0
      MiMo-V2.5-Pro            0.1
    
    DeepSWE v1.1

      DeepSeek V4.1 Flash     74.2
      Claude Opus 5           74.0
      GPT 6 Astra             74.0
      MiMo-V2.6-Pro           71.9
      Claude Fable 5          70.0
      MiMo-V2.6-Flash         67.9
      MiMo-V2.5-Pro           19.0
    • mokre 2 hours ago
      Maybe you should not trust any of the benchmarks!
    • dom96 1 hour ago
      Why not? In my own benchmark Opus 5 does in fact come out on top[1]

      1 - https://bench.killswitch-lang.org/

      • user43928 19 minutes ago
        Good question, maybe I am underestimating it based on its absolutely horrible writing style.
    • varispeed 2 hours ago
      They match my experience. Astra and Fable I rate below Sonnet. They are incredibly poor. They were excellent for a couple of days after release and then plummeted.

      Maybe I am being routed to more quantised versions or less capable models with system prompt to fake Astra or Fable.

  • jjcm 37 minutes ago
    Here's an image->html test for it using 2.6 Pro Ultraspeed, along with comparisons for grok 4.7 and Astra.

    Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

    MiMo 2.6 Pro Ultraspeed (36min): https://html.non.io/annui-mimo/

    Grok 4.7 (25min): https://html.non.io/Annui-grok/

    Astra (19min): https://html.non.io/annui/

    Overall this felt like the weakest of the three. Ultraspeed was fast as far as tokens per second goes, but it overthought quite a lot of things resulting it in having one of the longest build times. That overthinking didn't lead to better results either - note the statue with the cropped off arm. It's also the worst implementation of the dynamic lighting effect / displacement effect - the background especially has some significant distortion. Astra was the only one that seemed to understand that displacement should happen less the further something is in the distance.

    Here's a vid of all 3 side by side with the source design: https://non.io/video/annui-comparison.mp4

    • faitswulff 35 minutes ago
      Astra's looks the worst to me on mobile, though
      • jjcm 33 minutes ago
        That's fair - worth noting that none of them were instructed to make a mobile variant or to test the mobile size.
  • toephu2 1 hour ago
    I said this years ago, LLMs are a commodity (or were becoming one at the time). They are dime a dozen. Even the frontier ones. OpenAI and Anthropic have no moat.

    No moat and competition is good for consumers though.

    • boelboel 1 hour ago
      No moat and competition isn't always preferable over a competitive oligopoly (with some differentiation). The former ends up with politicians intervening way more like with solar, steel, agriculture ....
    • Marciplan 1 hour ago
      [flagged]
  • vatsachak 3 hours ago
    Wow, the chinese labs are getting good at advertising model releases. The moat is thin.

    Some features of the release I like:

    - Demonstration of diverse tasks, such as using a DAW

    - Graphs from various benchmarks and price ranges

    - Real world use of the model in scientific environments

  • volf_ 2 hours ago
    I've got a working recipe to run this model on Dual DGX Spark: https://github.com/volfco/spark-vllm-docker/blob/main/recipe...

    Averages ~25-35tok/s which isn't bad for a first attempt.

  • wkcheng 41 minutes ago
    Are people using MiMo models as their daily drivers in a company setting? If so, how? I know they're available over OpenCode and directly from Xiaomi, but those are not great options. OpenCode Go straight up doesn't give any guarantees about training on your data, and Xiaomi says that they won't but it's unclear.

    With some models you can find hosting companies based in the EU or US, but then you don't know how they're quantizing the models, so you're not sure about the actual output quality.

    How are people actually using this? Or are people just experimenting with side projects?

    • ricardobeat 4 minutes ago
      I only use it for personal projects, but their european Token Plan [1] has ZDR and is hosted in Amsterdam. The DC also likely runs on 100% solar wind/power, with heat recovery systems used to heat nearby university buildings.

      [1] https://platform.xiaomimimo.com?ref=UKV2FC (invite link = 10% off)

    • pimeys 19 minutes ago
      Fireworks is amazing for DeepSeek, Kimi and GLM.

      Hope they bring MiMo for tests.

  • GodelNumbering 1 hour ago
    Mimo has been one of those models that I have been rooting for since the first I used it, the 2.5 pro which I have used quite a bit, was very concise, very aware of how much context needs to be read for which tasks and would always keep the context tight. Also surprisingly good at strategic thinking. I had published a comparison between it and Terra where Terra was found to be using much more avg context for similar tasks https://dirac.run/posts/gpt-5-6-vs-mimo-2-5-pro-context-bloa...

    Interesting but not surprising trend across the board seems to be, the flash models seems to have caught up with the pro-sized models of H1'26. No surprise all labs are rushing to bigger models.

    EDIT: Wow, took a detailed look at the benchmarks. Mimo 2.6 pro, the 1T model leads Kimi K3, a 2.8T param model in 14 out of 15 benchmarks (and the last one is near tie)!! Good to see they also kept the price the same, and landed in the greenest quardrant of the intelligence vs speed of AA.

  • dom96 42 minutes ago
    Very capable model. I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.

    KillSwitch-Bench 1.0

      Claude Opus 5           66.9
      GPT-6 Astra             57.9
      Claude Fable 5.1        46.7
      MiMo-V2.6-Pro           38.8
      Muse Spark 1.3          36.5
    
    1 - https://bench.killswitch-lang.org/
  • bonsai_spool 41 minutes ago
    Are folks working in biology / cybersecurity seeing more limits in what Opus (not Fable) is allowing? This has happened quite suddenly for me and I’m stuck in the middle of a project that would have otherwise called for use of Claude.

    I’ll be trying these models out and may end up switching my subscriptions if this craziness continues

  • thrownawaysz 2 hours ago
    >Night 0.8x Usage, 00:00-08:00 -UTC+8

    It's because offpeak electricity is cheaper?

    Funnily it's perfect if you are in the Pacific Time Zone because you can use it daytime 9am to 5pm

    • eunos 36 minutes ago
      Same as DeepSeek, non busy time for UTC+8, maybe also cheaper electricity during night
    • SSLy 39 minutes ago
      I believe it's a function of their primary user base being in china
  • syntaxing 3 hours ago
    All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
    • verdverm 2 hours ago
      • trvz 2 hours ago
        That’s for toy GPUs, like the 5090.
        • verdverm 2 hours ago
          there are many tasks (increasingly more each day) where small models are more than enough
    • brcmthrowaway 2 hours ago
      Is there a gamechanger around the corner to reduce DRAM requirements?
      • zozbot234 2 hours ago
        You could always stream from SSD storage. Especially effective if you get a cheap old-gen HEDT with lots of PCIe slots to add NVMe storage to and reasonable overall PCIe bandwidth.
        • jkingsman 1 hour ago
          That nearly certainly boots you to secs-per-tok land (as opposed to tok/s). Plausible if you are willing to wait hours to days for responses for simple testing, but not (debatably) "usable".
      • stymaar 2 hours ago
        n-gram per-layer embeddings[1][2] might be it.

        [1] https://sebastianraschka.com/llm-architecture-gallery/per-la...

        [2]: See DS 4.1-Flash and Qwen-3.8-Next.

        • verdverm 2 hours ago
          this is to offload VRAM to DRAM (for GP comment), and makes no difference for URAM
          • zozbot234 2 hours ago
            You can definitely offload n-gram embeddings to storage; they're very sparsely used (only a few KB fetched per token) so this is quite effective. Loading to DRAM only becomes necessary if they are a bottleneck to overall performance (which might happen if you're doing very wide batches and everything else uses super fast VRAM/HBM).
            • verdverm 2 hours ago
              I was looking at the qwen-next-flash, and the weights would fill my OEM Spark on their own, before the n-gram. I'm unclear if offloading to disk can work here, is that what you are implying is possible?!
              • girvo 2 hours ago
                Check out eugr’s TP=1 sparkrun recipe :)

                It’s an NVFP4 quant, but it fits, and is surprisingly capable.

                • verdverm 2 hours ago
                  do you have a HF link? HF search is not uncovering it for me

                  (or is it somewhere else)

                  • girvo 2 hours ago
                    https://github.com/spark-arena/eugr-recipes/blob/main/recipe...

                    This one!

                    I'd recommend pointing your agent at it (after installing sparkrun), and asking it to research the absolute latest in TP=1 Flash-Next - mine grabbed particular vLLM nightlies and mods to improve performance, and it was well worth it.

                    • verdverm 1 hour ago
                      I have a quirky vLLM on k8s on 2x OEM sparks setup with 9 models available to me. I'm not keen to run nightly vLLM, too many issues with it in the past. Going the qwen-next path means displacing things I use daily :/

                      I have a watchful eye on the diffusion ~ Jev/Kev PR

                      https://github.com/vllm-project/vllm/pull/57250

                      • girvo 1 hour ago
                        For what it's worth, Flash Next outperforms every other model that is available to us on the GB10 in all of my testing; though if you have two sparks then the TP=2 version is even better and easier (I don't think you'll need the nightly for that at all, just use the recipe)

                        I'm so tempted to buy a second one...

                        • verdverm 1 hour ago
                          prices have gone up quite a bit...

                          I'm running embedding, reranking, and policy tuned models too, and a Jev/Kev when that's landed. Flash Next is not a substitute for those

                          I have OpenCode/Fireworks to access big models

                  • verdverm 2 hours ago
          • girvo 2 hours ago
            Nah, I’m streaming ngrams off NVMe on my Spark-alike right now. Works surprisingly well (except for when I accidentally bottlenecked it through my NAS)
            • jkingsman 1 hour ago
              What kind of throughput do you see on what models?
              • girvo 1 hour ago
                GB10 boxes have way more compute than they have memory bandwidth, which nicely fits medium sized MoE models with speculative execution (MTP, DSpark/DFlash, etc)

                Qwen 3.8 Flash Next (what I'm running basically entirely now) sees 30 / 35.0 / 45 tk/s for prose, analysis and code respectively for actual use (not short context benchmarking) with Pi. Thinking blocks are ~35tk/s or so.

                The GB10 having so much compute is great for prefill too, 2000-3000/s for 14k to 64k token prompts (cold cache too) in the quick benchmark I did. 3500tk/s for warm cache which is nice :)

                When I accidentally streamed my ngrams over the 2.5Gb/s network, it cut all the throughput down in half basically. Especially notable for the time-to-first-token, which is what clued me in that I'd messed up somehow!

                For Qwen 3.8 27B, I got it up to a consistent 20tk-25tk/s but 27B thinks so much that it was honestly too painful: Flash Next is as smart, as useful, but much faster for real agentic dev usage IMO

                Laguna S 2.1 saw similar numbers to Flash Next if I remember right, but their latest updates means it doesn't quite fit a GB10 128GB anymore at full context which is a shame.

                Note: these are all NVFP4 quants (usually a dynamic one where some tensor layers are left at full precision though)

                • verdverm 1 hour ago
                  I personally stopped caring as much about the tok/s as the agents are largely in the background, and so have also moved preference from MoE to dense

                  I want to see about fine-tuning these models a bit on the GB10 to tame that over thinking and some other behaviors (like using tools I don't use)

                  qwen 3.8 seems to have been trained with some `rkt` that messes with tool outputs to "save tokens"

              • verdverm 1 hour ago
                check out the spark arena website, its the raison d'etre
            • verdverm 2 hours ago
              interesting, peer comment seems to indicate this is a possibility as well, will have to take a deeper look
          • petu 2 hours ago
            n-grams can be kept on SSD, no need to hold them in any kind of RAM (at least w/o batching)
          • stymaar 2 hours ago
            Am I missing a joke? WTF is URAM?
            • verdverm 2 hours ago
              unified memory, not sure if anyone uses URAM, I human hallucinated it
  • stemlord 30 minutes ago
    Stupid question: in the benchmark diagrams I'm assuming the values are percentiles, so what does 100% represent?
  • pulkitsh1234 1 hour ago
    Anyone knows what they used to create the videos ? Is the model driving a program like Davinci Resolve / After Effects ? or is the model writing code to then generate these videos via some library.
  • drob518 1 hour ago
    Conspicuous that there’s no reference to GLM 5.3/Flash in the reported benchmarks. Just Deepseek and Kimi.
  • ddxv 3 hours ago
    This looks great in terms of cost and capabilities, truly pushing the frontier forward in terms of open weight light weight models.
  • eriquesito 2 hours ago
    Funny that all but one video has audio, the house 3D model one, where you can hear (what I assume are) Xiaomi's engineers talking about who knows what.
  • MisterMunchkin 2 hours ago
    I really liked MiMo 2.5, it was really affordable and actually had vision, unlike DeepSeek. (DeepSeek has only recently added it)

    Just tried 2.6 flash on a really niche topic I specialise in and it has done a really good job. They’ve definitely polluted their training data with claudeslop, but looking past the slop there is a decent model.

    • perrygeo 1 hour ago
      Can we afford to look past it? If/when claudeslop starts infecting every new model to such an extent, that model will produce its own slop, infecting new models... At what point do we lose all reliable methods for establishing "truth"? This is epistemic collapse waiting to happen. I honestly thought it would take longer... holding out for a coherent shared reality in 2030 seems optimistic.
    • omani 2 hours ago
      how do you recognize "claudeslop"?
      • Bluestein 2 hours ago
        It's an honest, load-bearing, simple thing.-
        • SSLy 38 minutes ago
          that's belt and suspenders too
          • pimeys 10 minutes ago
            a smoking gun
  • DanMcInerney 3 hours ago
    This is a big week. Probably getting next OpenAI and Anthro models, Grok 4.7, Mimo, etc. These open source model releases are why I can't take the "slow down" crowd seriously. I pitted older Mimo, qwen, step, gpt-oss, and other models against each other playing games like Werewolf and Sketch.io-like games where I let them talk shit while they played against each other. Mimo was by far pareto frontier of game-playing for the models that were <$0.15/m input tokens on OpenRouter. Qwen was pareto frontier in the shit talking game though. Qwen's hilarious. https://www.tiktok.com/@clankerfights/video/7642862917582425...
  • algoth1 3 hours ago
    Finally a lab that doesn't cheat on the charts
  • esafak 1 hour ago
    It tops the intelligence vs cost Pareto frontier and, uniquely for a Chinese model, does well in response time too.

    https://artificialanalysis.ai/models/mimo-v2-6-pro#intellige...

    That's pretty fast; I think I'll try it: https://openrouter.ai/xiaomi/mimo-v2.6-flash

    One concern I have is that they allegedly do not discount cached tokens: https://www.reddit.com/r/opencodeCLI/comments/1t37dz3/xiaomi...

    Can anyone comment?

  • alfalfasprout 2 hours ago
    The moat for OAI and anthropic seems to be very quickly shrinking. Chinese labs are now using RSI-like approaches and even without resorting to heavy distillation they're catching up in a couple of months vs. what would have been 6-12 months a year prior.

    And as these models get better the pace of training is quickly speeding up too.

    This doesn't bode particularly well for anthropic/OAI after they go public.

    • verdverm 2 hours ago
      token vendors are headed to the same place mobile data vendors went, this is good for everyone but those who thought they could maintain exorbitant prices
  • bertili 2 hours ago
    They mixed up DeepSeek 4.1 Flash with something else on this page, possibly DeepSeek 4.1 Flash means Gemini 3.8 Flash.
  • varispeed 2 hours ago
    These benchmark are useless as they don't say whether they were done before or after Fable and Astra got nerfed.
  • gigatexal 2 hours ago
    Leaning into what it cost to train is hilarious and an obvious shot at US frontier labs spending tens to hundreds of millions or more to train their models.
  • NooneAtAll3 2 hours ago
    does anyone know what unnamed model is on paretto frontier picture right between MiMo 2.5 and 2.6?

    so weird to acknowledge someone being on the front edge, but not name it

  • spwa4 2 hours ago
    As for the stats that everyone wants:

    MiMo-V2.6-Flash-310B-A15B roughly GPT-5.6 Luna / Claude 4.9 according to benchmarks MiMo-V2.6-Pro-1.02T-A42B roughly GPT-5.6 Sol / Opus 5 according to benchmarks

    Perhaps with IQ2 flash will run on 128G M5?

  • 16t96 2 hours ago
    [flagged]
  • nlcs 3 hours ago
    [flagged]
  • nofpu 2 hours ago
    [dead]
  • unpopularopp 3 hours ago
    [flagged]
    • InsideOutSanta 2 hours ago
      It's funny, I have the exact opposite reaction. This is probably misguided on my part, but Xiaomi is one of the very few major tech companies that I don't have an immediate strong negative reaction to. Everything I've bought from them, from robot vacuum to mobile phone, has been reasonably well designed, didn't break, and was priced fairly. I also think their car looks badass.

      I'm sure they're doing all kinds of terrible things, like all major companies. I just can't help but like them. Also, this model looks great, and I'll give their subscription a shot next month.

      • A_D_E_P_T 2 hours ago
        I must second this.

        I'm in Europe, and here the options for home appliances are usually German (e.g. Philips), Balkan (e.g. Gorenje), or Xiaomi. Xiaomi is the best by far, and it's honestly not even close.

        Their home appliances are so rock-solid that they actually still surprise me. For example: I've gone from having to replace electric water kettles every six months to buying one from Xiaomi and never replacing it. (Nigh on three years now.)

        I really have a very positive impression of them.

    • platinumrad 3 hours ago
      It's a big company, like Microsoft or Google. Some of their products are good and some are bad.
      • verdverm 2 hours ago
        ironic to this thread, I have less bloatware and ads since I switched from Verzion to Pixel on Fi (many years ago)

        Curious if Verizon / ATT still force apps on your phone, eg. NFL and Amazon apps, Fi service is subpar

    • algoth1 2 hours ago
      I still have a xiaomi mi 11 lite, my wife has a 15t. The cameras are the best for the price. The way they chove ads down your throat at every opportunity should be illegal though
    • bel8 2 hours ago
      Which model did you have?

      I ask because my wife has the 15T and the camera is better than my iPhone 17 Pro. And while toying around with it I didn't notice any bloat.

      Plus hers support native split screen which I kinda need to multitask on the go.

      I'm so pissed at how bad Siri is compared to her android phone that I'm thinking about selling the iPhone to get a Huawei Pura Ultra.

  • omani 3 hours ago
    ah, would you look at that. I was wondering why mimo 2.5 became "dumber" the last weeks. I was speculating they are probably about to release a new version of the model. because the model really acted out a lot. especially the last two weeks. dont know, was just a feeling, highly speculative.

    but now I got my "proof".

    • sandblast 2 hours ago
      I guess that would only be possible if your provider was Xiaomi itself?
      • omani 2 hours ago
        yes. I use opencode and opencode uses Xiaomi as a provider.
  • jwpapi 2 hours ago
    In the chart they use "Pareto Line", which I think is wrong. Pareto is 20% effort leading to 80% results. Which could be interpreted as models costing 20% having 80% of peak intelligence, but that’s not what it looks like to me.

    It looks like the "Frontier Line" to me, which is also often misinterpreted. frontier does not mean the best models. It means all models that are not strictly dominated, meaning in most cases: Not same price or cheaper and more intelligent.

    I personally would like the word frontier to be used with more criterias: Open Weights, per use-case, etc etc. This would make model selection easier, but I understand it’s not an easy thing to do.