Beam: Reflection's 501B open-weight model

(reflection.ai)

534 points | by Philpax 1 day ago

52 comments

  • Ariarule 1 day ago
    Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

    Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...

    • extr 1 day ago
      Yeah I remember when the original post about this came out. Def not recent. Though I think their point survives in that they didn't exactly RL on this.
      • glitchc 19 hours ago
        It could be old but still not be part of the training set.
        • stingraycharles 9 hours ago
          Yes, but then you can’t use “it’s 2 days old” to assert that it can’t be in the training set.

          Seems sloppy.

        • jiggawatts 23 hours ago
          I'd love to see these tests repeated for the current frontier models...
          • az226 15 hours ago
            Seems rather sloppy to not validate it wasn’t in their training or RL data.
            • charlieyu1 1 day ago
              I don't think age of the puzzle even matters, all models have search capacities these days
              • criemen 1 day ago
                > all models have search capacities these days

                one would hope that they disable websearch and internet access (maybe all tools?) when doing generalization testing?

                • zakisaad 22 hours ago
                  Model weights (what is being tested here) don't inherently "access the web" when inference is running. If the model has access to a web search tool, that's a different story.
                  • dexwiz 1 day ago
                    Is search part of the model or the harness?
                    • Cycl0ps 23 hours ago
                      Maybe that's a rhetorical question but just in case - the search would always be part of the harness. A model is only handling next-token prediction for a given input. That token may be something like [[web search]] to invoke a tool call but the actual call would be handled by the harness.
                      • dexwiz 20 hours ago
                        Yeah it was rhetorical. Search would be implemented as a tool call. Pure intelligence tests would likely have limited tools. But maybe they would have a python sandbox to solve issues like Rs in strawberry.
                    • Barbing 18 hours ago
                      If they made that statement and knowingly had search enabled, it would essentially be fraudulent.
                  • htrp 1 day ago
                    > Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

                    > Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.

                    Early access, no weights no tech details, just a sign up here for info

                    • wronglebowski 1 day ago
                      I'm all for more open models, but talk is cheap and this is a rather pointless announcement without anything backing it up. Publish your weights and HF repo or shut up IMO.
                      • antoniojtorres 21 hours ago
                        Is that all that a company about to give away the product of 10,000 GPUs running for a month gets to be now? give it away without a single promotional post or shut up? I support open source as much as the person but this is pretty caustic.
                        • mewse-hn 18 hours ago
                          I am also starting to take issue with "we've developed a new cutting edge model, and nobody can use it" announcements. Most recently with Google's Argon, at least they were using it internally and they'd be slowly rolling it out. This one is from a company I've never heard of, they're not releasing weights or offering API access at this point, it feels like a fairly worthless announcement.
                          • 40four 5 hours ago
                            It says in the very beginning that everything will be opened up by the end of the month. Maybe slightly annoying to you but not worthless. I for one am very excited to see some new American open models since we have called so far behind the Chinese. Looking forward to trying it out.
                          • ssivark 14 hours ago
                            Yes -- because they're late to the party (high performing open weights models have been a thing for a couple of years now), and therefore will be compared with all the other open-weights model providers they are competing against.

                            It's not just that they are doing users a favor with weights; they are just as much seeking favors with attention and usage (in a crowded market!).

                            • nl 17 hours ago
                              There have been previous "we will release the model" when the model release never comes in the past.

                              I think the comment you are replying to is unnecessarily hostile too though.

                              • RobotToaster 11 hours ago
                                > give it away without a single promotional post or shut up?

                                I think the point is that people are happy to see promotional posts when they actually release it, but only then and not before.

                                Unfortunately pinky promises from corporations to release something at some indeterminate time in the future aren't worth the bytes they're stored in, especially in the AI industry which is full of grifters and charlatans.

                                • ew-dev 14 hours ago
                                  The open weight community really is an odd one. Millions of Dollars for pre and post training given away for free and most often with very permissive licenses that allow commercial use (be it US, EU or mostly Chinese)

                                  Yet going by the comments on localLlaMA or HN, those companies are the devil :-D. Colour me surprised.

                                  • TeMPOraL 14 hours ago
                                    You know, the other night I trained a model on my secret stash of GPUs, that now outperforms Opus 5.5 on pelican benchmark and Jev on classification speed, while running on a potato.

                                    Will release weights soon.

                                    • 40four 5 hours ago
                                      I don’t understand what is wrong with some people lol. It says in the very beginning everything will be opened up by the end of the month. It’s like a bunch of toddlers throwing their milk cup on the ground and pouting because they can’t have their new toy right now!
                                      • autoexec 10 hours ago
                                        In the case of Facebook, no good thing they do will ever undo the evil things they've done, let alone the evil things they're still doing right now. How convenient for the devil that a single act of "charity" should make him immune from all criticism. By all means, praise whatever good Facebook does in the world, but don't kid yourself about what Facebook is.
                                        • crimsoneer 13 hours ago
                                          You don't get claim to be open and then not release your actual product.
                                        • octoberfranklin 15 hours ago
                                          Show me the mone^H^H^HWEIGHTS
                                        • adrian_b 13 hours ago
                                          They claim that they will release the model as open weights later this month.

                                          That means that they have the 31th of October as the deadline to make true their claims.

                                          The fact that they give early access to some may mean that they want some beta testers before the public release.

                                        • zelphirkalt 1 day ago
                                          And also a "proprietary data set" hahaha... Probably just means they don't want to show it, and it is data, that either they shouldn't have, or that there is nothing special about their training data and it is just meant to sound like there is some secret ingredient, while there is none.
                                          • janalsncm 1 day ago
                                            Not sharing the data is pretty standard because 1) it tends to get the lawyers involved and 2) good data is critical for getting good results.

                                            Imo you can get better results with great data and generic modeling techniques than with incredible modeling techniques and crappy data. Because if you have crappy data, you won’t even know if your model is good because your evals will also be bad.

                                            This is why Anthropic is throwing a fit about the Chinese distillation “attacks”. Clean reasoning traces are gold.

                                            • nl 17 hours ago
                                              This isn't true.

                                              Companies pay lots of money for proprietary agentic trajectories which are used during RL. These are things like "Task: summarize stock levels for months end accounting" which then traces the task though using SAP to look at different SKU stock levels, exporting them and generating summary Excel spreadsheets.

                                              This is very different to the "scrape the internet" datasets that a table stakes for training a LLM.

                                              Xiaomi released a fairly developer-centric dataset like this here: https://huggingface.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss

                                              SpreadsheetRL is another fairly specialized dataset: https://spreadsheet-rl.github.io/

                                              • mlmonkey 1 day ago
                                                Data has copyright issues, so one can't share it generally without getting permissions from all of the copyright holders. The data is not theirs to share, anyways. The derived (learned) weights are a different matter.
                                                • ordersofmag 19 hours ago
                                                  True for the pre-training data scraped from diverse sources. Less so for the later stage data for RL which is by all account more of a differentiator. In most cases the labs themselves produced the data so they are the copyright holders (or they are borrowing it from other labs via distillation). A lot of it is synthetic data, and since you can't copyright AI output, it becomes less about copyright and more about trade secrets.
                                                  • Borealid 21 minutes ago
                                                    If I train a model on a song's lyrics, and the user asks the model to recite the song, and it does, so in RLHF I downvote the generation because it refused-to-refuse to regurgitate the song's lyrics... is that RLHF session copyright-encumbered?
                                                • vanuatu 1 day ago
                                                  this is very normal for frontier lab companies. you need good data either synthetic or labelled (all the chinese open source models have their own armies of data labelers)
                                                • Loquebantur 1 day ago
                                                  > Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model.

                                                  > We will release the weights, technical report, model card, and developer artifacts later this month.

                                                • berkes 11 hours ago
                                                  What kind of company or organization is Reflection?

                                                  I think it is ever more important to realize who is releasing models rather than what the models do and how they compare.

                                                  Because models iterate at breakneck speed, looking at today's benchmarks is only useful for someone using the models today. Whereas if one builds a product on top of it, or commits to one for a project or team, the company or organization behind it, is far more important. Will they exist in a few months? Do they need a business-model? Are they subsidizing usage with venture capital and how long can they keep this up?

                                                  Is reflection a company? University lab? NGO?

                                                  • derivagral 8 hours ago
                                                    Looks like general reasoning is their target. As for who:

                                                    > The startup was launched in March 2024 by Misha Laskin, who led reward modeling for DeepMind’s Gemini project, and Ioannis Antonoglou, who co-created AlphaGo, the AI system that famously beat the world champion in the board game Go in 2016.

                                                    with the obligatory:

                                                    > Investors in Reflection AI’s latest round include Nvidia, Disruptive, DST, 1789, B Capital, Lightspeed, GIC, Eric Yuan, Eric Schmidt, Citi, Sequoia, CRV, and others.

                                                    • efficax 9 hours ago
                                                      • whatsdowndog 8 hours ago
                                                        Can't open that website with an ad blocker.
                                                        • robot_jesus 7 hours ago
                                                          Curious. FWIW, worked with no problems for me running ublock origin on Firefox/Mac.
                                                      • SadErn 10 hours ago
                                                        [dead]
                                                      • wren6991 23 hours ago
                                                        I thought it would be interesting to look at some key figures vs another contemporary model in the same weight class (DeepSeek V4.1 Flash)

                                                                                        DS V4.1F            Beam
                                                            LM total params             552B                501B
                                                            LM active params (prefill)  8B                  23B
                                                            LM active params (decode)   16B                 23B
                                                            N-gram/PLE params           196B                0
                                                            Pretrain tokens             45T                 28T
                                                            Disk KV bytes/token (FP4)   890                 No information
                                                            Vision                      Yes (pretrain)      No
                                                            Weights available           Yes (launch day)    "This month"
                                                            Weights licence             MIT                 Apache 2.0
                                                        
                                                        At first blush the benchmarks are impressive, but to paraphrase Linus: "Talk is cheap, show me the weights." :-)
                                                        • laybak 23 hours ago
                                                          yeah agreed, let's normalize "show me the weights" in this era
                                                        • NorwegianDude 1 day ago
                                                          Bigger and still worse than existing free Chinese models that are smaller? Open weight models are nice, but at this point it seems western models are very far behind Chinese ones, despite Chinese companies publishing a lot of their findings. I hope we get more open models and more providers, as being stuck with a model from China or US with no competition is risky.

                                                          Google does do a great job with Gemma models. It's one of the few language models actually good at language. OpenAI's top closed models can't even write norwegian correctly.

                                                          • mirekrusin 1 day ago
                                                            It takes time / few iterations to get it right (and it's moving target), but yes, expensive trial, my personal feeling is that they went a bit too high, at the same time who knows, maybe good move – as they're saying RL didn't plateau. It feels like they had something like $100M budget for it?
                                                            • irishcoffee 13 hours ago
                                                              [flagged]
                                                              • jubilee33 12 hours ago
                                                                I hear this so much from westerners who have no connection to the people or the country. -China passed a law forbidding companies from replacing workers with AI. -China regularly punishes CEOs and corrupt government officials who do bad things. China promotes open source to the world enabling everyone in every country to have equitable access. From living in china for over 20 years and talking to many many people, the majority is happy with the trajectory of their country. Could any for these be said of the west? I won't get into specifics on which countries regularly throw bombs on children, use starvation and blockades as weapons, and completely disregard the massive dissatisfaction of their citizens...but it sure isn't china. Please clean your own house first before you complain about your neighbor.
                                                                • yorwba 7 hours ago
                                                                  China didn't pass a law forbidding companies from replacing workers with AI. There was an unfair termination lawsuit that has been misconstrued as such http://english.scio.gov.cn/m/chinavoices/2026-04/30/content_... but it was applying existing law. If they had followed the provisions of Article 41 (3) of the Labor Contract Law https://english.court.gov.cn/2015-08/17/c_761484_6.htm which explicitly authorizes layoffs if "The enterprise changes its line of production, introduces a major technological updating or adjusts its business method, and, after modification of the labor contracts, still needs to reduce its personnel," there would have been no problem.
                                                                  • irishcoffee 6 hours ago
                                                                    Account made 3 months ago. You’re a shill.
                                                                    • xquce 39 minutes ago
                                                                      Comment going back a month on that account, not a single one even mention China, much less shilling.

                                                                      I guess you the shill then :)

                                                                    • CMay 9 hours ago
                                                                      [flagged]
                                                                    • goolz 12 hours ago
                                                                      This sentiment was really common circa 2005, or at least general anti-China rhetoric, where I lived in the Tri-state. Funnily enough, I have never felt more a cog, living here in the U.S., than I do right now. Maybe you are right, but my guess is wherever you live, yours is a bit of a glass house as well.
                                                                      • irishcoffee 5 hours ago
                                                                        I am right, and whatever glass house bullshit you’re talking about is a flimsy shim of disagreement.
                                                                      • MadameMinty 13 hours ago
                                                                        Yeah, we should only promote, like... Finnish, Norwegian, maybe Swiss products, like Apertus. Not this unethically-made baby oil[1] from repressive torment nexuses of USA or China.

                                                                        [1] https://i.imgur.com/rJSG019.jpeg

                                                                    • springtimesun 11 hours ago
                                                                      I feel like this is a marketing miss. If they had held their announcement until the model was released, I would have grabbed it and started running it through my benchmarks. It probably doesn’t get a place in the rotation based on their own description of its performance, but now the weights live on the server, I’m probably following them on HF and I will remember to check in every time I ls the models folder. With the announcement only, none of that happens and I’m likely to forget about this by the time it actually gets released.

                                                                      The email harvest move just doesn’t fit where we/they are in the cycle. There are established players and a buffet of models to choose from (plus a ton of empty hype). The first move at this point for any new entrant should be to show, not tell. Even an API only release with the promise to open weight would be better (actually probably all around better since most people can’t run this locally).

                                                                      I wish this lab and all the labs releasing the best. It’s a brutal landscape to sink millions of dollars into for a guaranteed “behind x model from a year ago” evaluation. But, I do believe there is genuine innovation left to uncover.

                                                                      • onlyrealcuzzo 1 day ago
                                                                        This appears to be larger than DeepSeek v4.1 Flash, more expensive to run, and worse on every measured metric.

                                                                        Am I missing something?

                                                                        • swiftcoder 1 day ago
                                                                          > Am I missing something?

                                                                          It's pretty clear from their framing ("Beam advances the Western open-weight frontier") that one of their main selling points is not being a Chinese lab.

                                                                          I can't imagine that mattering to many individuals, but I guess someone out there has a government contract that forbids the use of foreign models

                                                                          • htrp 1 day ago
                                                                            Reflection raised on the idea of creating the "American Deepseek Project"
                                                                            • mirekrusin 1 day ago
                                                                              Who's funding this?
                                                                              • Looks like Nvidia, Eric Schmidt, Sequoia, and a host of others https://techcrunch.com/2025/10/09/reflection-raises-2b-to-be...
                                                                                • ipsum2 1 day ago
                                                                                  I still don't get why, after over a decade on HN, people refuse to Google very simple questions.
                                                                                  • jauntywundrkind 23 hours ago
                                                                                    Reciprocally, people probably do go find out. The great filter is also how many of them come back to post the useful interesting information.
                                                                                    • hadlock 20 hours ago
                                                                                      I still don't get why, after 25 years of slashdot, still ask why people DRTFA and complain about it as meta commentary.

                                                                                      InB4: kids these days :shakes-fist-at-cloud:

                                                                                      • nickpsecurity 18 hours ago
                                                                                        Sometimes we just enjiy having a conversation with people. It's how we got many answers before Google exists. I believe such choices have many, positive effects on people that society is losing.
                                                                                        • dominotw 23 hours ago
                                                                                          quesiton is more like "lets analyze the motivations behind funding this"
                                                                                          • kelnos 22 hours ago
                                                                                            And a much better post would be "Just checked, and X, Y, and Z are funding this. My bet is that Y is funding it because $REASON..."

                                                                                            Asking an easily-searchable question is just lazy.

                                                                                            • calmworm 18 hours ago
                                                                                              Profit. Always.
                                                                                            • NamlchakKhandro 22 hours ago
                                                                                              because my templeos goes straight from ring0 after bios straight into a ui for hackernews that only lets me scroll, click into comments and type comments.
                                                                                              • JSR_FDED 20 hours ago
                                                                                                Try using the same Internet that you used to download templeos to search for stuff.
                                                                                              • nozzlegear 21 hours ago
                                                                                                I don't get why anyone posts anything on HN when they can just have an LLM generate an entirely self-contained conversation now.

                                                                                                /s

                                                                                                The conversation is why.

                                                                                          • mirekrusin 1 day ago
                                                                                            Multiple independent approaches are cool and all but fully open source model training (datasets, pipeline, checkpoints) should be taking advantage of being open and share runs/budget between different entities.
                                                                                          • dotancohen 1 day ago
                                                                                            We're still at the stage where every new entrant is welcome in my opinion. Doesn't need to be record-breaking upon initial release.
                                                                                            • flockonus 1 day ago
                                                                                              It depends! If a startup is entering with a large model to face other larger models, it must be better at least in 1 meaningful dimension.

                                                                                              500B params performing worse than other OSS of the same size is pretty meaningless if no one will use it.

                                                                                              • bot41 14 hours ago
                                                                                                Look at meta, while they went from open to closed, they got better as time went on
                                                                                                • harpersealtako 3 hours ago
                                                                                                  I mean, the obvious way it's better than the Chinese open-weight models is literally the fact that it's not Chinese and is therefore less likely to be banned or restricted. Chinese models cannot be used on certain government systems already in the US, and regulators are actively considering expanding these restrictions more generally (such as adding it to the Entity List).
                                                                                                • halJordan 1 day ago
                                                                                                  I disagree. Sure let them play and see if they can improve. But this model has more compute and more training data than the predecessors it fails to surpass. That only means their training regime is inferior if their predecessors did so much more with so much less. That inferiority should not be encouraged.
                                                                                                  • kelnos 22 hours ago
                                                                                                    You don't just magically do better than everyone else on every metric on your first go at something. Doing worse than others and refining is how pretty much everything works.
                                                                                                    • halJordan 19 hours ago
                                                                                                      I think i captured that in my first (second?) sentence
                                                                                                    • janalsncm 1 day ago
                                                                                                      The reality is they trained a model and it looks worse on benchmarks than Qwen or GLM. I don’t see how sharing the weights hurts anyone? Even when Llama 4 came out and it was a dumpster fire, it didn’t affect me personally.

                                                                                                      > That only means their training regime is inferior if their predecessors did so much more with so much less

                                                                                                      Hard to imagine how that wouldn’t be the case. They probably missed the boat on distilling Claude (or their lawyers said no), they probably didn’t hire an army of math PhDs to write reasoning traces, they don’t have millions of DAUs in a coding agent to train from, and they probably have less money, less experience, fewer top tier researchers, and fewer resources for experiments. They are an underdog without a doubt.

                                                                                                      None of that means they shouldn’t release their model.

                                                                                                      • halJordan 19 hours ago
                                                                                                        Them releasing the weights doesn't hurt anyone. It's the peanut gallery clamoring to put them onto the same pedestal as actual tier 1 companies simply because they aren't named openai or anthropic that is hurtful.
                                                                                                      • thinkcontext 23 hours ago
                                                                                                        I wonder if that's an indication that they are not distilling which limits how good they can get.
                                                                                                        • halJordan 19 hours ago
                                                                                                          Openai, grok, and Anthropic aren't distilling. Theyre just second class. It's not a big deal, we just shouldn't be lauding them for being second class.
                                                                                                          • girvo 16 hours ago
                                                                                                            > Openai, grok, and Anthropic aren't distilling

                                                                                                            Says who? We know Grok does at the least. They admitted it openly.

                                                                                                            • thinkcontext 19 hours ago
                                                                                                              Musk said under oath that they use distillation for Grok.
                                                                                                              • disgruntledphd2 14 hours ago
                                                                                                                And it still sucks. My apologies to the Cursor team but that's just very very poor performance.

                                                                                                                Alternative explanation is that the Chinese have far more technical talent than anyone else, along with the infra and capital to build out these models.

                                                                                                                My money is on the latter explanation, tbh.

                                                                                                      • vanuatu 1 day ago
                                                                                                        Reflection is explicitly marketed as the 'US' DeepSeek

                                                                                                        seems like they are aiming to provide both inference and RLaaS for american companies and western govts. even if they never fully beat deepseek if they get close enough the fact that they're American will help them close deals

                                                                                                        • efficax 9 hours ago
                                                                                                          deepseek is from the evil east, this is from the virtuous west
                                                                                                          • jstummbillig 1 day ago
                                                                                                            Apparently there is more to making good models than copying everything on the internet.
                                                                                                            • chews 20 hours ago
                                                                                                              12 yards long, 2 lanes wide, 65 tons of American Pride! Canyonero! Canyonero!
                                                                                                              • aizk 1 day ago
                                                                                                                Yes it's (hopefully) not distilled from every single major American provider.
                                                                                                                • Centigonal 1 day ago
                                                                                                                  new entrant in this weight class, US lab.
                                                                                                                  • martini333 1 day ago
                                                                                                                    Beam goes brrrr
                                                                                                                  • michaelkdev 1 day ago
                                                                                                                    Is it worse than the top open-weight Chinese models? Yes, it is, but at least the West has joined the party, and hopefully they will iterate on this and keep up the pace. The Chinese labs will certainly release new and powerful versions soon, so it's all about relative pace right now.
                                                                                                                    • drubs 1 day ago
                                                                                                                      I remember being in the room with pretraining day 1 to help monitor the training job launch. Watching this model train from day 1 has been an amazing experience!
                                                                                                                      • jeremyjh 1 day ago
                                                                                                                        What sort of outputs or telemetry is monitored on a large pre-training run?
                                                                                                                        • drubs 1 day ago
                                                                                                                          Outside of ML metrics, you're monitoring the health of every piece of hardware in the system. You need to make sure that you have every GPU, every CPU, the PCIe buses, the networking fabric are all working without any errors. You need to ensure that you can respond as fast as possible to any possible error. One bad component can bottleneck the entire job.
                                                                                                                          • costco 21 hours ago
                                                                                                                            https://github.com/facebookresearch/metaseq/blob/main/projec...

                                                                                                                            I really enjoyed reading the log book from the training of OPT-175B at Meta… I guess it’s all classified info but it’d be fun to read a blog post about the crazy day to day issues you run into when doing things at this scale :)

                                                                                                                            • spindump8930 20 hours ago
                                                                                                                              GPU failures are frequent enough that at a certain scale, you constantly have workers dropping out. Designing systems that can still keep training is very interesting!
                                                                                                                      • efficient_dairy 18 hours ago
                                                                                                                        I don't find open-weight models that impressive anymore. MiMo-V2.6 already showed that you can have a not so crazy architecture and enough compute, the bottleneck is then just the data. OAI and Anthropic are largely the frontier models because of the synthetic data they made. They have a large customer base and have the user's data as well as knowing what tasks their customers use the models for and what domain they should get synthetic data for.
                                                                                                                        • computerdork 17 hours ago
                                                                                                                          Interesting. Have heard about the need to create synthetic data for LLM's, but didn't realize it was such an important factor. Although, how big a factor synthetic data is the next question, but guess there is no way to truly verify how much difference it makes with these closed models.
                                                                                                                          • efficient_dairy 17 hours ago
                                                                                                                            This is literally what the labs have been doing over the past year. The whole idea of emergence is a lie, there is some interpolation and superhuman long evaluations the models can do, but almost all the gains are from synthetic data. They hire thousands of professionals and pay them as much as 200$/hour to create many tasks that they want the model to perform and use these to teach the model on how to do it with RL. OAI had 30,000 contractors from Mercor for Sol 5.6. When you see Opus suddenly becoming great at blender or some other 3D graphics, that's because they hired professionals and had them do similar tasks that people are looking for. They keep having better professionals at each iteration and so the quality improves. There is no emergence or "General" intelligence. The model doesn't learn to become better at a task because of scaling laws or emergence or whatever they might wanna say, it is literally RL on tasks that they want the model to perform well on.
                                                                                                                            • computerdork 16 hours ago
                                                                                                                              Also interesting. Do you happen to know if once they've used those thousands of contractors to teach the model something like Blender (or some similar app) using RL, when they training a new model, do they need to use same contractors again to teach that new model the same behaviors?

                                                                                                                              The reason I ask is even though it can take hundreds or thousands of contractors to teach a model a certain behavior, wonder if they really only need to do it once for each desired behavior? (of course, future models might expand and refine this previous training) Because if that is the case, then wow, then future models can really expand their capabilities really very fast.

                                                                                                                              ...Am wondering if they somehow record a training session so they can play it back whenever they need to train a new model with the same info? Or maybe the new models can just use distillation from the old model to relearn the old behaviors?

                                                                                                                              • efficient_dairy 16 hours ago
                                                                                                                                The contractors don't directly teach the model. They create datasets. Mostly they create tasks within an environment that the current generation of models wouldn't be able to do, they then write maybe a solution, a set of rules for evaluating the response, and whatever is needed for the RL. These tasks form a dataset. You can see examples of a task in the Mimo dataset that was open-sourced recently. The dataset would then be used by engineers for post-training of whatever model. Some of the model iterations that are released every month tend to be just a further post-training of the same base model that was pre-trained months ago. OAI recently has been doing a lot more pre-training, but for a long time they had the same pre-trained base model. This is why you see so many releases done so fast by the labs, they just need to post-train the same base model on whatever task they think would be better suited, I would speculate based on what users want and what the benchmarks test for.

                                                                                                                                Regarding your second question, I think if they want to further post-train a model using new data, they wouldn't feel the need to re-train it again on the data that it has already being trained on. But you never know. If the model is a completely new pre-trained base model, then they could either train the model using all the data and/or use a previous model to teach it. There is definitely a bunch of tricks they do to evaluate the models and check the performance or whatever their recipe is. It's really up to what the engineers would prefer. But you get the core idea, the models are not suddenly coming up with how to use the Blender on their own, they are explicitly being trained on a dataset curated by a professional that teaches the model how to use Blender. Surely there is another aspect that if the model gets better at coding, then it also helps it become better at Blender, and you have that transfer learning. However, there is no emergence or a deity popping up. But you get people who were evaluating theses models on blender use and suddenly seeing the model ace their tasks and they think they are dealing with a super-intelligence. They then undergo an AI psychosis once they try and extrapolate the (super)-exponential improvement in that one task over the next few months and across all other domains.

                                                                                                                                Regarding your third question, I already answered at it. But, when it comes to training, they definitely freeze the weights after each run just in case an issue arises and they need to address it (a GPU not working or the loss value blowing up).

                                                                                                                              • saberience 13 hours ago
                                                                                                                                What you're talking about is not synthetic data. Synthetic data is data generated by an LLM. If the data is generated by a human, it's not synthetic.

                                                                                                                                You're talking about real data created and curated by humans to help in training LLMs.

                                                                                                                                • efficient_dairy 50 minutes ago
                                                                                                                                  This is true and thank you for correcting me. I make a mistake of using synthetic data for what is human-generated data. There are augmentation tricks they could use on the human data but still human-generated data shouldn't be mistaken with synthetic data. I believe my core arguments still stand. Specifically that these models are learning their capabilities through specialized data that human professionals curate. There is no emergence or "general" intelligence. This is really what the open-weight models lack.
                                                                                                                          • algoth1 19 hours ago
                                                                                                                            Does this one also routes to Claude under the hood like Reflection 70B did? I recall they even run a basic regex to remove "Claude" from the output. Then they promised to be completely transparent on the postmortem (they claim they had no idea what had happened), but the postmortem never came
                                                                                                                            • eldenring 18 hours ago
                                                                                                                              Reflection AI is a completely different company with no relation (as far as I can tell) to the model you are referencing from 2024
                                                                                                                              • algoth1 9 hours ago
                                                                                                                                If that’s the case i take it all back. I thought this was another venture by Matt Shummer’s Reflection Ai
                                                                                                                            • ContinuityLab 13 hours ago
                                                                                                                              Releasing open-weight models at this scale is a massive milestone for the community. Access to transparent model internals is foundational for trustworthy systems.
                                                                                                                              • astrostl 22 hours ago
                                                                                                                                Beggar choosing: my kingdom for more 90B-133B MoE local models. Especially with disk offload, that is a function/performance sweet spot for Mac workstations with 64GB-128GB of RAM.
                                                                                                                                • reissbaker 18 hours ago
                                                                                                                                  Instead of yet another mediocre but fully-made-in-the-West open model (alongside Mistral, Trinity, Poolside, Inkling, etc etc) I'd really love for a Western neloab start the same way Qwen did: by focusing on post-training. Qwen's first release was a Llama 1 finetune [1]! Once they made it useful, they started working their way back in the stack to also do their own pretraining, etc. Starting with pretraining feels like such a waste: there's millions of dollars of crystallized compute and data sitting around in the Chinese model weights. Why not start with one of those, and only work your way back to pretraining once you've released something you can prove is useful?

                                                                                                                                  1: https://en.wikipedia.org/wiki/Qwen

                                                                                                                                  • redox99 18 hours ago
                                                                                                                                    I think there aren't any recent base models to post train on. Labs don't release them any more.
                                                                                                                                    • reissbaker 18 hours ago
                                                                                                                                      At least for RL, you don't need a base model — the rollouts are run in an inference engine with an instruction-tuned model using a chat template. You can start with just that!
                                                                                                                                      • redox99 16 hours ago
                                                                                                                                        Yeah but they are already RL'd too. No doubt you can further improve an RL'd model by doing your own better RL on top of that. But it's not going to be the same as starting with a base model or instruction tuned model.
                                                                                                                                      • volf_ 18 hours ago
                                                                                                                                        • barrrrald 16 hours ago
                                                                                                                                          This is point of Nemotron series
                                                                                                                                      • eaf7e281 1 day ago
                                                                                                                                        > Where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time.

                                                                                                                                        It's great to see a company that acknowledges it still needs improvement instead of making false claims.

                                                                                                                                        • Iolaum 16 hours ago
                                                                                                                                          I am really curious why after Qwen-3.8-flash and deepseek-v4.1-flash newer open weight models (or proprietary but they don't tell) don't use n-grams. It looks (to me) like they are a cheep way to add more knowledge to the model.

                                                                                                                                          What do I mean by cheap? You can rely on the SSD to retrieve the relevant tokens as no computation is needed meaning you can leverage storage (or cpu ram if you don't have unified memory) to serve part of the model which (to my understanding) is much cheaper to get than GPU RAM.

                                                                                                                                          Anyone know what am I missing? Or is it that the pace of iteration for labs slow enough that they can't actually leverage it yet?

                                                                                                                                          • wren6991 10 hours ago
                                                                                                                                            DeepSeek Engram paper published: 12th January

                                                                                                                                            Qwen3.8 Flash Next release date: 26th August

                                                                                                                                            DeepSeek V4.1 Flash release date: 10th September

                                                                                                                                            Current date: 6th October

                                                                                                                                            I think they'll become more popular in the coming months. Also Gemma 4 PLE (April) is similar to DeepSeek Engram in a lot of ways, just with 1-grams.

                                                                                                                                            On the proprietary model point: I'm personally curious about whether heavy n-gram offload is one reason Anthropic keep driving down their token vocabulary size (the other reason being eliminating the LM head gradient bottleneck).

                                                                                                                                          • Phineas_here 11 hours ago
                                                                                                                                            Hey since it's an open-weight model, I would love to know: how much model safety alignment have you done? Did you do any sort of post-training and what restrictions are in place

                                                                                                                                            Ideally i would like to place my own restrictions and align from scratch, currently I am resolved to do harness alignment using tools like Prismor but would love to do my own post training alignment

                                                                                                                                            • aeetes 1 day ago
                                                                                                                                              the performance chart puts the better open source models behind the fold making it seem like it outperforms them... but it doesn't! all for open source models but this announcement is misleading
                                                                                                                                              • brumbelow 1 day ago
                                                                                                                                                Yes. All the link made me realize is that I should checkout Deepseek 4.1 flash
                                                                                                                                              • xutopia 6 hours ago
                                                                                                                                                What kind of machine does one need to run this model? What's the usual setup someone would have for running this?
                                                                                                                                                • segmondy 1 day ago
                                                                                                                                                  Any time a new lab shows up, folks complain about how their models are worse. Really? It would be nice if a new comer comes from no where and beats everyone, but that's rarely the case. The good thing is that other labs/people are figuring out how to build this, and if they keep at it then this is as bad as it gets for them and it would hopefully get better. A new entrant to the market is good for everyone.
                                                                                                                                                  • spindump8930 20 hours ago
                                                                                                                                                    Reflection isn't really "from no where", they have huge financial backing, 10K of the latest GPUs, and many ex leads from the established labs.
                                                                                                                                                    • NamlchakKhandro 22 hours ago
                                                                                                                                                      honest question: how honest do you think people are about their improvements and performance compared to objective results when all you do is praise them?

                                                                                                                                                      participation awards are not helpful.

                                                                                                                                                    • hypfer 1 day ago
                                                                                                                                                      Someone should name their next model "Workhorse" just for SEO reasons.

                                                                                                                                                      It's interesting how the industry converged to this very term, given that very less work is being done by horses since quite a while.

                                                                                                                                                      • latentsea 19 hours ago
                                                                                                                                                        I guess future AI agents might market things as a "workman" for the same reason, despite less work being done by people :)
                                                                                                                                                        • hnedeotes 1 day ago
                                                                                                                                                          Thankfully there wasn't mistagging, we could have ended with workjackass.
                                                                                                                                                        • gitowiec 11 hours ago
                                                                                                                                                          I'm on the waiting list... Couldn't find any download option, so I suppose it is only obtainable through their API. Strange way to distribute open weights model.
                                                                                                                                                          • shingoshoji 16 hours ago
                                                                                                                                                            I'm a big fan of open-weight models.

                                                                                                                                                            It's true that no benchmark communicates the whole picture, and we won't really know how it behaves until weights are out, but the performance here doesn't seem particularly groundbreaking just based on the benchmark.

                                                                                                                                                            • soltanov 14 hours ago
                                                                                                                                                              The meaningful test begins after the weights arrive: fixed harness, network disabled, repeated trials, and real latency, memory, energy and cost per completed task.
                                                                                                                                                            • Marcuss2 15 hours ago
                                                                                                                                                              Very little in terms of the layers they use. Calling it now, they are using global layers everywhere, making the model basically unusable due to high KV Cache use.
                                                                                                                                                              • Tepix 12 hours ago
                                                                                                                                                                If Beam "rivals GLM 5.2 on reasoning" does it mean it's as good as GLM 5.3 Flash? (a much smaller model)
                                                                                                                                                                • antonyragleap 13 hours ago
                                                                                                                                                                  Open-weight at 501B is huge. What was the eval setup - was web search disabled to test generalization?
                                                                                                                                                                  • hidelooktropic 22 hours ago
                                                                                                                                                                    Nitpicking but I really wish this benchmarks table were easier to read. Should show which columns win in each row and should not require horizontal scrolling to see across.
                                                                                                                                                                    • minimaxa 22 hours ago
                                                                                                                                                                      That was exciting... Will come back later when the weights are out and gguf'd.

                                                                                                                                                                      Access is currently limited. We'll contact you if early access becomes available.

                                                                                                                                                                      • JeanCampos 13 hours ago
                                                                                                                                                                        Minimum hardware to run this model?? doable in consumer hardware?
                                                                                                                                                                        • Tepix 13 hours ago
                                                                                                                                                                          2x 128GB unified memory machines with RDMA could run this at Q4. Not very fast.

                                                                                                                                                                          The price of that kit is rapidly approaching $10k.

                                                                                                                                                                      • vcryan 1 day ago
                                                                                                                                                                        This is like an ad for how great Deepseek V4.1 Flash is.
                                                                                                                                                                        • zopper 1 day ago
                                                                                                                                                                          Open model that is not yet open or widely accessible via API. Primarily comparing to non-SOTA models like Inkling and GLM 5.2. Included comparison to GLM 5.3 and DeepSeek V4.1 Flash in the table, but not in the charts (I assume they would make them look bad). Also no results from AA Index or Arena.
                                                                                                                                                                          • keeganpoppen 1 day ago
                                                                                                                                                                            very curious to see more about what kinds of hardware you can run this on and the perf. characteristics… on the face of it, it seems like optimizing for inference speed might(?) be good for running on smaller hardware, but i suppose it could be the other way around and it is actually much resource-hungrier for the number of parameters, etc. …
                                                                                                                                                                            • andai 18 hours ago
                                                                                                                                                                              Noticed some good questions made dead at the bottom of this thread. Odd...
                                                                                                                                                                              • arjunchint 17 hours ago
                                                                                                                                                                                ok so they are comparing themselves to and claim to be beating GLM's last generation GLM 5.2 model, GLM 5.3 Flash is a monster, this is honestly embarassing
                                                                                                                                                                                • sharktheone 1 day ago
                                                                                                                                                                                  Am I the only one who thought of the BEAM VM after the first word of thee title?
                                                                                                                                                                                • TheArcane 1 day ago
                                                                                                                                                                                  If you don't buy into "America good, China bad" narrative, this new entrant & release by Inclusion Ai is a lot more exciting by every measurable metric.

                                                                                                                                                                                  https://github.com/inclusionAI/Ling

                                                                                                                                                                                  • spindump8930 20 hours ago
                                                                                                                                                                                    There are many measurable metrics and I don't see any that seem impressive. Care to share the ones you found exciting?
                                                                                                                                                                                    • jonathaneunice 23 hours ago
                                                                                                                                                                                      That seems to be from ~1y ago. What am I missing?
                                                                                                                                                                                      • moelove 19 hours ago
                                                                                                                                                                                        You can directly check the link to Huggingface in their readme; they are constantly updating it on HF.

                                                                                                                                                                                        BTW, this is also an AI lab from China

                                                                                                                                                                                      • JSR_FDED 20 hours ago
                                                                                                                                                                                        I don’t buy into that narrative, but I don’t understand why Ling’s release is more exciting?
                                                                                                                                                                                      • ronfriedhaber 14 hours ago
                                                                                                                                                                                        Congrats to the reflection team
                                                                                                                                                                                        • tesnorindian 16 hours ago
                                                                                                                                                                                          I am a great fan of open weights model, but off late I am starting to loose track on the capabilities of the latest models released and now a days most of the open weights models are above 100 B params which is not going to run in our laptops. What happened to Jev hype? A 501B Beam model is not going to help with it (Non AR Schema driven responding under 1 second).

                                                                                                                                                                                          I am more interested with a SOTA Frontier 8B-10B model. Is this even possible?

                                                                                                                                                                                          • wg0 1 day ago
                                                                                                                                                                                            Suppose I inherited a data center spanning several hundred acres full of GPUs and free electricity.

                                                                                                                                                                                            Where do I get the data?

                                                                                                                                                                                            I mean, this many models. They have to start somewhere.

                                                                                                                                                                                            • petu 1 day ago
                                                                                                                                                                                              I guess public datasets on HuggingFace and some shadow libraries content is enough to start.

                                                                                                                                                                                              e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb

                                                                                                                                                                                              • lucrbvi 1 day ago
                                                                                                                                                                                                There are a lot of open-research on pre-training, post-training and RL data mixtures and sourcing.

                                                                                                                                                                                                I recommend checking papers from Datalogy, Nvidia Nemotron, Ai2 (Ollmo, Tulu, ...) and the recent model from Aleph Alpha if you want to learn more.

                                                                                                                                                                                                • spindump8930 20 hours ago
                                                                                                                                                                                                  All of those are good references. Other folks in the thread are missing distinctions between pre training (~the internet + curated sources) and post training (~instructions and RL)
                                                                                                                                                                                                • wren6991 22 hours ago
                                                                                                                                                                                                  I heard you should ask Claude about this. Preferably with thousands of accounts, routed through residential proxies
                                                                                                                                                                                                  • altcognito 1 day ago
                                                                                                                                                                                                    If you ask a model, they will generally tell you where to get data. Modern frontier models have the large advantage of having tens if not hundreds of millions of users providing use cases to train against to improve their responses.
                                                                                                                                                                                                    • ttul 1 day ago
                                                                                                                                                                                                    • blourvim 22 hours ago
                                                                                                                                                                                                      You can also hire teams to create data for you for higher quality.
                                                                                                                                                                                                      • Hamuko 1 day ago
                                                                                                                                                                                                        Get data from Claude. That's what the Chinese (allegedly) do.
                                                                                                                                                                                                        • kbwal7 1 day ago
                                                                                                                                                                                                          Note that this sort of distillation is NOT for pre-training data (which is tens of trillions of tokens). I think the allegations against Chinese companies by Anthropic is more so that they distill SFT data (which is good for post-training, but you still need a strong base model)
                                                                                                                                                                                                        • forget the data....sell it and go live your life!
                                                                                                                                                                                                        • nullc 14 hours ago
                                                                                                                                                                                                          fauxpen like OpenAI? says open but no weights. Feels like a bid to get a buyout before the weights go public. Doesn't even have technical details.
                                                                                                                                                                                                          • webbrainiac 1 day ago
                                                                                                                                                                                                            [flagged]
                                                                                                                                                                                                            • lin7c 15 hours ago
                                                                                                                                                                                                              [flagged]
                                                                                                                                                                                                              • derin-picment 23 hours ago
                                                                                                                                                                                                                [flagged]
                                                                                                                                                                                                                • meherabhossain 14 hours ago
                                                                                                                                                                                                                  [flagged]
                                                                                                                                                                                                                  • hmartin 1 day ago
                                                                                                                                                                                                                    [dead]
                                                                                                                                                                                                                    • alexx-devv 10 hours ago
                                                                                                                                                                                                                      [dead]
                                                                                                                                                                                                                      • yalok 21 hours ago
                                                                                                                                                                                                                        pretty wild that, per public sources, Reflection AI raised over $4 billion and hit a $25 billion valuation while operating in total stealth, without ever releasing a single public product until now. Beam (501B) seems be their first-ever model drop. Or am I missing something?