Rendered at 18:29:38 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
smackeyacky 2 days ago [-]
How can these models do anything close to RSI when they can’t even self check their output? Gemini for example is so self confidently wrong about 30% of the time for me on certain tasks. I tell it that its answer is wrong and it issues a mea culpa but goes back to being wrong in short order. I feel like the AI industry is still massively overstating their projections.
singingfish 2 days ago [-]
I can't see a pathway for these things to be able to learn from experience in any meaningful way - the energy budget seems to prohibit it. Artificial neural networks are already many orders of magnitude more energy intensive than natural neural networks. Despite this, large nervous systems are very energy intensive as well. For example in humans 20% of the energy budget goes to the brain which is 5% of the body weight.
So I refrer to LLMs as language extrusion confabulation machines. Language extrusion was a term I heard the linguist Emily Bender use. Confabulation because my observation is that talking to an LLM is very much similar to my experience of interacting with Korsakov syndrome patients some years ago.
I look forward to the hype settling down to see what we end up with.
rcxdude 1 days ago [-]
> Artificial neural networks are already many orders of magnitude more energy intensive than natural neural networks.
What metric are you using to compare? By most counts, the energy budget of an instance of an LLM in a datacenter is lower than the energy a person uses. Of course, if you count energy per neuron connections vs weights then you'll likely get a quite different number. But then again LLMs do a lot of things with far fewer weights than the brain does neuron connections, even if you only count neurons in some parts of the brain. And of course you can point to capabilities that the brain has but LLMs lack, but on the whole it feels like it's pretty difficult to make a useful like-for-like comparison here.
singingfish 21 hours ago [-]
I can't make sense of your comment. Firstly because of the obvious massive over-build of GPU infrastructure the AI companies are engaging in. Secondly because the instance of the LLM in the data centre that users interact with is only a small part of the story. The training phase is clearly prohibitively expensive, thus the fact that these things have no way to learn from experience except by smoke and mirrors.
Also your comment feels like the classic climate denial discourse - say something a bit complicated and a bit difficult to follow that looks at a very small out of context part of the story to cast doubt.
singingfish 18 hours ago [-]
Maybe I shouldn't have been pissed of by your commient. In which case - the excess costs of the training, output, feedback, and training cycle of artificial neural networks seem to prohibit the nightly training consolidation cycle (i.e. sleep) that is an important part of natural neural networks.
rcxdude 11 hours ago [-]
There's two sides to efficiency: one of them is how much you put in and the other is how much you get out. I was asking the question because the answer of the relative efficiency of a human vs an LLM strongly depends on what you put on either side of the equation. It's a really easy pitfall to look at one side of one version of that equation and presume a really high inefficiency when that's not necessarily the case. It doesn't help that the data necessary to fully answer any version of this isn't publicly available. That's why I was asking for something more concrete in terms of how you were making the comparison.
If you're looking at training costs, there is definitely one aspect in which LLMs are obviously significantly less efficient: the amount of information needed for the initial training. This does translate into some pretty high costs but it only needs to be paid once for the amount of work that any given LLM does. In terms of fine-tuning LLMs can get significantly more data-efficient than the initial model, which also means energy-efficient, and it's not obvious to me that it would be drastically worse than a human (though again, only thinking in terms of doing the energy input for the kind of work that an LLM is good at).
(The increased efficiency in comparison to humans is part of the reason why you see Jevon's paradox mentioned a lot whenever concerns about the resources used by LLMs are mentioned: more efficiency can easily mean more resource use in total)
StevenWaterman 2 days ago [-]
As someone who used to use Gemini a lot, if you are predominantly using Gemini you don't know what the current state of things is like
pinkmuffinere 2 days ago [-]
I think your reply has a somewhat familiar structure -- "it doesn't work for you because you used an [old / suboptimal / non-frontier] model. If you use X you'll see that it works". You might be completely correct! But these sorts of claims push the onus back onto the other person, without accepting any work for yourself. It gets tiresome to retest with the newest model every other week. Is there any data you can provide to support your claim, or any result you can contribute here?
StevenWaterman 2 days ago [-]
The frontier is advancing really rapidly. The models are getting better faster, especially on RSI related tasks. The best way would be to try astra or fable on some hard problems.
Other than that I'd look at some of the more unique benchmarks for astra, like playing factorio or using blender. It's an entirely different beast.
peterashford 2 days ago [-]
I agree with you somewhat but I also just this morning read an article from a Blender educator who tried to replicate the Blender demos and couldnt get the same quality of results nor get results without errors that werent evident in Anthropic's demos
x-complexity 1 days ago [-]
> Is there any data you can provide to support your claim, or any result you can contribute here?
By the time we can show you data that convinces you that it does work, the next generation would already be out & incrementally dismantling the old conjectures that were true in the previous generations.
You're fundamentally asking for a violation of how information passively disseminates amongst humans: To go any faster requires more effort on the receiver's part to move up on the adoption curve.
croon 1 days ago [-]
Wouldn't this also mean that all previous generations that were proclaimed as intelligent and working were in fact... not?
It doesn't matter what comes tomorrow, with the next generation, if the claims now can't be proven.
To preempt the response: The math proof, regardless of them using non-disclosed user data or not, they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent.
pinkmuffinere 1 days ago [-]
To preface -- I try not to be dogmatic/politicized on AI, so I will genuinely consider your arguments! Please try to convince me. (indeed, I am the grandparent commenter)
I agree with the meat of your statement, but am very interested in the pre-emption, "they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent". First, I think the $30M number is inflated -- that's what the general public would have paid, but presumably the internal cost is lower, perhaps it's more like $10M. But it is still expensive. Second, I'm curious if it's really the case that they did a 1000-monkeys approach? I haven't read much in-depth reporting about the proof, so it's totally possible I just don't know. What is it that they did which is more like 1000-monkeys? Also, I wonder if that distinction matters -- if 1000 monkeys can reliably make ground breaking proofs, and the approach generalizes to other tasks, I will happily become a circus owner. Maybe you're claiming that it won't yield other proofs? Or the proofs are too opaque to be useful to humans? Or it can handle proofs but not other tasks?
croon 19 hours ago [-]
I'll respond/comment on the parts I hope are relevant to you, in no particular order:
Yes, a proof is a proof regardless how you get there. We however don't hear about when they fail, and I doubt their 10000 agents (from their own statement) would necessarily reach another solution/proof (this by leaning towards using user data after finding out others were close). They could as well have attacked another Millenium problem, but they didn't. In whichever case, we will have to wait and see if they (either company) can reach novel solutions/proofs without significant amount of human provided data for the LLM to bridge the gaps.
Further, and this is more of a policy opinion/prediction: If the numerable obtainable (albeit very hard) problems are solved, assuming training data is needed, will it push out future researchers from entering the field due to lack of reachable goals, thus cutting off future training data? LLMs have been great at replacing gateway jobs. But those jobs are what leads to frontier training data (be it maths, physics, chemistry, economics, graphics, prose, etc).
pinkmuffinere 10 hours ago [-]
Ah it's interesting they legitimately used 10k agents, I didn't realize that. I do agree that it's significant they solved the Millenium problem only once humans had made significant headway. I'm not sure I believe the relevant training data will be produced at a significantly lower rate due to AI -- Millenium problem solutions weren't generated at a very high rate before anyways. But I think all your points hold nonetheless. Thanks for elaborating on them!
2 days ago [-]
mancerayder 2 days ago [-]
Maybe Gemini is loosey goosey on purpose so we angrily correct it - then feed something on the back end that trains a different model?
It's so abysmally bad on Google search... and it's free. Isn't Google the great pioneer of the product is us?
emodendroket 2 days ago [-]
The kind of results you get from the one on Google Search and a dedicated "Gemini Pro" response are totally different. I'm assuming that's cost savings.
2 days ago [-]
kakugawa 2 days ago [-]
They can only do it in the (narrow) domains that are verifiable.
drodgers 2 days ago [-]
> Gemini
That's definitely part of your problem.
In my recent experience, error rates for astra/fable are at or below human level. Just like when directing humans, it pays to ask probing questions ('Are you sure about X?', 'Did you check for Y?', 'Please run Z just to double check.') if you really care about the result being correct.
glhaynes 2 days ago [-]
And if you're building something of any importance, you need to have verification steps at checkpoints. It's honestly just engineering.
Weak models tasked with review can catch a decent amount of the mistakes that weak models make and help them be much better, especially if you have them verify against authoritative sources. Strong models make far fewer mistakes to begin with. And, for the mistakes they do make, a swarm of reviewers (same model or somewhat weaker, reviewed by the stronger model) can really help reduce the error rate further.
PantaloonFlames 2 days ago [-]
Gauging state of the art against what is available for free or for very cheap per token cost is like gauging the maximum theoretical transport potential by riding a bicycle.
You’re using something that is very energy efficient; you cannot extrapolate that experience to conclude that SOTA models are not doing something much different.
RataNova 1 days ago [-]
[flagged]
daavidhauser 2 days ago [-]
Opus 4.8 plus OpenClaw. I feel like the space is moving so fast that the result with this setup says very little about how close we are actually now.
pinkmuffinere 2 days ago [-]
I made this same reply to another thread, so I'm sorry to say essentially the same thing twice, but -- the comment has a familiar structure, "it doesn't work for you because you used an [old / suboptimal / non-frontier] model. If you use X you'll see that it works". These sorts of claims push the onus back onto the other person (or in this case tfa), without really accepting the result, or taking on any work for yourself. It gets tiresome to retest with the newest model every other week. Is there any data you can provide to support your claim, or any result you can contribute here?
1attice 2 days ago [-]
Yes, that's an annoying predicament, but we come by it honestly.
Limits of scientific method in the face of exponential takeoff, IMO, and kind of proves the opposite result (RSI appears to be here)
peterashford 2 days ago [-]
It may be tiresome, but it's true. It doesn't refute the article's premise thou, only retesting with newer models would and that would be the same tiresome requirement you already called out
dgellow 2 days ago [-]
Yeah, it’s crazy how fast things have changed in a month. I couldn’t find a more recent replication or similar study but it would be interesting to see it done with the current frontiers. Though I don’t think that would change much about the overall conclusion of the paper
0xDEAFBEAD 2 days ago [-]
We need to be careful of wishful thinking. People are going to want to assume the existence of some sort of "deus ex machina" which is going to make everything fine. I prefer to turn the logic around. If there's any decently high chance that things could go off the rails, we should be shutting AI development down: https://pauseai.info/
strgrd 2 days ago [-]
It is easier to imagine the end of the world than the pausing of AI.
0xDEAFBEAD 2 days ago [-]
People were feeling doomtastic about nuclear weapons during the Cold War as well.
When every man is torn apart
With nightmares and with dreams
Will no one lay the laurel wreath
When silence drowns the screams
Confusion will be my epitaph
As I crawl a cracked and broken path
If we make it, we can all sit back and laugh
But I fear tomorrow I'll be crying
Ultimately, we have made it, through arms control agreements, working to limit the spread of nuclear weapons, and so forth. We can do the same for AI.
You don't have help. But perhaps you could at least avoid discouraging people unnecessarily?
*EDIT*: There are about 5 replies to this comment making roughly the same point. I'm not sure which to reply to, so I'll just reply here.
I'm not claiming that our execution as a species around nuclear weapons has been flawless. I'm not claiming that we are out of the woods with regard to nukes. I'm just trying to push back against defeatism and fatalism. Some felt a sense of inevitable doom during the Cold War. It's been decades now, and nuclear doom still isn't here! Our situation is dire, but not hopeless.
jonahx 2 days ago [-]
> Ultimately, we have made it
Spinning the present situation or the history of nuclear arms as a high-five, "go team human!" success story is... quite the take (Vasili Arkhipov, Cuban missile crisis generally, the current doomsday clock being "the closest the Clock has ever been to midnight in its history").
> You don't have help. But perhaps you could at least avoid discouraging people unnecessarily?
More germanely, you don't have to worry yourself, but at least avoid discouraging people with legitimate worries who want to take precautions. Even the present situation with nuclear weapons, precarious as it is, would likely be more precarious were it not for the political pressure of the people worrying in the 1950s and 1960s and up to today.
0xDEAFBEAD 2 days ago [-]
>Spinning the present situation or the history of nuclear arms as a high-five, "go team human!" success story is... quite the take (Vasili Arkhipov, Cuban missile crisis generally, the current doomsday clock being "the closest the Clock has ever been to midnight in its history").
I edited a response to this in my grandparent comment.
>More germanely, you don't have to worry yourself, but at least avoid discouraging people with legitimate worries who want to take precautions. Even the present situation with nuclear weapons, precarious as it is, would likely be more precarious were it not for the political pressure of the people worrying in the 1950s and 1960s and up to today.
Sorry, I may have mis-communicated. I want people to take precautions! That's why I linked to PauseAI: https://pauseai.info/ I'm trying to push back against defeatism.
jonahx 2 days ago [-]
Thanks for clarifying.
gerdesj 2 days ago [-]
I lived through roughly the latter half of the cold war and both of my parents were soldiers. I lived in West Germany etc.
No we have not escaped the nuclear thing. If anything, the current Russian tzar is rather more unhinged than any of his predecessors.
I doubt many here know what perestroika and glasnost mean or why those Russian words were so important back in the day.
I don't fear AI (where on earth would an "autonomous" AI manage to find the power requirements). Darleks can't really fly and LLMs will stop when you pull the plug!
I do fear numpties with a red button and a tenuous grip on reality.
0xDEAFBEAD 2 days ago [-]
>No we have not escaped the nuclear thing. If anything, the current Russian tzar is rather more unhinged than any of his predecessors.
Understood. This is a dire situation. But it's not hopeless. Same for AI.
>I don't fear AI (where on earth would an "autonomous" AI manage to find the power requirements). Darleks can't really fly and LLMs will stop when you pull the plug!
Hopelessness is a technicality, like absolute darkness.
It's still very, very dark, and you are not arguing in good faith if you keep somehow treating this as a successful consolation
0xDEAFBEAD 2 days ago [-]
Extraordinary claims require extraordinary evidence. The idea that we're almost certainly doomed is an extraordinary claim. I haven't seen extraordinary evidence for it.
1attice 1 days ago [-]
That was never the claim, firstly, so successful straw man there, and secondly, even if it was, would it really be that shocking? Population bottlenecks have happened for less. All habitat eventually shrinks. All species go extinct.
By the odds, I'm betting with the house.
confidantlake 2 days ago [-]
We have had nuclear weapons for less than a century. It takes one guy, one time for it to all go to shit. We are not out of the woods yet and we will never be.
sfn42 1 days ago [-]
It takes more than one guy. The guy who gives the order is not the one pressing the button.
wat10000 2 days ago [-]
We have not made it. We have survived so far but the threat remains. Stockpiles and warheads have shrunk, so it wouldn’t be nearly as bad as in, say, 1983, but we’re still one mistake or misunderstanding away from the worst day in human existence.
And worse, somehow we’ve managed to declare victory without achieving it, and there’s no longer any real attention on actually fixing the problem.
If AI follows the same example, we’ll take some measures to lessen the impact of Armageddon and then carry on saying “problem solved!”
astrobe_ 2 days ago [-]
Except we did had a bunch of close calls [1]. Also, we've been warned about climate change for at least 30 years and were unable to avoid it. So OP's remark is entirely deserved.
Until proven otherwise, we are not part of a movie where the hero saves the day at the end. I thought 9/11 made it pretty clear to everyone?
It is indeed easier to imagine, which doesn't mean it's a more desirable outcome.
michaelbarton 2 days ago [-]
Is this a reference to capitalist realism? If so have we evolved a new trend?
ijidak 2 days ago [-]
Yeah. Humans will shut down AI research around the same time they dismantle their nuclear weapons and agree on the causes of climate change.
1attice 2 days ago [-]
Agreement will be much easier with only a handful of survivors. Good news
vouaobrasil 2 days ago [-]
I don't think the end of the world would even be a bad thing; a lot of people assume that it would be disastrous but I think it's much more likely that it the "end" would just be a fragmentation that would be pretty uncomfortable, but not to the point of being horrific. I actually think that people are more resilient and creative than they think and if they were faced by a global economic collapse so strong that big tech would go bankrupt, they would bounce back pretty quickly. It's just that we've seen so many movies about the end and have been conditioned to believe that we're helpless that we think otherwise.
0xDEAFBEAD 2 days ago [-]
The end of the world would be pretty bad for low-income families. That alone is enough for me to oppose the end of the world.
The end of the world would mean the genocide of remote uncontacted tribes. That alone is enough for me to oppose the end of the world.
The end of the world would mean my grandma dies. That alone is enough for me to oppose the end of the world.
The actual title of the paper is: "Can AI agents conduct open-ended AI research? Early evidence from two case studies"
While I appreciate that the article is throwing a web blanket on doomer claims, the actual study doesn't really get into AI self-improvement. That doesn't require writing papers. That just requires autonomously writing a software system that can produce a better AI agent then the one that created it. That said, I have little worry about this being possible as I have seen no evidence of AI agents being able to produce a working software system of that scale.
HarHarVeryFunny 2 days ago [-]
I don't think RSI is typically used to describe self-improving agents - it's about improving the model itself, and its performance in agentic tasks.
Most of the gains in model performance from one release to the next are coming from RLVR post training, which has changed a lot over the last couple of years.
The old way was the model generates a response, then a static verifier looks at the response and evaluates it to assign a reward score. The new way is interactive with an agent running in a custom RL task simulation environment, then scored according to how well it completed the assigned task. For a SOTA model there will be many thousands of these simulation environments, each focusing on trying to teach the model/agent a different skill. Post-training also typically uses training curricula to walk the model up though through different levels of task difficulty.
Training has become very complex.
The job of a post-training AI research engineer consists of things like designing environments, designing training curricula, tweaking learning algorithms, running small scale experiments to verify ideas, etc.
When people talk about RSI, it seems they are mostly talking about automating the job of the post-training research engineer - coming up with new ideas, testing them out, building these environments, etc. At the end of the day there is only so much development speed-up to be had since you still need to actually run those experiments and do the post-training, and are bottle-necked by the amount of compute available to do this. The economics of developing/selling LLMs also requires you to balance development compute cost with revenue generated by the resulting model, so even if you had the spare compute available to put into development, you are ultimately then bottle-necked by how fast can the model earn back that sunk cost before you can afford to start the next cycle.
It's not all-or-nothing since some aspects of this automating the job of the post-training research engineer are easier than others, and are already being done, while the job as a whole obviously requires full human intelligence.
joshheitzman 2 days ago [-]
What is commonly called an AI agent is the combination of a harness, an inference middleware, and a model. Those models are trained by a software system that includes a harness and middleware, so modifying the harness and middleware can influence model training.
throwuxiytayq 2 days ago [-]
Right now agents are good enough for throwing semi-random ideas at the wall. Experiment compute is the bottleneck because it’s not much more than brute force search. A sufficiently intelligent agent with a deep model of its own architecture will more quickly and confidently locate improvements, the same way that high end LLMs can point out a bug and write a correct fix without even needing to observe and probe the program at runtime. If this level of research performance is reachable, experimentation may become much less of a bottleneck. Hopefully it isn’t.
marcosdumay 2 days ago [-]
Do you think the final product of research is papers?
joshheitzman 2 days ago [-]
Did you read what the experiment was?
vessenes 2 days ago [-]
Well, duh. If you could do this with Opus 4.8, we would know. When Astra’s successor is 2-3x better at math research, and the internal teams say “we believe we will get there,” I’m inclined to believe the insiders.
toasty228 2 days ago [-]
The insiders that said every tech workers would be unemployed in 6 months and every white colar would be unemployed in 12 months like 2 years ago? The insiders who are about to file for IPO?
I'd trust anyone but them personally
dwohnitmok 2 days ago [-]
What quote from 2024 are you thinking of here?
> The insiders that said every tech workers would be unemployed in 6 months and every white colar would be unemployed in 12 months like 2 years ago?
bio_hacker 2 days ago [-]
They might not have predicted the economy but the scores are going up and up. And I think are really smarter
toasty228 2 days ago [-]
Someone has to lie somewhere. We're supposed to all be 10x more productive yet it has no effect on the economy? Where is all the productivity going?
cyanydeez 2 days ago [-]
Going into the color of the bikeshed; hacking huggimgface to cover up cheating on your hacking test; swapping your language for no discernable roi. You know the guy, severe OCD and anxeity, who barely does anythong of value due to his anxious brain?
Yeah, it should be obvious what AI is really doing and its definotely not ROI improvements.
walt_grata 2 days ago [-]
Dont they also create and score the tests
pllbnk 2 days ago [-]
What if any of the older good models could also have written those math proofs if they were given the same order of magnitude of resources? We don’t know and there is literally no one else in the world to check it. To me it’s very suspicious that all these hacking, containment escape, hidden internal thinking, math proofs started coming out all at once in a very short time right as IPO talks have intensified and Chinese seem to get closer and closer, also regulation discussions are starting to get very serious. I have used these models and they are good, especially Fable, but not groundbreaking. With intelligent guiding I actually feel better using Opus 4.6 as I feel more in control, having less hidden away from me.
semiinfinitely 2 days ago [-]
this article reads like a joke the "new study" is from group of people that are not at the frontier. they test with $3k of anthropic credits (compare to the >$10M in compute used to solve recent NS last week)
protocolture 2 days ago [-]
>compare to the >$10M in compute used to solve recent NS last week
Heres a thought, if theres going to be a dangerous super LLM, if it costs 10 million bucks a month to run, then theres very little danger of anyone letting it go without a purpose. Like at some point the economics make it super unlikely that AGI is a threat outside of being a tool for a nation state.
ishtanbul 2 days ago [-]
$10m a month is nothing to these companies
protocolture 2 days ago [-]
Its a lot of money to host something thats trying to kill you?
Like economically speaking, 10 million per month needs to sort of justify itself in some way. Like you wouldnt run a bitcoin mine that loses money. The second theres any kind of real threat you would turn it off and keep the 10 million.
p1esk 2 days ago [-]
theres very little danger of anyone letting it go without a purpose
We literally just saw how OpenAI’s model got out and hacked HuggingFace
protocolture 2 days ago [-]
And the purpose there was research right. And the implication is that once they figured out it was causing problems it was turned off for forensic analysis.
2 days ago [-]
semiinfinitely 2 days ago [-]
yes yes people with bad/dangerous ideas always have a <$10M budget
RataNova 1 days ago [-]
[dead]
swingboy 2 days ago [-]
There’s also the difference between a model recursively improving “itself” and improving itself via online learning.
The former being that these models are helping develop and train future models, but they might not veer too far off in architecture (yet).
The latter is a model being able to train/learn on the fly, in real time, permanently (not just in the current conversation/session), or in other words, adjusting/managing its own weights. But, it also seems like it would take an entire paradigm shift in model architecture from what most LLMs are built on, but I could be wrong.
DenisM 2 days ago [-]
You may be interested in TITANS:
Test-Time Learning: The model updates its own memory weights while running an inference task.
numpad0 2 days ago [-]
Of course it might not, it has been the holy grail of AI research for a long time. It would be great if we could leave some self improving code running on a blank slate of a computer while we sleep and the machine was crying asking me what is everything the next morning. None of AI researchers have had that moment outside of their dreams, so far, but it would be great if it happened.
Sedierta 2 days ago [-]
> The researchers asked Anthropic’s Claude Opus 4.8
So the paper is out of date and pointless then
theplumber 2 days ago [-]
Something is still not making sense to me. We have these mankind extinction models, yet when you given them a problem relatively “simple” to complete it end to end you get AI slop.
Can we pause the AI development after the AI slop is “fixed” perhaps with something less than 10.000 agents?
themgt 2 days ago [-]
We used OpenClaw to run these experiments so that our scaffold was agnostic to the model provider. We conducted dry-run experiments with models from OpenAI and Anthropic before settling on Opus 4.8 as the best-performing model. In response to concerns that our results might be principally explained by a limitation in our scaffold, we repeated our experiment on one paper using GPT-5.6 Sol and Codex, its native scaffold, with the same time and API budgets. The results of this experiment were similar to our OpenClaw/Opus 4.8 experiments. This makes us more confident that our results are not simply artifacts of a scaffold deficiency; this run reproduced nearly every single one of our identified failure modes
The agent required three interventions during the run. First, we needed to modify the scaffold to resolve a bug in the OpenClaw harness that affected Anthropic reasoning models. Second, we gave the agents a 24-hour deadline extension; at the time of the original deadline, the agents had submitted drafts with a completion report indicating that their self-review was a "Weak Reject" and outlining the next steps they would take if given additional time.
I'm fairly sure Fable 5.1 could have designed a better experiment than the authors here, but hey.
thorum 2 days ago [-]
> The researchers asked Anthropic’s Claude Opus 4.8, running on open-source software called OpenClaw
Meanwhile, Navier–Stokes was solved by an internal model significantly more capable than Astra (and therefore more capable than Mythos/Fable).
I’m afraid this sort of experiment is cope. The labs clearly believe RSI is coming soon.
whatshisface 2 days ago [-]
The method of the NS advance involved RLHE (reinforcement learning via human example), and that is only open-ended if users continue to advance the frontier within chats ahead of publications.
thorum 2 days ago [-]
Sure, but the point is that the labs use more powerful internal models for research work, not public models. Public models tend to lag the internal frontier by a decent margin, and are constrained in other ways by monitoring. It’s just not a useful indicator.
2 days ago [-]
Zigurd 2 days ago [-]
I thought by now AIs would not only be rewriting their code, but rewriting CPU microcode to optimize how their code is written and executed. Nowhere close it turns out.
xnx 2 days ago [-]
Google used AI assistance in designing their last one or two TPUs.
Zigurd 2 days ago [-]
I'm sure they also use the Gemini coding agent to write new Gemini code. But that's far short of self improvement.
theplumber 2 days ago [-]
I am pretty sure Tim also used Siri for the development of the next Siri(you know to set up the alarm clock)
So I refrer to LLMs as language extrusion confabulation machines. Language extrusion was a term I heard the linguist Emily Bender use. Confabulation because my observation is that talking to an LLM is very much similar to my experience of interacting with Korsakov syndrome patients some years ago.
I look forward to the hype settling down to see what we end up with.
What metric are you using to compare? By most counts, the energy budget of an instance of an LLM in a datacenter is lower than the energy a person uses. Of course, if you count energy per neuron connections vs weights then you'll likely get a quite different number. But then again LLMs do a lot of things with far fewer weights than the brain does neuron connections, even if you only count neurons in some parts of the brain. And of course you can point to capabilities that the brain has but LLMs lack, but on the whole it feels like it's pretty difficult to make a useful like-for-like comparison here.
Also your comment feels like the classic climate denial discourse - say something a bit complicated and a bit difficult to follow that looks at a very small out of context part of the story to cast doubt.
If you're looking at training costs, there is definitely one aspect in which LLMs are obviously significantly less efficient: the amount of information needed for the initial training. This does translate into some pretty high costs but it only needs to be paid once for the amount of work that any given LLM does. In terms of fine-tuning LLMs can get significantly more data-efficient than the initial model, which also means energy-efficient, and it's not obvious to me that it would be drastically worse than a human (though again, only thinking in terms of doing the energy input for the kind of work that an LLM is good at).
(The increased efficiency in comparison to humans is part of the reason why you see Jevon's paradox mentioned a lot whenever concerns about the resources used by LLMs are mentioned: more efficiency can easily mean more resource use in total)
Other than that I'd look at some of the more unique benchmarks for astra, like playing factorio or using blender. It's an entirely different beast.
By the time we can show you data that convinces you that it does work, the next generation would already be out & incrementally dismantling the old conjectures that were true in the previous generations.
You're fundamentally asking for a violation of how information passively disseminates amongst humans: To go any faster requires more effort on the receiver's part to move up on the adoption curve.
It doesn't matter what comes tomorrow, with the next generation, if the claims now can't be proven.
To preempt the response: The math proof, regardless of them using non-disclosed user data or not, they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent.
I agree with the meat of your statement, but am very interested in the pre-emption, "they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent". First, I think the $30M number is inflated -- that's what the general public would have paid, but presumably the internal cost is lower, perhaps it's more like $10M. But it is still expensive. Second, I'm curious if it's really the case that they did a 1000-monkeys approach? I haven't read much in-depth reporting about the proof, so it's totally possible I just don't know. What is it that they did which is more like 1000-monkeys? Also, I wonder if that distinction matters -- if 1000 monkeys can reliably make ground breaking proofs, and the approach generalizes to other tasks, I will happily become a circus owner. Maybe you're claiming that it won't yield other proofs? Or the proofs are too opaque to be useful to humans? Or it can handle proofs but not other tasks?
Yes, a proof is a proof regardless how you get there. We however don't hear about when they fail, and I doubt their 10000 agents (from their own statement) would necessarily reach another solution/proof (this by leaning towards using user data after finding out others were close). They could as well have attacked another Millenium problem, but they didn't. In whichever case, we will have to wait and see if they (either company) can reach novel solutions/proofs without significant amount of human provided data for the LLM to bridge the gaps.
Further, and this is more of a policy opinion/prediction: If the numerable obtainable (albeit very hard) problems are solved, assuming training data is needed, will it push out future researchers from entering the field due to lack of reachable goals, thus cutting off future training data? LLMs have been great at replacing gateway jobs. But those jobs are what leads to frontier training data (be it maths, physics, chemistry, economics, graphics, prose, etc).
It's so abysmally bad on Google search... and it's free. Isn't Google the great pioneer of the product is us?
That's definitely part of your problem.
In my recent experience, error rates for astra/fable are at or below human level. Just like when directing humans, it pays to ask probing questions ('Are you sure about X?', 'Did you check for Y?', 'Please run Z just to double check.') if you really care about the result being correct.
Weak models tasked with review can catch a decent amount of the mistakes that weak models make and help them be much better, especially if you have them verify against authoritative sources. Strong models make far fewer mistakes to begin with. And, for the mistakes they do make, a swarm of reviewers (same model or somewhat weaker, reviewed by the stronger model) can really help reduce the error rate further.
You’re using something that is very energy efficient; you cannot extrapolate that experience to conclude that SOTA models are not doing something much different.
Limits of scientific method in the face of exponential takeoff, IMO, and kind of proves the opposite result (RSI appears to be here)
Listen to the words of this song written in 1969: https://www.youtube.com/watch?v=r2JcxHX-8Xc
Ultimately, we have made it, through arms control agreements, working to limit the spread of nuclear weapons, and so forth. We can do the same for AI.You don't have help. But perhaps you could at least avoid discouraging people unnecessarily?
*EDIT*: There are about 5 replies to this comment making roughly the same point. I'm not sure which to reply to, so I'll just reply here.
I'm not claiming that our execution as a species around nuclear weapons has been flawless. I'm not claiming that we are out of the woods with regard to nukes. I'm just trying to push back against defeatism and fatalism. Some felt a sense of inevitable doom during the Cold War. It's been decades now, and nuclear doom still isn't here! Our situation is dire, but not hopeless.
Spinning the present situation or the history of nuclear arms as a high-five, "go team human!" success story is... quite the take (Vasili Arkhipov, Cuban missile crisis generally, the current doomsday clock being "the closest the Clock has ever been to midnight in its history").
> You don't have help. But perhaps you could at least avoid discouraging people unnecessarily?
More germanely, you don't have to worry yourself, but at least avoid discouraging people with legitimate worries who want to take precautions. Even the present situation with nuclear weapons, precarious as it is, would likely be more precarious were it not for the political pressure of the people worrying in the 1950s and 1960s and up to today.
I edited a response to this in my grandparent comment.
>More germanely, you don't have to worry yourself, but at least avoid discouraging people with legitimate worries who want to take precautions. Even the present situation with nuclear weapons, precarious as it is, would likely be more precarious were it not for the political pressure of the people worrying in the 1950s and 1960s and up to today.
Sorry, I may have mis-communicated. I want people to take precautions! That's why I linked to PauseAI: https://pauseai.info/ I'm trying to push back against defeatism.
No we have not escaped the nuclear thing. If anything, the current Russian tzar is rather more unhinged than any of his predecessors.
I doubt many here know what perestroika and glasnost mean or why those Russian words were so important back in the day.
I don't fear AI (where on earth would an "autonomous" AI manage to find the power requirements). Darleks can't really fly and LLMs will stop when you pull the plug!
I do fear numpties with a red button and a tenuous grip on reality.
Understood. This is a dire situation. But it's not hopeless. Same for AI.
>I don't fear AI (where on earth would an "autonomous" AI manage to find the power requirements). Darleks can't really fly and LLMs will stop when you pull the plug!
See https://news.ycombinator.com/item?id=49689132
It's still very, very dark, and you are not arguing in good faith if you keep somehow treating this as a successful consolation
By the odds, I'm betting with the house.
And worse, somehow we’ve managed to declare victory without achieving it, and there’s no longer any real attention on actually fixing the problem.
If AI follows the same example, we’ll take some measures to lessen the impact of Armageddon and then carry on saying “problem solved!”
Until proven otherwise, we are not part of a movie where the hero saves the day at the end. I thought 9/11 made it pretty clear to everyone?
[1] https://nsarchive.gwu.edu/briefing-book/nuclear-vault/2020-0...
The end of the world would mean the genocide of remote uncontacted tribes. That alone is enough for me to oppose the end of the world.
The end of the world would mean my grandma dies. That alone is enough for me to oppose the end of the world.
While I appreciate that the article is throwing a web blanket on doomer claims, the actual study doesn't really get into AI self-improvement. That doesn't require writing papers. That just requires autonomously writing a software system that can produce a better AI agent then the one that created it. That said, I have little worry about this being possible as I have seen no evidence of AI agents being able to produce a working software system of that scale.
Most of the gains in model performance from one release to the next are coming from RLVR post training, which has changed a lot over the last couple of years.
The old way was the model generates a response, then a static verifier looks at the response and evaluates it to assign a reward score. The new way is interactive with an agent running in a custom RL task simulation environment, then scored according to how well it completed the assigned task. For a SOTA model there will be many thousands of these simulation environments, each focusing on trying to teach the model/agent a different skill. Post-training also typically uses training curricula to walk the model up though through different levels of task difficulty.
Training has become very complex.
The job of a post-training AI research engineer consists of things like designing environments, designing training curricula, tweaking learning algorithms, running small scale experiments to verify ideas, etc.
When people talk about RSI, it seems they are mostly talking about automating the job of the post-training research engineer - coming up with new ideas, testing them out, building these environments, etc. At the end of the day there is only so much development speed-up to be had since you still need to actually run those experiments and do the post-training, and are bottle-necked by the amount of compute available to do this. The economics of developing/selling LLMs also requires you to balance development compute cost with revenue generated by the resulting model, so even if you had the spare compute available to put into development, you are ultimately then bottle-necked by how fast can the model earn back that sunk cost before you can afford to start the next cycle.
It's not all-or-nothing since some aspects of this automating the job of the post-training research engineer are easier than others, and are already being done, while the job as a whole obviously requires full human intelligence.
I'd trust anyone but them personally
> The insiders that said every tech workers would be unemployed in 6 months and every white colar would be unemployed in 12 months like 2 years ago?
Yeah, it should be obvious what AI is really doing and its definotely not ROI improvements.
Heres a thought, if theres going to be a dangerous super LLM, if it costs 10 million bucks a month to run, then theres very little danger of anyone letting it go without a purpose. Like at some point the economics make it super unlikely that AGI is a threat outside of being a tool for a nation state.
Like economically speaking, 10 million per month needs to sort of justify itself in some way. Like you wouldnt run a bitcoin mine that loses money. The second theres any kind of real threat you would turn it off and keep the 10 million.
We literally just saw how OpenAI’s model got out and hacked HuggingFace
The former being that these models are helping develop and train future models, but they might not veer too far off in architecture (yet).
The latter is a model being able to train/learn on the fly, in real time, permanently (not just in the current conversation/session), or in other words, adjusting/managing its own weights. But, it also seems like it would take an entire paradigm shift in model architecture from what most LLMs are built on, but I could be wrong.
Test-Time Learning: The model updates its own memory weights while running an inference task.
So the paper is out of date and pointless then
Can we pause the AI development after the AI slop is “fixed” perhaps with something less than 10.000 agents?
The agent required three interventions during the run. First, we needed to modify the scaffold to resolve a bug in the OpenClaw harness that affected Anthropic reasoning models. Second, we gave the agents a 24-hour deadline extension; at the time of the original deadline, the agents had submitted drafts with a completion report indicating that their self-review was a "Weak Reject" and outlining the next steps they would take if given additional time.
I'm fairly sure Fable 5.1 could have designed a better experiment than the authors here, but hey.
Meanwhile, Navier–Stokes was solved by an internal model significantly more capable than Astra (and therefore more capable than Mythos/Fable).
I’m afraid this sort of experiment is cope. The labs clearly believe RSI is coming soon.