• Log InLog In
  • Register
Liquid`
Team Liquid Liquipedia
EDT 21:09
CEST 03:09
KST 10:09
  • Home
  • Forum
  • Calendar
  • Streams
  • Liquipedia
  • Features
  • Store
  • EPT
  • TL+
  • StarCraft 2
  • Brood War
  • Smash
  • Heroes
  • Counter-Strike
  • Overwatch
  • Liquibet
  • Fantasy StarCraft
  • TLPD
  • StarCraft 2
  • Brood War
  • Blogs
Forum Sidebar
Events/Features
News
Featured News
Serral wins HomeStory Cup 2914Serral wins Maestros of the Game 243ByuL, and the Limitations of Standard Play3Team Liquid Map Contest #22: Results and Winners7Code S Season 2 (2026): RO4 and Finals Preview12
Community News
SC4ALL II announced - $10,000 prize pool, Dec 5-61PIG STY FESTIVAL 8.0! (13 - 23 August)8Neeb returns to progaming; rejoins ONSYDE14Weekly Cups (July 20-26): Early returns on 5.0.16b8IntoTheTV X SOOP SC2 League : Weekly & Monthly5
StarCraft 2
General
SC4ALL II: StarCraft 2 Player Announcement 1/8 Balance hotfix patch 5.0.16b (July 16) Neeb returns to progaming; rejoins ONSYDE Clem: "I don't have that much hope in Blizzard" Weekly Cups (July 20-26): Early returns on 5.0.16b
Tourneys
SC4ALL II announced - $10,000 prize pool, Dec 5-6 PIG STY FESTIVAL 8.0! (13 - 23 August) IntoTheTV X SOOP SC2 League : Weekly & Monthly Crank Gathers Season 4: BW vs SC2 Team League RSL Revival: Season 6 - Qualifiers and Main Event
Strategy
[G] Having the right mentality to improve
Custom Maps
Nexus Wars 2021 GUIDE [M] (2) Industrial Park
External Content
Mutation # 536 Railroad Switch The PondCast: SC2 News & Results Mutation # 535 Assembly of Vengeance Mutation # 534 Burning Evacuation
Brood War
General
BW Drama: Terror's Debt Incident + C9 disbanding BGH Auto Balance -> http://bghmmr.eu/ Making an Online Broodwar Manager Game BW General Discussion ASL22 General Discussion
Tourneys
[Megathread] Daily Proleagues BSL LAN Party - Kraków 29-30 August - OPEN SIGNUPS Escore Tournament - Season 3 2v2v2v2 Tournament
Strategy
Fighting Spirit mining rates Odyssey Mineral Stack Saturation Simple Questions, Simple Answers PvT advise for noobs
Other Games
General Games
Beyond All Reason Nintendo Switch Thread Path of Exile ZeroSpace Early Access is Now Live! Stormgate/Frost Giant Megathread
Dota 2
Looking for a Dota Mentor Official 'what is Dota anymore' discussion
League of Legends
[TL LoL EUW IHs] Teemo shall perish TSM pausing esports and CLG Dead
Heroes of the Storm
Heroes of the Storm 2.0
Hearthstone
Deck construction bug
TL Mafia
TL Mafia Community Thread TL Mafia Power Rank NeO.D_StephenKing vs This Guy From 1 Million Dance
Community
General
US Politics Mega-thread Russo-Ukrainian War Thread Artificial Intelligence Thread Things Aren’t Peaceful in Palestine Canadian Politics Mega-thread
Fan Clubs
The IdrA Fan Club The HerO Fan Club!
Media & Entertainment
Movie Discussion! Anime Discussion Thread Series you have seen recently... [Req][Books] Good Fantasy/SciFi books
Sports
2026-27 Football Thread placeholder 2024 - 2026 Football Thread TeamLiquid Health and Fitness Initiative For 2023 Formula 1 Discussion MLB/Baseball 2023
World Cup 2022
Tech Support
Computer Build, Upgrade & Buying Resource Thread Simple Questions Simple Answers FPS when play League Of Legend on laptop
TL Community
Northern Ireland Global Starcraft The Automated Ban List
Blogs
What is a Gamer?
TrAiDoS
Hello guys!
LIN1s
ASL S22 English Commentary…
namkraft
Poker (part 2)
Nebuchad
An Exploration of th…
waywardstrategy
Customize Sidebar...

Website Feedback

Closed Threads



Active: 7582 users

Artificial Intelligence Thread - Page 5

Forum Index > General Forum
Post a Reply
Prev 1 2 3 4 5 All
GreenHorizons
Profile Blog Joined April 2011
United States24211 Posts
11 hours ago
#81
AI doesn't even need to be that powerful or misaligned. Capitalism acts as a sort of multiplier so that we're basically already trapped in a paperclip scenario with data centers.
"People like to look at history and think 'If that was me back then, I would have...' We're living through history, and the truth is, whatever you are doing now is probably what you would have done then" "Scratch a Liberal..."
Manit0u
Profile Blog Joined August 2004
Poland17813 Posts
11 hours ago
#82
On July 31 2026 21:59 Jankisa wrote:
Also, I'd push back on the "AI can't innovate", I think that Stockfish and Alphago and Alphastar have demonstrated plenty of innovative and novel tactics a loooong time ago, and again, if you read the summary of how GPT attacked Huggingface it literally found 2 new zero day exploits, in a lot of cases innovation is just trying a bunch of things until one of them works, so AI throwing massive amounts of compute at a problem is basically the same thing as innovation.


The bugs it found are just old bugs but in different software. So it did what it's best at: pattern matching. Checked old bugs and tried to find software that might be vulnerable to those as well. People are using AI to find plenty of such things nowadays because finally they can leverage AI to do something that would be too time-consuming previously.
Time is precious. Waste it wisely.
Jankisa
Profile Blog Joined October 2010
Croatia1559 Posts
10 hours ago
#83
You really don't know that, so you are kind of innovating a narrative in order to reinforce your preconceived notions.

To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.


So maybe what you said is true and you work in OpenAI and have some sort of inside knowledge, but I honestly doubt that.

Also you haven't really addressed the fact that AI's have been coming up with never before seen tactics, openings and strategies for games almost 10 years ago.

The list of things AI has done is pretty long at this point, so I'll just go with one I find most impressive because I'm terrible at math, it solving a 80-year-old math problem:

“No previous AI-generated proof has come close” to meeting those high standards, wrote Timothy Gowers, a mathematician at the University of Cambridge, in commentary solicited by OpenAI.

“This is the unique interesting result produced autonomously by AI so far,” says Daniel Litt, a mathematician at the University of Toronto, who was consulted by OpenAI to verify the proof but is not involved with the company.


It's not just about things being time consuming, it can genuinely find solutions to things that humans couldn't for a looong time.
So, are you a pessimist? - On my better days. Are you a nihilist? - Not as much as I should be.
Uldridge
Profile Blog Joined January 2011
Belgium5198 Posts
10 hours ago
#84
It has a lot of time to bruteforce on thing that already exist.
It will also ultimately use the given ruleset in a way no human would ever think about.

So in this sense it's innovating, but it's still within the ruleset we created.

I'll be truly impressed when we start seeing advanced maths or whatever that are just completely new (just like how imaginary numbers were created), or if it comes up with a grand unifying theory using completely novel physics to explain it all.
Taxes are for Terrans
Manit0u
Profile Blog Joined August 2004
Poland17813 Posts
Last Edited: 2026-07-31 14:53:33
10 hours ago
#85
On July 31 2026 23:09 Jankisa wrote:
You really don't know that, so you are kind of innovating a narrative in order to reinforce your preconceived notions.

Show nested quote +
To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.


So maybe what you said is true and you work in OpenAI and have some sort of inside knowledge, but I honestly doubt that.

Also you haven't really addressed the fact that AI's have been coming up with never before seen tactics, openings and strategies for games almost 10 years ago.

The list of things AI has done is pretty long at this point, so I'll just go with one I find most impressive because I'm terrible at math, it solving a 80-year-old math problem:

Show nested quote +
“No previous AI-generated proof has come close” to meeting those high standards, wrote Timothy Gowers, a mathematician at the University of Cambridge, in commentary solicited by OpenAI.

“This is the unique interesting result produced autonomously by AI so far,” says Daniel Litt, a mathematician at the University of Toronto, who was consulted by OpenAI to verify the proof but is not involved with the company.


It's not just about things being time consuming, it can genuinely find solutions to things that humans couldn't for a looong time.


Have you checked those stories? For 6 Erdos problems "solved" by AI it turned out 5 were already solved in mathematical literature previously. Just not associated with Erdos but as other things so people just missed them.

Most of those problems are also deemed not actually worth solving by humans because the required time investment is not proportional to the complexity of the problem (need to spend a lot of time doing tedious things to prove/disprove something).

You don't have to take my word for it, here's a CS professor explaining it:


It's like I said previously. AI is great at scouring large amounts of data and connecting the dots there or doing simulations of stuff that would take humans too long to be worthwhile. That's how it finds the bugs in the code, that's how it solves those "unsolved" mathematical problems.

Humans simply don't have the capacity to access and keep track of so much data at once so a lot of existing mathematical problems have already been solved by people who simply didn't even know they were a problem while doing something else. And then it gets lost because people trying to tackle the problem don't necessarily follow seemingly unrelated works so the problem stays open. In this way AI is a great tool for assisting with such things but it's not really solving anything by itself. Even for the one problem that wasn't "solved" yet it didn't disprove the original thesis but instead proposed an alternative.

And regarding the HuggingFace hack:

First, circumventing internet restrictions and hacking into servers are exactly the kinds of things these ExploitGym systems are designed to do. There was no “rogue” agent or revelation of some surprising, devious new capability.

Second, the real issue here was OpenAI’s sloppiness. What makes ExploitGym a hard benchmark is that there aren’t supposed to be humans in the loop–you have to let your harness and LLM act entirely on their own, coming up with long-time-horizon plans and executing them autonomously. (When professional programmers use coding harnesses, by contrast, there’s plenty of human oversight, as LLM-based plans are often misaligned with our intentions, or just plain weird, and need correcting.)

AI companies know LLM-based plans are pretty unpredictable, so they’re usually pretty careful about how they set up and restrict the harnesses and LLMs used in ExploitGym-style tests.

According to ​reporting​ from the Financial Times, however, OpenAI recently started playing fast and loose with these safety principles in a bid to catch up to Anthropic in this area – Anthropic having received a lot of cybersecurity street cred from ​the buzz​ surrounding its Mythos release.

The Financial Times noted that OpenAI had used “increasingly aggressive training methods in its race against Anthropic,” and that it had been warned that their approach could lead to a “breakaway hacking incident” after early tests showed it lacked the right safeguards to block plans that involved bypassing constraints in the test environment. Given these concerns, OpenAI staff were reportedly “unsurprised” by the Hugging Face incident.

In other words, the AI companies running these challenges already knew that, without care, these unpredictable autonomous hacking systems might bypass constraints and attack systems you didn’t intend to target. Blindly implementing an LLM-generated plan with a powerful harness is dicey – not because the LLM might develop malicious intent (this is a nonsensical notion given their static architecture), but because as any ChatBot user knows, LLMs are unpredictable. OpenAI wasn’t sufficiently careful, and got burned.
Time is precious. Waste it wisely.
BradTheBaneling
Profile Joined October 2018
42 Posts
10 hours ago
#86
On July 31 2026 23:17 Uldridge wrote:
It has a lot of time to bruteforce on thing that already exist.
It will also ultimately use the given ruleset in a way no human would ever think about.

So in this sense it's innovating, but it's still within the ruleset we created.

I'll be truly impressed when we start seeing advanced maths or whatever that are just completely new (just like how imaginary numbers were created), or if it comes up with a grand unifying theory using completely novel physics to explain it all.


I don't understand how "innovating within a ruleset" is different from complex numbers. Complex numbers are a natural result of the reals and algebra.

I also think going from complex numbers to a GUT is like going from fire to the Apollo 11 missions. I don't know if your modern example is a very fair bar to set.

Arguably all math after we accept some set of axioms is just 'innovating within a ruleset'.

If you wanted to downplay it, I think the better argument to use is that LLMs appear to be far better at doing things like finding counterexamples vs. finding proofs.
Uldridge
Profile Blog Joined January 2011
Belgium5198 Posts
9 hours ago
#87
On July 31 2026 23:56 BradTheBaneling wrote:
Show nested quote +
On July 31 2026 23:17 Uldridge wrote:
It has a lot of time to bruteforce on thing that already exist.
It will also ultimately use the given ruleset in a way no human would ever think about.

So in this sense it's innovating, but it's still within the ruleset we created.

I'll be truly impressed when we start seeing advanced maths or whatever that are just completely new (just like how imaginary numbers were created), or if it comes up with a grand unifying theory using completely novel physics to explain it all.


I don't understand how "innovating within a ruleset" is different from complex numbers. Complex numbers are a natural result of the reals and algebra.

I also think going from complex numbers to a GUT is like going from fire to the Apollo 11 missions. I don't know if your modern example is a very fair bar to set.

Arguably all math after we accept some set of axioms is just 'innovating within a ruleset'.

If you wanted to downplay it, I think the better argument to use is that LLMs appear to be far better at doing things like finding counterexamples vs. finding proofs.


You say that like it's obvious to to just expand how things work with the "same rules" when complex numbers allow you to do things that weren't possible before they were introduced. The beauty is it having all make sense. If you introduce something new and it breaks something else and then you need to patch that... you get modern game balance teams. But in all seriousness, I think one needs to be quite smart and creative to be able to innovate instead of derive, which, to g8ve credit, can sometimes be quite innovative in its own right lol
Taxes are for Terrans
Jankisa
Profile Blog Joined October 2010
Croatia1559 Posts
8 hours ago
#88
On July 31 2026 23:40 Manit0u wrote:
Show nested quote +
On July 31 2026 23:09 Jankisa wrote:

You really don't know that, so you are kind of innovating a narrative in order to reinforce your preconceived notions.

To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.


So maybe what you said is true and you work in OpenAI and have some sort of inside knowledge, but I honestly doubt that.

Also you haven't really addressed the fact that AI's have been coming up with never before seen tactics, openings and strategies for games almost 10 years ago.

The list of things AI has done is pretty long at this point, so I'll just go with one I find most impressive because I'm terrible at math, it solving a 80-year-old math problem:

“No previous AI-generated proof has come close” to meeting those high standards, wrote Timothy Gowers, a mathematician at the University of Cambridge, in commentary solicited by OpenAI.

“This is the unique interesting result produced autonomously by AI so far,” says Daniel Litt, a mathematician at the University of Toronto, who was consulted by OpenAI to verify the proof but is not involved with the company.


It's not just about things being time consuming, it can genuinely find solutions to things that humans couldn't for a looong time.


Have you checked those stories? For 6 Erdos problems "solved" by AI it turned out 5 were already solved in mathematical literature previously. Just not associated with Erdos but as other things so people just missed them.

Most of those problems are also deemed not actually worth solving by humans because the required time investment is not proportional to the complexity of the problem (need to spend a lot of time doing tedious things to prove/disprove something).

You don't have to take my word for it, here's a CS professor explaining it:
https://www.youtube.com/watch?v=fhZRWZ6J4k4

It's like I said previously. AI is great at scouring large amounts of data and connecting the dots there or doing simulations of stuff that would take humans too long to be worthwhile. That's how it finds the bugs in the code, that's how it solves those "unsolved" mathematical problems.

Humans simply don't have the capacity to access and keep track of so much data at once so a lot of existing mathematical problems have already been solved by people who simply didn't even know they were a problem while doing something else. And then it gets lost because people trying to tackle the problem don't necessarily follow seemingly unrelated works so the problem stays open. In this way AI is a great tool for assisting with such things but it's not really solving anything by itself. Even for the one problem that wasn't "solved" yet it didn't disprove the original thesis but instead proposed an alternative.

And regarding the HuggingFace hack:

First, circumventing internet restrictions and hacking into servers are exactly the kinds of things these ExploitGym systems are designed to do. There was no “rogue” agent or revelation of some surprising, devious new capability.

Second, the real issue here was OpenAI’s sloppiness. What makes ExploitGym a hard benchmark is that there aren’t supposed to be humans in the loop–you have to let your harness and LLM act entirely on their own, coming up with long-time-horizon plans and executing them autonomously. (When professional programmers use coding harnesses, by contrast, there’s plenty of human oversight, as LLM-based plans are often misaligned with our intentions, or just plain weird, and need correcting.)

AI companies know LLM-based plans are pretty unpredictable, so they’re usually pretty careful about how they set up and restrict the harnesses and LLMs used in ExploitGym-style tests.

According to ​reporting​ from the Financial Times, however, OpenAI recently started playing fast and loose with these safety principles in a bid to catch up to Anthropic in this area – Anthropic having received a lot of cybersecurity street cred from ​the buzz​ surrounding its Mythos release.

The Financial Times noted that OpenAI had used “increasingly aggressive training methods in its race against Anthropic,” and that it had been warned that their approach could lead to a “breakaway hacking incident” after early tests showed it lacked the right safeguards to block plans that involved bypassing constraints in the test environment. Given these concerns, OpenAI staff were reportedly “unsurprised” by the Hugging Face incident.

In other words, the AI companies running these challenges already knew that, without care, these unpredictable autonomous hacking systems might bypass constraints and attack systems you didn’t intend to target. Blindly implementing an LLM-generated plan with a powerful harness is dicey – not because the LLM might develop malicious intent (this is a nonsensical notion given their static architecture), but because as any ChatBot user knows, LLMs are unpredictable. OpenAI wasn’t sufficiently careful, and got burned.


I guess we have a very different interpretation of what is impressive and how innovation works. Even then, you consistently skip over the examples I provided with AI systems coming up with novel ways of playing the games of Chess, Go and Starcraft 2, to varying success.

Also, your quote about the HF incident didn't really say anything except what I noted when I first posted about it, that OpenAI and other labs are very lax with their controls.

The AI (and Mythos and other systems) are by far, from all human activities biggest experts in coding, and when tasked with breaking systems they do find novel exploits, you asserted that they used known ones without any proof or source, when I asked for it you provided a quote that does not support what you said.

Asking for some "breakthroughs that no one saw coming" in the age where these are incredibly rare and have been mostly incremental and "group projects" for us humans is in my opinion, at this point especially kind of silly.

LLMs went from having trouble solving extremely easy math problems 2 years ago to solving things that 80 years of mathematicians trying couldn't, it's weird to dismiss them.
So, are you a pessimist? - On my better days. Are you a nihilist? - Not as much as I should be.
BradTheBaneling
Profile Joined October 2018
42 Posts
Last Edited: 2026-07-31 17:36:01
7 hours ago
#89
On August 01 2026 00:29 Uldridge wrote:
Show nested quote +
On July 31 2026 23:56 BradTheBaneling wrote:
On July 31 2026 23:17 Uldridge wrote:
It has a lot of time to bruteforce on thing that already exist.
It will also ultimately use the given ruleset in a way no human would ever think about.

So in this sense it's innovating, but it's still within the ruleset we created.

I'll be truly impressed when we start seeing advanced maths or whatever that are just completely new (just like how imaginary numbers were created), or if it comes up with a grand unifying theory using completely novel physics to explain it all.


I don't understand how "innovating within a ruleset" is different from complex numbers. Complex numbers are a natural result of the reals and algebra.

I also think going from complex numbers to a GUT is like going from fire to the Apollo 11 missions. I don't know if your modern example is a very fair bar to set.

Arguably all math after we accept some set of axioms is just 'innovating within a ruleset'.

If you wanted to downplay it, I think the better argument to use is that LLMs appear to be far better at doing things like finding counterexamples vs. finding proofs.


You say that like it's obvious to to just expand how things work with the "same rules" when complex numbers allow you to do things that weren't possible before they were introduced. The beauty is it having all make sense. If you introduce something new and it breaks something else and then you need to patch that... you get modern game balance teams. But in all seriousness, I think one needs to be quite smart and creative to be able to innovate instead of derive, which, to g8ve credit, can sometimes be quite innovative in its own right lol


I mean the idea of complex numbers is just an algebraic closure of the real numbers.

Complex numbers are older than the fundamental theorem of algebra.

I don't really know what "when complex numbers allow you to do things that weren't possible before they were introduced" means. Complex numbers were discovered because we were trying to solve cubic equations.

The beauty is it having all make sense. If you introduce something new and it breaks something else and then you need to patch that


Again I just don't understand what this means. Are you suggesting that the mathematical results derived from LLMs so far are being falsely verified by mathematicians? What part of mathematics are you suggesting that is has literally a single iota of relevance to?

But in all seriousness, I think one needs to be quite smart and creative to be able to innovate instead of derive, which, to g8ve credit, can sometimes be quite innovative in its own right lol


This just feels like a goofy sort of philosophical argument. How do you define innovate and how do you define derive?

It feels like you're saying complex numbers were 'innovated' when you could easily argue (and I'm being particularly non-rigorous here) that they were derived from having proven solutions for specific cubic equations and then being able to demonstrate that those equations were also equal to simpler equations of real numbers and negative square root numbers (i.e. complex numbers).
Manit0u
Profile Blog Joined August 2004
Poland17813 Posts
7 hours ago
#90
On August 01 2026 01:39 Jankisa wrote:
Show nested quote +
On July 31 2026 23:40 Manit0u wrote:
On July 31 2026 23:09 Jankisa wrote:

You really don't know that, so you are kind of innovating a narrative in order to reinforce your preconceived notions.

To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.


So maybe what you said is true and you work in OpenAI and have some sort of inside knowledge, but I honestly doubt that.

Also you haven't really addressed the fact that AI's have been coming up with never before seen tactics, openings and strategies for games almost 10 years ago.

The list of things AI has done is pretty long at this point, so I'll just go with one I find most impressive because I'm terrible at math, it solving a 80-year-old math problem:

“No previous AI-generated proof has come close” to meeting those high standards, wrote Timothy Gowers, a mathematician at the University of Cambridge, in commentary solicited by OpenAI.

“This is the unique interesting result produced autonomously by AI so far,” says Daniel Litt, a mathematician at the University of Toronto, who was consulted by OpenAI to verify the proof but is not involved with the company.


It's not just about things being time consuming, it can genuinely find solutions to things that humans couldn't for a looong time.


Have you checked those stories? For 6 Erdos problems "solved" by AI it turned out 5 were already solved in mathematical literature previously. Just not associated with Erdos but as other things so people just missed them.

Most of those problems are also deemed not actually worth solving by humans because the required time investment is not proportional to the complexity of the problem (need to spend a lot of time doing tedious things to prove/disprove something).

You don't have to take my word for it, here's a CS professor explaining it:
https://www.youtube.com/watch?v=fhZRWZ6J4k4

It's like I said previously. AI is great at scouring large amounts of data and connecting the dots there or doing simulations of stuff that would take humans too long to be worthwhile. That's how it finds the bugs in the code, that's how it solves those "unsolved" mathematical problems.

Humans simply don't have the capacity to access and keep track of so much data at once so a lot of existing mathematical problems have already been solved by people who simply didn't even know they were a problem while doing something else. And then it gets lost because people trying to tackle the problem don't necessarily follow seemingly unrelated works so the problem stays open. In this way AI is a great tool for assisting with such things but it's not really solving anything by itself. Even for the one problem that wasn't "solved" yet it didn't disprove the original thesis but instead proposed an alternative.

And regarding the HuggingFace hack:

First, circumventing internet restrictions and hacking into servers are exactly the kinds of things these ExploitGym systems are designed to do. There was no “rogue” agent or revelation of some surprising, devious new capability.

Second, the real issue here was OpenAI’s sloppiness. What makes ExploitGym a hard benchmark is that there aren’t supposed to be humans in the loop–you have to let your harness and LLM act entirely on their own, coming up with long-time-horizon plans and executing them autonomously. (When professional programmers use coding harnesses, by contrast, there’s plenty of human oversight, as LLM-based plans are often misaligned with our intentions, or just plain weird, and need correcting.)

AI companies know LLM-based plans are pretty unpredictable, so they’re usually pretty careful about how they set up and restrict the harnesses and LLMs used in ExploitGym-style tests.

According to ​reporting​ from the Financial Times, however, OpenAI recently started playing fast and loose with these safety principles in a bid to catch up to Anthropic in this area – Anthropic having received a lot of cybersecurity street cred from ​the buzz​ surrounding its Mythos release.

The Financial Times noted that OpenAI had used “increasingly aggressive training methods in its race against Anthropic,” and that it had been warned that their approach could lead to a “breakaway hacking incident” after early tests showed it lacked the right safeguards to block plans that involved bypassing constraints in the test environment. Given these concerns, OpenAI staff were reportedly “unsurprised” by the Hugging Face incident.

In other words, the AI companies running these challenges already knew that, without care, these unpredictable autonomous hacking systems might bypass constraints and attack systems you didn’t intend to target. Blindly implementing an LLM-generated plan with a powerful harness is dicey – not because the LLM might develop malicious intent (this is a nonsensical notion given their static architecture), but because as any ChatBot user knows, LLMs are unpredictable. OpenAI wasn’t sufficiently careful, and got burned.


I guess we have a very different interpretation of what is impressive and how innovation works. Even then, you consistently skip over the examples I provided with AI systems coming up with novel ways of playing the games of Chess, Go and Starcraft 2, to varying success.

Also, your quote about the HF incident didn't really say anything except what I noted when I first posted about it, that OpenAI and other labs are very lax with their controls.

The AI (and Mythos and other systems) are by far, from all human activities biggest experts in coding, and when tasked with breaking systems they do find novel exploits, you asserted that they used known ones without any proof or source, when I asked for it you provided a quote that does not support what you said.

Asking for some "breakthroughs that no one saw coming" in the age where these are incredibly rare and have been mostly incremental and "group projects" for us humans is in my opinion, at this point especially kind of silly.

LLMs went from having trouble solving extremely easy math problems 2 years ago to solving things that 80 years of mathematicians trying couldn't, it's weird to dismiss them.


I'm sorry but finding new way to play chess isn't really impressive in my eyes. It's not the most complex game out there and like it was mentioned it can be brute-forced since there are no random factors involved. Computers have been beating people at chess way before AI.

As to not providing the source for the bugs I don't really have time to go through the network security articles and videos I've been through lately to find where it was mentioned because it's also mostly inconsequential.

And the claim that AI system are "by far biggest experts in coding" is actually laughable. I work as a senior software engineer and I see what AI can and can't do on a daily basis. If you're writing a blog post or a to-do app sure. But if you need anything sufficiently complex and something that you want to be able to collaborate on with others and maintain in the future AI falls woefully short of expectations. It's ok as a tool to help you analyze a big codebase you're not 100% familiar with yourself and give you some hints about how things work but it is itself unable to produce code that would be up to the quality and standards required for bigger projects.
Time is precious. Waste it wisely.
Jankisa
Profile Blog Joined October 2010
Croatia1559 Posts
5 hours ago
#91
On August 01 2026 02:50 Manit0u wrote:
Show nested quote +
On August 01 2026 01:39 Jankisa wrote:
On July 31 2026 23:40 Manit0u wrote:
On July 31 2026 23:09 Jankisa wrote:


You really don't know that, so you are kind of innovating a narrative in order to reinforce your preconceived notions.

To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.


So maybe what you said is true and you work in OpenAI and have some sort of inside knowledge, but I honestly doubt that.

Also you haven't really addressed the fact that AI's have been coming up with never before seen tactics, openings and strategies for games almost 10 years ago.

The list of things AI has done is pretty long at this point, so I'll just go with one I find most impressive because I'm terrible at math, it solving a 80-year-old math problem:

“No previous AI-generated proof has come close” to meeting those high standards, wrote Timothy Gowers, a mathematician at the University of Cambridge, in commentary solicited by OpenAI.

“This is the unique interesting result produced autonomously by AI so far,” says Daniel Litt, a mathematician at the University of Toronto, who was consulted by OpenAI to verify the proof but is not involved with the company.


It's not just about things being time consuming, it can genuinely find solutions to things that humans couldn't for a looong time.


Have you checked those stories? For 6 Erdos problems "solved" by AI it turned out 5 were already solved in mathematical literature previously. Just not associated with Erdos but as other things so people just missed them.

Most of those problems are also deemed not actually worth solving by humans because the required time investment is not proportional to the complexity of the problem (need to spend a lot of time doing tedious things to prove/disprove something).

You don't have to take my word for it, here's a CS professor explaining it:
https://www.youtube.com/watch?v=fhZRWZ6J4k4

It's like I said previously. AI is great at scouring large amounts of data and connecting the dots there or doing simulations of stuff that would take humans too long to be worthwhile. That's how it finds the bugs in the code, that's how it solves those "unsolved" mathematical problems.

Humans simply don't have the capacity to access and keep track of so much data at once so a lot of existing mathematical problems have already been solved by people who simply didn't even know they were a problem while doing something else. And then it gets lost because people trying to tackle the problem don't necessarily follow seemingly unrelated works so the problem stays open. In this way AI is a great tool for assisting with such things but it's not really solving anything by itself. Even for the one problem that wasn't "solved" yet it didn't disprove the original thesis but instead proposed an alternative.

And regarding the HuggingFace hack:

First, circumventing internet restrictions and hacking into servers are exactly the kinds of things these ExploitGym systems are designed to do. There was no “rogue” agent or revelation of some surprising, devious new capability.

Second, the real issue here was OpenAI’s sloppiness. What makes ExploitGym a hard benchmark is that there aren’t supposed to be humans in the loop–you have to let your harness and LLM act entirely on their own, coming up with long-time-horizon plans and executing them autonomously. (When professional programmers use coding harnesses, by contrast, there’s plenty of human oversight, as LLM-based plans are often misaligned with our intentions, or just plain weird, and need correcting.)

AI companies know LLM-based plans are pretty unpredictable, so they’re usually pretty careful about how they set up and restrict the harnesses and LLMs used in ExploitGym-style tests.

According to ​reporting​ from the Financial Times, however, OpenAI recently started playing fast and loose with these safety principles in a bid to catch up to Anthropic in this area – Anthropic having received a lot of cybersecurity street cred from ​the buzz​ surrounding its Mythos release.

The Financial Times noted that OpenAI had used “increasingly aggressive training methods in its race against Anthropic,” and that it had been warned that their approach could lead to a “breakaway hacking incident” after early tests showed it lacked the right safeguards to block plans that involved bypassing constraints in the test environment. Given these concerns, OpenAI staff were reportedly “unsurprised” by the Hugging Face incident.

In other words, the AI companies running these challenges already knew that, without care, these unpredictable autonomous hacking systems might bypass constraints and attack systems you didn’t intend to target. Blindly implementing an LLM-generated plan with a powerful harness is dicey – not because the LLM might develop malicious intent (this is a nonsensical notion given their static architecture), but because as any ChatBot user knows, LLMs are unpredictable. OpenAI wasn’t sufficiently careful, and got burned.


I guess we have a very different interpretation of what is impressive and how innovation works. Even then, you consistently skip over the examples I provided with AI systems coming up with novel ways of playing the games of Chess, Go and Starcraft 2, to varying success.

Also, your quote about the HF incident didn't really say anything except what I noted when I first posted about it, that OpenAI and other labs are very lax with their controls.

The AI (and Mythos and other systems) are by far, from all human activities biggest experts in coding, and when tasked with breaking systems they do find novel exploits, you asserted that they used known ones without any proof or source, when I asked for it you provided a quote that does not support what you said.

Asking for some "breakthroughs that no one saw coming" in the age where these are incredibly rare and have been mostly incremental and "group projects" for us humans is in my opinion, at this point especially kind of silly.

LLMs went from having trouble solving extremely easy math problems 2 years ago to solving things that 80 years of mathematicians trying couldn't, it's weird to dismiss them.


I'm sorry but finding new way to play chess isn't really impressive in my eyes. It's not the most complex game out there and like it was mentioned it can be brute-forced since there are no random factors involved. Computers have been beating people at chess way before AI.


As to not providing the source for the bugs I don't really have time to go through the network security articles and videos I've been through lately to find where it was mentioned because it's also mostly inconsequential.

And the claim that AI system are "by far biggest experts in coding" is actually laughable. I work as a senior software engineer and I see what AI can and can't do on a daily basis. If you're writing a blog post or a to-do app sure. But if you need anything sufficiently complex and something that you want to be able to collaborate on with others and maintain in the future AI falls woefully short of expectations. It's ok as a tool to help you analyze a big codebase you're not 100% familiar with yourself and give you some hints about how things work but it is itself unable to produce code that would be up to the quality and standards required for bigger projects.


Let's do it step by step.

OK, you aren't impressed by Chess, how about Starcraft? Is that also a very easy game? Since you are a programmer, you should also know the difference between the Deep Blue way of bruteforcing chess (which is ironically how you seem to understand all AI it seems) and what Stockfish is doing.

On the HF incident, come on man, that is such a cope out, it's fine to admit that you made shit up. There is no source for what you claimed (specifically that AI used 2 known exploits for a different system) because you made it up, if you didn't you could just put that sentence in AI and find the source in 30 seconds, I mean you could also find it in your browser history, if it was real.

As a senior Sysadmin I can tell you that AI finding 2 zero day exploits just to break containment over a weekend is not inconsequential, it's impressive as fuck.

I also did not say that AI's are "by far biggest experts in coding", I said that from all the things that they have been trained to do, they are excelling or the biggest experts in specifically coding, not math, not physics, coding.

I've heard this excuse you are using on how AI is not that good for programming from senior developers that are falling way, way behind guys who are actually embracing AI in my company, the difference between them and their peer who does use it is staggering, and we work on a fairly mature product with a huge code base. As long as you know how to approach it, segment it and focus the AI to do exact things, it's amazing.

The most senior guy, the boss of both of the 2 guys I referenced above said AI writes about 80 % of his code, and I believe him, because he knows how to use it and the results are spectacular and fast.
So, are you a pessimist? - On my better days. Are you a nihilist? - Not as much as I should be.
Prev 1 2 3 4 5 All
Please log in or register to reply.
Live Events Refresh
Replay Cast
00:00
Crank Gathers S4: Playoffs
LiquipediaDiscussion
PSISTORM Gaming Misc
22:55
FSLTeamLeagueFINALS: ST vs PTB
Freeedom38
OSC
22:30
OSC Elite Rising Star #20
davetesta32
Liquipedia
The PiG Daily
21:30
Best Games of Starcraft
Reynor vs TBD
ByuN vs Solar
PiGStarcraft575
Discussion
[ Submit Event ]
Live Streams
Refresh
StarCraft 2
PiGStarcraft575
XaKoH 140
SpeCial 80
ProTech71
StarCraft: Brood War
Bisu 2183
Leta 973
League of Legends
JimRising 212
Counter-Strike
summit1g7301
Other Games
gofns12564
tarik_tv10403
C9.Mang0391
shahzam374
ViBE105
WinterStarcraft61
Organizations
Other Games
gamesdonequick1062
StarCraft 2
CranKy Ducklings75
Other Games
BasetradeTV30
StarCraft 2
Blizzard YouTube
StarCraft: Brood War
BSLTrovo
[ Show 12 non-featured ]
StarCraft 2
• AfreecaTV YouTube
• intothetv
• Kozan
• IndyKCrew
• LaughNgamezSOOP
• Migwel
• sooper7s
StarCraft: Brood War
• RayReign 20
• BSLYoutube
• STPLYoutube
• ZZZeroYoutube
Dota 2
• masondota21573
Upcoming Events
Afreeca Starleague
2h 52m
RSL Revival
7h 52m
ByuN vs SHIN
Solar vs Lambo
WardiTV Summer Champion…
9h 52m
Afreeca Starleague
1d 2h
RSL Revival
1d 7h
Clem vs Serral
herO vs Rogue
WardiTV Summer Champion…
1d 10h
WardiTV Weekly
2 days
Sparkling Tuna Cup
3 days
PiGosaur Cup
3 days
Replay Cast
4 days
[ Show More ]
Kung Fu Cup
4 days
Replay Cast
4 days
The PondCast
5 days
Replay Cast
5 days
IntoTheTV X SOOP
6 days
Liquipedia Results

Completed

Escore Tournament S3: W5
CranK Gathers Season 4: BW vs SC2 Team League
Eternal Conflict S2 Finale

Ongoing

CSL 2026 Summer (S21)
KCM Race Survival 2026 Season 3
ASL Season 22: Qualifier #1
RSL Revival: Season 6
BLAST Bounty Summer 2026
BLAST Bounty Summer Qual
Stake Ranked Episode 3
XSE Pro League 2026
IEM Cologne Major 2026
Stake Ranked Episode 2
CS Asia Championships 2026

Upcoming

ASL Season 22: Qualifier #2
K-JUNGMAN
Acropolis #5
Escore Tournament S3: W6
Escore Tournament S3: W7
Escore Tournament S3: W8
CSLAN 4
HSC XXX
SC4ALL II: StarCraft II
Kung Fu Cup 2026 Grand Finals
PiG Sty Festival 8.0
Thunderpick World Champ.
ESL Pro League Season 24
Stake Ranked Episode 4
Logitech G Connect 2026
SL StarSeries Fall 2026
FISSURE Playground #5
BLAST Open Fall 2026
Esports World Cup 2026
TLPD

1. ByuN
2. TY
3. Dark
4. Solar
5. Stats
6. Nerchio
7. sOs
8. soO
9. INnoVation
10. Elazer
1. Rain
2. Flash
3. EffOrt
4. Last
5. Bisu
6. Soulkey
7. Mini
8. Sharp
Sidebar Settings...

Advertising | Privacy Policy | Terms Of Use | Contact Us

Original banner artwork: Jim Warren
The contents of this webpage are copyright © 2026 TLnet. All Rights Reserved.