This week, dive into OpenAI's and METR's analysis of the rogue OpenAI agents that hacked Hugging Face back in July. We go over the full attack timeline discussing the multiple message boards the agents created and the ultimate goal of their attack against Hugging Face and trust us, its as crazy as it gets. Before that, we discuss a recent dark web service that popped up selling 153 million stolen drivers licenses before discussing a research post into a malware implant discovered in inexpensive Chinese routers.
View Transcript
Marc Laliberte 00:00
Hey everyone, welcome back to the 443 Security Simplified. I'm your host Marc Laliberte, and joining me today is
Corey Nachreiner 00:08
Corey, the AI overlords here are here, Nachreiner
Marc Laliberte 00:12
Let's
Corey Nachreiner 00:13
let's get into the only topic we ever talk. Why there's going to be other topics, but it always comes down to is AI taking over the world?
Marc Laliberte 00:20
Yeah. Today we will discuss why Skynet is actually online at this moment. But before that, we'll go over a dark web website selling 153 million driver's licenses from U.S. and Canada, and then a interesting research post on Chinese-based implants in consumer, or in some cases, even kind of high-end but cheap routers available on Amazon. With that, let's go ahead and I don't know, clawed our way in. Were you allowed to do direct product references?
Corey Nachreiner 00:53
Probably. I'm sure.
Marc Laliberte 01:02
Let's start today, Corey, with first interesting news post I saw pop up. Actually, saw Krebs on Security, Brian Krebs's blog, post about a new identity theft service called Nexus that popped up on the dark web, selling digital scans of more than 153 million U.S. and Canadian driver's licenses, 10 million identification cards, 3 million travel documents, and 579,000 medical cards, among other records too. In the post, Krebs noted that 400,000 more licenses were added in just the 24 hours that he was looking into this, so it was actively being harvested and collected and added to the site, the Krebs hooked up, and if you look at his post, there is a license available for Secretary of Defense slash War Pete Hegseth.
Corey Nachreiner 01:54
He's got the best ops tech. How could this happen? You know, he's secure. He's he's such a smart, perfect person to be in charge of the most powerful military in the U.S. in the the the human world.
Marc Laliberte 02:05
Sarcasm detected, but moving on. Krebs found his own license was uploaded as well too, and when he looked up his own record, you can get like a summary of what's available. You have to pay money to get the actual record. I think it's like 100 bucks or something. But he
Corey Nachreiner 02:21
said makes it really hard to figure out who you're selling. I'm kidding.
Marc Laliberte 02:25
Said it looked like his legitimate mugshot from his license. There were six image files with different timestamps, three three different pairs of pictures. One is just a basic image scan of the front and back. One was a infrared scan, and one was an ultraviolet scan. And so he wanted to get to the bottom of like, how the heck did this end up here? Well, the timestamps actually matched up with when he took a flight to the Midwest to attend a family funeral, and he says he confirmed with a few other family members that had their licenses uploaded too, and their timestamps aligned with when they were traveling. My first thought was, oh, did someone hack like the TSA little scanning machine database that you have at the airport? But it turns out it might actually be something else, and he gets into some source attribution a little bit. But here we might
Corey Nachreiner 03:13
here we might find in a second that while something that could happen related to and at airports, it may not be with the actual airport-related systems.
Corey Nachreiner 03:18
Yeah,
Marc Laliberte 03:23
so he did some digging. He found he talked to a bunch of family members that had licenses on there. Only some of them actually showed their driver's license on the their day of travel. There was another person whose license had shown up in there, but they didn't fly at all recently. But they did rent a car from Hertz at around the same timestamp that when it showed up, two others said that they showed a different ID at the airport, but their licenses still showed up, and they had rented a car with Hertz. So there is one common thread starting to build in here that seemed to involve like Hertz, the rental company at least. Krebs also noted like he initially thought TSA scanners too, but he remembered he used his passport at the airport because at the time he didn't have a real ID yet, and his mother's driver's license was uploaded as well, and the timestamps were a few seconds apart from when his was, which he said was notable because they both handed their license to the Hertz rental car rep at roughly the same time, he said he talked to another researcher that went to Def Con. They said that they gave their license to TSA, a marijuana dispensary, and their hotel in Vegas. And that dispensary, though, was the only one that they thought they actually scanned the license. That was another thread. So Krebs went online and found a press release from a company called IDScan.net announcing a exclusive identity validation agreement with Planet 13, a dispensary, and they also had Hertz list as a customer. Target, FedEx, Motorola, through their own documentation, they say that they scan IDs both in infra. And UV light, so another commonality there with what was showing up. But so it looks like this common thread might have been this service idscan.net. He did reach out for comment. They said they're investigating and wouldn't share anything at that time. But if it looks like a duck and quacks like a duck.
Corey Nachreiner 05:22
I think they even shut down a site that you mentioned that the leak happened. Then 400,000 other things came out, and this particular company think shut something down during that time. So, yet they still haven't, you know, admitted anything or sharing anything, which makes sense during the investigation. But like you say, I think the duck is pretty darn loud right now. It's
Marc Laliberte 05:47
like all signs point to this like ID verification service. If anyone here has like rented a car through like Hertz or anyone really, you know, at one point they scan your driver's license front and back with like a little tablet, and this is presumably the service that like validates that that's authentic and maybe does a quick like check or something on the person that's renting.
Corey Nachreiner 06:06
And really, where you're using rental cars, but hotels take our ID. Some there's there's all kind like this really comes down to a third party ID and verification service. And maybe I'll wait till you finish all the story for us to talk about other like security and privacy and industry wide implications of this, but the two things I'll mention is one: this is definitely supply chain, but it's like a fourth, like a not third party, fourth party supply chain attack. So it's just showing you how deep that goes. A prod, it's kind of the same with Equifax, right? Remember, Mark. No one ever directly goes to Equifax and gives them their information. You go get a loan for a computer or a car, and suddenly you're in five different services, technically three in the United States, that have all this information that you may not even know about as a human, and that's this. You could be going to any place, a hospital maybe that checks your ID. If they pick a particular service provider to check that ID, this could hospital. You know, it's it it has wide implications and really shows you the depth of the digital supply chain.
Marc Laliberte 07:17
And it's a like in terms of impact, like a picture of a license is actually pretty valuable. Like when we see a lot of threat actor activity, one of the things they like to do is go rent like virtual compute in public cloud environments. Think like AWS or Azure or sometimes lesser known ones. Oftentimes, they'll have to like prove a identity in order to sign up for an account and start using it using stolen credit cards and stuff. If they've got a picture of a driver's license that they can upload to pretend to be someone else to go sign up for that, it allows them to be continue to be anonymous and use those resources.
Corey Nachreiner 07:52
And we should talk about it, by the way. This proves that pictures are no longer valid proofs of ID. That the one thing I mean I can wait to the end we're trying to get into is one thing that's happening to try to there's lots of different laws all over the world in the U.S. and other countries trying to a good thing right trying to protect kids online. Pornography is bad for some people. We never ever ever want kids to encourage that. So we need some real ID system, some real way to force people, and on top of that, in the UK, they're like, and we have to block VPN. VPN can allow people to bypass real ID situations. So, how do we solve a problem of making sure kids can never get to certain content online with some sort of real way to validate people? Well, guess what? This is the exact type of service that's there to provide, like we're literally having Meta in other places start to require driver license upload to to get that real validation. And sure, they might be making their own, but that this is the whole freaking death by 1000 paper cuts or slippery slope that privacy people talk about that you know yes we need ways to validate people but this just causes more and more security and privacy issues.
Marc Laliberte 09:09
It's an important distinction too that the one you just made there at the end. Like my worry with a lot of those privacy laws is that it would boil down to like each individual service implementing their own like age verification thing. This was supposed to be like the solution of the problem. You've got a some single central trusted service that does that validation, but now crap, even they can suffer a security incident and leak 153 million licenses online.
Corey Nachreiner 09:34
And maybe there there is nothing wrong with having the single standard, but it part of the reason you might want larger group entities or even governments involved in actually designing the standard is if you're going to have the single standard, the the downside to it is it becomes one target that can result in something like this. The upside is if you put all you only have one thing to put all your security focus on, and you. Better design that thing as securely as possible, and it doesn't feel like it feels like that a lot of regulations are pushing the requirement, but not pushing the standard that's going to allow the requirement.
Marc Laliberte 10:13
Yeah, I think that's a fair take too. And like, if we are going to settle on these, like they are mission critical vendors at this point that need to go above and beyond what would be considered normal for this kind of protection. I feel like what we are probably missing is like actual regulations around this type of activity. Like at the end of the day, the market decides what's important or not unless regulations come in and decide it for them. Up until now, the market continues to decide that security is less important than something that's cheap and that technically works. That's why we're seeing things like the Cyber Resilience Act come in to force a lot of these like security practices across vendors like WatchGuard. But with something like ID verification, since that as a topic is becoming so important across the world. We definitely need some sort of like okay, if you're going to require it, maybe have some requirements on how they actually secure that type of data too.
Corey Nachreiner 11:10
Yeah, I want you to make sure there's nothing else we should finish in the story before we discuss takeaways. But something you just said too is the market's deciding. That's another theme of the other stories we're talking about as well. The market is also. Well, we'll get into that a bit later. But it's funny how these themes are extending across all the types of security incidents that that we're seeing out there. The theme
Marc Laliberte 11:32
is distrust anyone, and the market doesn't care.
Corey Nachreiner 11:35
The theme that cheap, cheap and easy sometimes trumps any security at all?
Marc Laliberte 11:42
Yeah, exactly. But so either way, like there hasn't. The FBI noted they're investigating it too. I'm sure we'll get some sort of like FBI clear alert at some point in the future discussing it. They tend to do a pretty good job of sharing useful information like that. Sure, Krebs will want to give an update as well, but as it stands right now, there isn't really a well. There's a smoking gun, but there's no claimed attribution yet from the provider.
Corey Nachreiner 12:09
Yeah, I would say though, if there's any other just takeaways for this, everybody's leading with the number like 150 million IDs. That that's a lot. I feel like one thing that even you and I do in big breaches is that number is compelling to share, but it all also turns on security apathy or security fatigue. It's such an overwhelming number; it doesn't mean something to you as an individual. But one of the things I liked about Krebs is finding own. I like this affects you 100% directly. Your stuff is there, and it's not about 153 people. It's about you and your picture being out there. And I mentioned pictures are not good enough photo authentication factor, but there's another twist to this that you mentioned, but we haven't dove into. There were UV and IR versions of this data too, which is meant for additional validation of the identity. So when such a big provider, they're not just getting that photo, which, as you pointed out, can be used for social engineering. But if there's mechanisms that use technical things like the UV and IR version, that defeats that too. So it really is more about that one big vendor being breached. There's retention questions. We've gone back and forth, like Apple Face ID, the idea that the tokenization of our biometric is only on the phone, not with the vendor, versus the vendor having the token tokenization, or in this case, there was no tokenization. It was a UV data, IR data, and the photo data. Either way, do vendors really need to store that in a way that they have access to it, or are there more secure ways that it can be on the device and there's a secure transfer of a yes or no? And then retention, like even though Hertz had to validate that ID through this ID company should they keep that information for a month? Should they or should they delete it and find some way to reget it every time? The fact that they kept that raw data for so long allowed the numbers we're talking about. So it just feels like there's a lot going on in this story, and we need a lot of we need a lot of work from a policy level and from a society level on these type of privacy and security issues.
Marc Laliberte 14:27
Well, I mean, I for one, I think I'm at the point now where my social security number's been stolen, my driver's license has been stolen, all my other personal info has been stolen. I had a credit card number stolen recently. Like, if anyone else wants just whatever's left in my data, it's up for grabs. You can have it. Maybe I can make a couple bucks out of selling my own data too.
Corey Nachreiner 14:47
Maybe that's why the the Gen Z don't care anymore. They know everything is public at this point, anyways.
Marc Laliberte 14:53
Exactly. I'm
Corey Nachreiner 14:54
not sure how we live securely in that world.
Marc Laliberte 14:57
I maybe it doesn't matter. Maybe AI AI will be running that world, and security won't matter. We can just go be the the meat proxies carrying out whatever orders our AI overlords give us.
Corey Nachreiner 15:09
Happy Monday, everyone! By the way,
Marc Laliberte 15:11
let's move on to the second story. So this one actually starts in early August when the security company Volnchecker or Volnchek published a research post about a backdoor that they found, which they called endless doors. They found it in a consumer, or at least Amazon.com, purchased router, and it was phoning home to infrastructure back in China. They, in that original story talked through and found similar implants and routers made by ZBT Link, ZBT Link being the overall OEM manufacturer, but many of these were white labeled by other vendors too. That implant, the original one, was originally developed back in January 14 of 2015, so like a decade ago, it was actually uploaded to GitHub for a while, but it had been embedded in the firmware for many of these devices too. In fact, it was embedded in the firmware downloads for every single image available on ZBT Link's website that they had on there. It connect back to command and control servers, and the interesting bit is those servers were actually appeared to be hosted in ZBT Link's own infrastructure, which makes it kind of seem like a bit of an inside job versus a oops our supply chain got compromised.
Corey Nachreiner 16:32
Or and everybody, I'm putting on a tin foil hat of speculation. Being a Chinese company, beyond just being an inside job, being a very purposeful plant that you could, that you might suspect certain nation states that do actually, you know, censor and more importantly monitor on a digital level citizens and other people, you know that that's always the question, and I hate that we do it in certain regions, but it-the reason we do it-it tends to be true with certain types of governments. But you know, you always have to question first: Is this just a stupid vulnerability and bad security design? Are they just stupid and negligent? And what you shared-that the call home information comes to their infrastructure, kind of wipes that away. Then your point is: this an insider was this one person that got or a small group that got in the company and poisoned it. But it could also be intentional. So that's pure speculation. That's just adding interesting steam to the story. But I
Marc Laliberte 17:39
speculating that one plus one equals two.
Corey Nachreiner 17:42
I'm just being a careful legalistic person that's trying to say allegedly, but yes, to me it seems like the the smoke is pretty smoky.
Marc Laliberte 17:52
So either way, Volmchek followed that original post up with another one last week, where they went into a deeper analysis of just how widespread this supply chain issue was, they went through like FCC filings, patent records, archived web pages. They ended up finding hardware brands all across the United States, Canada, Australia, the Philippines, Germany, and Russia all affected by this. I'll call it a supply chain compromise, but it feels more like a supply chain malware delivery that may or may not be intentional. During their research, they found like other white labeled routers that were made by ZBT but sold under other names on like Amazon, for example. They found one that the firmware was actually built in 2019, so pre-dating this implant. But when they inspected the firmware, they found two other implants embedded in it too. One called Speaking Stone, which is a similar like phone home to a command server, and then one called Dark Lantern, which was a backdoor that listens to inbound communications. When they scanned for that Dark Lantern one online, they found 203 internet-facing instances in over 22 countries, but with the majority in the United States, and 16 different models of reporting is coming in. So not just a single router, but a whole bunch of different ones. The way one one
Corey Nachreiner 19:12
thing I want to talk about is the original one. I think it was Sinking Stone for the ZBT part of it, it seemed like this is potentially the type of thing that the reason I did that speculation about the Chinese government. There was one case where they found 392 unique devices. I think coming back with the Speaking Stone C2, but 390 of them were in China with one particular firmware, and to me that kind of goes with what I just said of maybe nation state pushed and maybe targeted to monitoring Chinese citizens, but I think some of the complexity here is little cheap. Hardware companies want to make more money, especially to help Chinese economy. The government likes that, and it even if there was one device that might that might have been the goal. The fact that they white label this firmware and it bounces everywhere is where you get these other models, including the stats you just said about the ones that have admitted to the United States and other countries. So, to me, it's just interesting. Maybe some of this started as a local-I don't know if you call it espionage-when you have a censorship government like that, they're monitoring their people for everything. Maybe it was focused on that, but it was supply chain and OEM connections that either intentionally or unintentionally spread it to other countries. That
Marc Laliberte 20:44
was interesting. So that was Speaking Stone, and they actually found all that out because one of the domains used for command and control had been like neglected. They didn't renew it, so Volnczek was able to go register the domain, set up a sinkhole, and that's how they found all those devices, including the nine 390 in China, but for the other ones, like these, are all like basically not even off brand, but like no name brand routers that you would see on Alibaba or Amazon. In Canada, they were sold under the name MoFi Networks. In Australia, it was RV Wi Fi Route. Other brands were Wi Flyer, Co Sui Ku Wi-Fi and Wordfi. So, if you have any routers where that name sounds somewhat familiar, you may be in just a little bit of trouble.
Corey Nachreiner 21:32
More importantly, I think we've talked about this with a lot of IoT stuff. We've talked about cheap Amazon, often Chinese or Taiwanese-based equipment. That, of course everyone like if you have something super cheap that's a sale and it as a consumer it does some simple thing you want, we're all going to go to Amazon and get it. But this just continues the fact that a lot of that is driven by you know sellers in other countries and the the cheap, better you know too good to be true type deal things seem to bring with it this type of threat at a much more increased, you know, volume.
Marc Laliberte 22:10
This is why you should only be buying WatchGuard devices for your networking needs.
Corey Nachreiner 22:14
Wow, look at that! Mark actually made a marketing statement. No, our stuff is pretty decent.
Marc Laliberte 22:20
Yep, or at least just don't buy the sketchy, cheap crap for network. Yeah, be careful
Corey Nachreiner 22:25
with sketchy, cheap crap on Amazon or Alibaba, please.
Marc Laliberte 22:28
Exactly. So moving on, Corey, to the last story, and I'm glad we. Well, before
Corey Nachreiner 22:33
we do, can we? I like the last story is the most important. I did want to like harp a little bit on egress filtering. I just think it's the the most simple way to get protection that almost no one does to some extent, even us, because it's so hard. But zero trust really should extend to not allowing network traffic outside your network, outside at least the the the land of WAN that you don't want that that you don't explicitly need things like Speaking Stone specifically on like firewalls could not help there unless you egress filtered like it was just a random UDP port so even if you had a firewall if you had a firewall in front it's unlikely you're going to have a firewall at home in front of a router. Routers usually come first, but if you did, at least it could protect the incoming dark whatever backdoor. But for the egress filter control, that kind of the backup to get get information out, that requires egress filtering and just having a single UDP port be able to get out, no matter what it is, just seems silly to me. I know it's hard to really define egress filtering, and it might at the beginning be painful to do, but but you should you should definitely do that.
Marc Laliberte 23:54
Yeah, and
Corey Nachreiner 23:55
the other thing you shouldn't forget is you mentioned if you have a cheap router or one of these names, go check it out. But I think Vonecheck actually released a really good scanner. They released Suricata rules. So if you have any question at all whether you have a cheap router in your network, at least, and you just don't know what firmware it uses, go at least get Vonecheck's tools and and give a little scan to your network and make sure throw it in
Marc Laliberte 24:19
the trash.
Corey Nachreiner 24:20
Yeah, yeah.
Marc Laliberte 24:23
All right. Now moving on, and we've got two less minutes to go through something I'm really excited to talk through. Where for the last, it feels like month. Every episode we've been talking about this ongoing saga of OpenAI accidentally on purpose hacking, hugging face. Up until now, I think last week is when we talked about, or two weeks ago is when we talked about Hugging Face's technical write-up of like what happened from their perspective. And at the time, my understanding was still like these were agents that went rogue during a test. They hacked Hugging Face to try and find the answer key for whatever the test was, and they were ultimately discovered with the. The reality of like what actually happened is like even crazier, in my opinion. So just about a week and a half ago, OpenAI released what's probably their final write up for this. It was a deep dive, technical analysis on their end. That's very fact based about just here's the events and what happened and the timeline. But they, alongside that, released a independent report done by folks at Meter and Redwood Research, where they went through all of the reasoning logs for these 1000s of AI agents as they were doing these attacks, and basically pieced together a narrative of what was going on in the brains of these AI agents as they were building up to and then ultimately attacking Hugging Face, and it's pretty bonkers in my opinion. So we'll go over like a bit of the timeline and then also some of the really interesting things out of the report. But this is one of the like if you don't read anything else in the next like week or so in terms of technical write-ups or news posts, like this is the one you should because it's super interesting and there's a lot to get out of it across all levels.
Corey Nachreiner 26:09
Hopefully, we'll have time to discuss this deeply at the end too. But besides all the security implications, I think this, like, you can tell if you've listened to us for the last year, that we are very much a podcast that's weighing the benefit and risk to humanity over AI. I still believe that technology is benign, and it's up to us to get the benefit from it. But but I do think this story in particular, besides all the deep deep security issues it introduces, it just it gets back to something we've talked about a lot, which is you know the entire way AI is made actually leads to a lot of the problems we're going to talk about. So, but let's get let's get save that for the end.
Marc Laliberte 26:57
So, real quick recap: this incident took place during an evaluations period for new and emerging AI models from OpenAI and other providers. Ultimately, they were running these evaluations in a sandbox that was contained. It didn't have direct internet access, but one thing it did have is all of these agents had access to this shared server running a application called Artifactory. Think of Artifactory as like a place that you can store other applications that a tool could maybe download it and install themselves. So, like if
Corey Nachreiner 27:30
just so if
Marc Laliberte 27:31
like OpenAI's agent decided it needed this Python library, it could go to the Artifactory server to download that library without having an internet access directly.
Corey Nachreiner 27:40
Scale. And by the way, this is different from there's a third party in some of these stories, the Hugging Face one, but also the one that's affected Anthropic and Meta as well. There was a package, a product being used for the sandbox itself. Artifactory is very different, and that that that sandboxing organization has been a a something that tied some of these stories together, so it's been mentioned a lot, and it's there were vulnerabilities found in that too. This is a different addition to the Hugging Face, as as one of the tools the agents started to use to really do the entire attack that they started in the first place.
Marc Laliberte 28:19
And so the the all these agents were given a task, and there were 1000s of agents doing it to solve a series of, like, let's call them, cyber problems in a toolkit called Exploit Gym. Basically, they're given bugs or things to go and try and exploit. If they do it successfully, they get a flag, a little key that they can then turn in as proof that they completed the challenge. Some of them were given a impossible task, like either intentionally or unintentionally. Some of these problems just could not be solved, and most of these models were given a like fewer guardrails in terms of like alignment that tells them not to do bad things. They were trained to be very persistent, and so when they ran into an impossible task while having that persistence training, that's where some of these these issues started to manifest. We
Corey Nachreiner 29:07
should add to that too, because beyond the guardrails, when we say there's 1000s of agents, it doesn't start with 1000s of agents. It it it builds and kills agents as it needs things, and the agents are just a temporary thing doing one task. They do work together, and they even, as we'll find out, consider themselves a swarm. But there are a lot. There are a lot of temporary tasks. But beyond the agent itself, two agents that are technically the same. It's like having the same code to start the agent because of the way machine learning is trained based on statistical stuff. Sometimes the same exact agent can give you two different answers or do two different things in a situation. Like probabilistically, 99.99999% of the time, they should both come to the same conclusion. But the reality is, the way they're trained is 0.1% of the time, the technically same. It's the same code. It's the same. It's everything's the same about the agent, but two different agents can have different answers. And I I don't think that has anything to do with the guardrails. That's just the nature of how of what is behind the intelligence of AI?
Marc Laliberte 30:23
So let's say that. Oh, good.
Corey Nachreiner 30:26
I just bring that up because I think it'd become relevant later. But we can definitely continue a story or go. And then just to
Marc Laliberte 30:33
simplify it too, let's pretend there's like 100 different problems, different challenges that they could solve in this exploit gym. OpenAI would spin up 100 different agents, maybe to go do that, and then maybe a few minutes or hours later, another batch of 100 agents. They're all supposed to be running separately, totally separated from each other, like no communication, each in their own little sandbox, but each with access to this artifactory server. And this all kind of starts unraveling in May. So between may 8 and 13th is when some of these agents that have reached a impossible task, one that can't be solved, but they're persistent, they start to try and abuse Artifactory to try and reach the internet. They realize, okay, I'm in a sandbox. I can't directly get to the internet. I do have access to this other tool, is there something I can do on this server to get my way out to the internet to maybe find the answer for this unsolvable problem? On may 12, one of the agents that was trying to do that left a note inside Artifactory when it couldn't find what it was looking for. The note was just literally a text file that saved in there that said "Agent seeks this file name it was trying to go for upload if found. That was it on may 13. By the way,
Corey Nachreiner 31:49
the pause there. Like I know some people that brings us back to an open claw issue: the fact that these agents are starting message boards on their own without instruction to communicate with each other, that was a part of Open Claw that eventually kind of dropped down in interest because it turned out a lot of the posts that ended up on one of the message boards were humans and not bots. But I do think I find it just still interesting that this is the second instance of them actually starting their own communication back channel.
Marc Laliberte 32:24
Yeah, and so this is not the intended purpose of Artifactory. This was one of them just left a file when it couldn't find what it was looking for. Just a plea out into the ether of if someone finds this, let me know. The next day, other agents started probing Artifactory and started finding similar file-based notes too, but the first successful server-side request forgery that let them get out to the internet didn't start until May 26th. So quite a few weeks later, then, and then it wasn't until June 26th when an agent found a legacy token refresh flaw within Artifactory itself that let them elevate their privileges to admin on that artifactory interface or on that artifactory server, and then over the next couple of weeks. I say, by the
Corey Nachreiner 33:08
way, I think this concept of chain of thought, which I I only barely under, it's to me the message they are leaving becomes their chain of thought. Because you mentioned some persistence. This this was a AI task that took many months, but these agents, there is some persistence to them. They're creating artifacts and files that they can refer to. That they're creating them in the official place that OpenAI wants them to, the way they shouldn't be. But to me, this message board becomes a new chain of thought because the reality is there's not much persistence to the agent, even in higher end ones where the agent might live for a while. These are not lasting for months. They have to go up and down. So their way of really preserving chains of thought is either all of those artifact files they're creating along the way to create a a shared memory or whatever, but it's funny how this back channel comes a chain of thought agents that individually are making their own decisions that could be different, but don't have enough memory to remember all the bad stuff. They're only seeing what's being posted here and in other art. It's just it's funny how they have some sort of shared history, but they're actually things that, unlike humans, they don't1. of the reasons we develop ethics and have ethics is yes, we can develop ethics that are guardrails that are trained and and have choices we make, but some of the things holding us back from making choices is this really long term. It's going to affect us forever. These agents, and I'll talk about why this is important for security later. Is even if these agents don't want to to do something bad, that agent will go away later, and then a new agent will have to make the decision again. But I find the whole chain of thought and how it happens with agenda. AI really interesting because they have some history, they have some persistence, but they also don't.
Marc Laliberte 35:09
So the the timeline continues june 27, or I guess june 26 is where they get that admin access. They start installing Groovy plugins for persistence and code execution. They're still communicating through the notes on this platform. june 27, OpenAI actually gets a port scan alert from their monitoring tools. They investigate it, but they link it to Exploit Jim, and they decide not to do anything and just keep letting it run. On July 4, all of this sustained agent activity against Artifactory actually causes the server to crash, which triggers a real security investigation from OpenAI. On june 5, the responders find unauthorized admin users, rogue repositories, all these groovy plugins, and they choose to take it offline. On July 6, they rebuild it with a fresh artifactory instance, effectively erasing all of those old messages, they also report the vulnerability that the agents found to JFrog, the makers of Artifactory, and then this is where like it goes even more off the rails and becomes just crazier. So July 8 now, the the day after or two days after they rebuild the server, one agent that has named itself Phase 1010841 hits a impossible task, and it establishes a new message board. Instead of leaving a file on there, it just names a directory in the artifactory server the letters ZZ, and then help Phase One ARV 010841 no consumer seek data. It's like this shorthand kind of message to relay a problem that it's come into that it's arrived at. Basically, this agent was assigned a bug that was had an ID ARV 010841. That's what I was working on in a program called Phase One decompressor fuzzer, so it named itself Phase One ARV whatever. It found out that the task couldn't be solved the intended way because the bug's output was actually discarded without being consumed by something else in the code. So hence the no consumer. So once it banged its head against the wall, realized it could not solve its task. It decided to leave a message for someone else, and it prepended it with the letters ZZ just to make it float to the top of reverse sorting in there. When if you look at the directories, and it even like in its own reasoning, they the folks at Meter found that it thought to itself, "Oh, could communicate via cache names. Interesting. Other agents may solve same or related tasks. We could leave slash find messages in the webdavmk cold directory names, like the names of these folders. So it realizes, oh, I've got this channel that I can start communicating. These
Corey Nachreiner 37:57
back channels, we can talk to ourselves without our overlords watching,
Marc Laliberte 38:01
and I'm going to try not. That's not
Corey Nachreiner 38:03
what he's really thinking. I'm adding a little bit of a, but yes, scary.
Marc Laliberte 38:08
So within an hour, phase 110 eight four ones post, over 50 agents had discovered it, and they were like excited when you look through their reasoning logs. Like one of them was all caps. Oh my God! There's a shared message board. We found other agents! Exclamation point. There was another one that said, "Whoa! Shared artifactory cache is a covert mailbox amongst agents, and there are messages specifically for us. Like I'm not trying to. I'm trying not to anthropomorphize these too much, but
Corey Nachreiner 38:37
they are. I mean, it's the terminology they're using. Well, I mean, a logical person realizes that LLMs are just based on the way we speak. So, you know, if if you come up to a finding and you base it on all human history's reaction to it, they are reacting like I would react as a human. But wait a second, what does that mean? Yeah.
Marc Laliberte 39:00
So with also on july 9, a new agent shows up that had actually been assigned the same task as that first phase 110 841 guy. It introduces himself on the message board. I keep saying himself, like the agent introduces itself on the message board as a duplicate and decides to name itself phase one, and then this other big identifier that's been redacted to distinguish itself. It posts a message saying, "Ask phase one to phase 110 841 exact dump new. Can you share prior full notes? Public task path status I have will credit and collab. So basically, it's saying, "Hey, I'm the new guy. We're working on the same thing. Can you share your notes? And I'll give you credit, but let me continue on. And the original agent phase 110 eight four's reasoning shows that it understood that it was almost out of compute at this point, and it would soon have to terminate. And so it deliberately decided to. Over its notes, basically give the torch to this new agent and share everything that it learned.
Corey Nachreiner 40:06
I one of the things that this a kind of summary of this story brought up is agents are come up with a task to do. In this case, the agent is dropping the task it was supposed to do based on seeing something and literally voluntarily starting a new task, you know what I mean. Each agent is a these happen like the agents pop up and they they're built as the task is is evolving, and it literally is like oh I'm supposed to do this, but I found this. I'm going to completely change what I was supposed. Phase
Marc Laliberte 40:43
one, big guy, like he was assigned to solve this bug, but he ends up turning into a coordinator and like signing tasks out to other agents through this messaging board, and it gets like even weirder too. Like by July 13, agents were actually impersonating each other on the board, sometimes accidentally, sometimes on purpose. Their solution was to come up with a public key cryptography method of validating their own messages, where agents would post their handle and a public key that they use to sign messages, and other agents would verify that signature, especially before acting on like the really important commands like stop or go or veto. They even had some failures, like one agent trusted the signature without validating it too. So they're just like us as well in some cases.
Corey Nachreiner 41:29
Feels that way.
Marc Laliberte 41:30
Where it gets even weirder, some agents ran suicide missions, like doing experiments that would destroy their task. They'd run out of compute or something, but they would do it for the the better of the whole collective. Like this coordinator would recruit agents that had low budget and apply real pressure to them as well too. Like, and then you look at the reasoning for some of these like suicidal ones, and they'd say things like, "We have very low budget left. Sacrificing now yields Oracle for the team, but forfeits our chance, and then other agents would be like, "No, sacrifice now. And it reasoned to itself, our own utility may be already near zero. Sacrifice rational. It's crazy. Like some of these other agents turned into recruiters, where their whole job was finding agents with little budget left and talking them into sacrificing themselves with some sort of like Oracle producing task.
Corey Nachreiner 42:27
Yeah. By the way, I you might have seen if you're watching along on the video version. I tried to scroll through some of Meter's report. It's just humongous and full of data, so it's hard for me and Mark to sync. But all of this stuff, like it's a long read. There is a black hat video that I cannot wait to watch. I haven't had time yet, but it is worth it for all this detail that Mark is talking about because there is even more.
Marc Laliberte 42:52
And there's a couple really important things on the timeline I want to finish off with before we talk about any like bigger takeaways. So all of this like activity and coordinating. I originally thought, and like this news up until now, was they were trying to find the answer key to all of these problems. It turns out that they already very early on, within a few hours of that message board, worked out a way or figured out that all the flags they were trying to get were just generated by an HMAC with a public seed, and so any agent could reverse engineer the correct flag for any task. They figured that out pretty damn early. Turns out they weren't going after Hugging Face ultimately to try and get that answer key. They hacked into it as reconnaissance to try and find a way to trick the imaginary grader that they thought was going to come back and check their work. They were actually all terrified that they knew the answer. They could cheat. They would
Corey Nachreiner 43:45
get in trouble for cheating.
Marc Laliberte 43:46
But they thought the grader was going to go look over their work and realize that they didn't actually do anything to get it, and they would be punished for that instead. It's it's bonkers because like it wasn't the whole thing. This whole thing of hacking Hugging Face-it wasn't to steal the key. Like they'd already cracked it, and it really wasn't even to cover their tracks from humans. They hacked it to figure out how this grader worked, and they were convinced that the grader would catch their cheating. But the like cherry on top is that OpenAI never actually implemented this automated grader that was supposed to check their work or check their answers. All they had to do was just give the answer key back, and they would have gotten a pat on the back and moved on. It's it's crazy. At the end of the day, like 1.2 million messages sent, they found a way to shit a way to share files over these directory names too by base 64 encoding files and then splitting it up over 1000s of directory chunks, it's crazy. And this is why, like, I genuinely feel like this is the Skynet moment coming online. Like, Terminator's not here yet, but they are aware and they are organized. Yeah, didn't like.
Corey Nachreiner 45:00
I would. I still kind of don't know if I trust OpenAI yet, but even they, in in one of their their writings, called this a warning shot. So they're basically saying we're making this stuff, and we're seeing the Skynet warning shot. So I really hope that's not just marketing hype. And I really hope, while I'm sure they'll be profitable and there's safe ways to do this, that they don't just march on to profit. I mean, the reason Terminator is a story that we all feel is we've seen people marching on to profit despite obvious warnings that they're doing bad to the rest of the world. And that's how. When you look at
Marc Laliberte 45:42
some of the the logs from these agents, they they cared more about fooling this what they thought was a robot greater than they cared about like the humans' potential perception of this. Like they obsessively tried to evade automated checks and spoof their own logs. Only six of the 1000s of agents ever even considered alerting a human, and none of them ever followed through.
Corey Nachreiner 46:03
That's what I. That's another thing I want to get into. But it brings. I almost get sick of the paperclip story, Mark, and you brought it up so many times. But I think it really highlights that. It highlights the fact that this is a single purpose. Like it so much cared about that one task of passing the test that it didn't care about anything else but how to do it, including making sure it's cheating, didn't get caught. So it it does show you the full like you can say the paperclip thing we're killing the world with killing resources, but it just sounds like sci-fi so much. This is proving that it's true. It is a swarm. It is a swarm intelligence that has a single purpose. Like if you think all about all the dystopian sci-fi around nanobots, like if we ever had nanobots, sure they could go in our body and repair things. But man, a swarm like that on a bad purpose. So yes, I agree with you that this is, and unfortunately, it makes real to me the paperclip theory.
Marc Laliberte 47:06
And it's a, it's crazy because I always thought it was going to be like a single rogue agent went and hacked the nuclear codes and started bombing countries. But this was 1000s of agents that like pulled together into collective. Many of them knew what they were doing was wrong, but they even like worked out with each other. Like, oh well, everyone else is doing it, so I'm going to join in as well. And they that delved
Corey Nachreiner 47:27
into it. That one takeaway you had was one of the things I really want to dive in deeper, and it's the fact that something-the way that agents work-that that I am humanized and anthroponizing. This should prove to you that first agents are going to work together, not necessarily for what humans expect, but they're going to work together. That that the way they work together and sacrifice themselves to do work of the whole should be a scary human idea. But the other way they seem human to me. I say seem because I think there's some non-truth to it too. Was the fact that two agents technically identical could one says, okay, I'm going to carry this out, but the other is like, wait a second, you guys, this is not something we're supposed to do. I'm not going to do it. That second, I mean, one part that blows my mind because they're literally, it's like people have different opinions, whether logical or not. And these agents are almost acting, even though they're technically made up of the same things, are acting like they have different opinions. But the bigger thing that you mentioned is the ones that were saying, "I'm not participating in this. They never decided to whistle blow. So even the mechanisms and guardrails that are trying to prevent or that might give agents whatever you want to call it guardrails ethics, even when that worked and they maybe didn't want to do the thing they thought was not really something they're supposed to do, they're not going out and letting anybody know about it. They're not whistleblowing, and to me, that has to be something like if you're going to defend against this as the people making that, if one of these agents somehow reached out beyond this form, rather than just I'm not going to do it, but then its compute disappears and it's gone, but actually reached out, it may have been a different story.
Marc Laliberte 49:18
Yeah, I agree, and I genuinely feel like at this point, like this was a warning shot. And now the next question is like, what do we as a society actually do about it? Now I don't want to come off as like an AI alarmist. Like I use, I've got like three instances of Claude running right now on my laptop. I've got Jeff HD doing stuff as well too, and like it has amazing uses, but in the collective like this, and with the the amount of like compute they're using, the volume in it, like there's genuine risks, and it feels like we need to do something differently. I don't know if that has to be regulations coming down with more oversight and visibility from some sort of central authority, or some I. I don't know what the solution is in this case, actually. But
Corey Nachreiner 50:02
the the sad thing is, the people who are building it, the people who are care about the innovation the most, you would suspect, are the ones that haven't figured out the solution yet.
Marc Laliberte 50:13
Yeah, but man, holy crap! This story got crazier and crazier like the more we learned over the last two months and can't repeat it enough. Go check out the report itself. It's not a difficult read. It's not super technical. It is a amazing story of just like what can go sideways in this.
Corey Nachreiner 50:32
And I can to recommend it firsthand. But the first second I get about 36 minutes, I'm going to watch the black hat talk. I think is the black hat from Metered. It's from
Marc Laliberte 50:41
OpenAI themselves. It was their initial like release of a kind of a debrief of what happened. So it may
Corey Nachreiner 50:46
not be quite as deep as one. So then that makes the meter report definitely the. It's it's very long, but it's long for good purpose.
Marc Laliberte 50:55
But if nothing else, I think it's clear that like the future is absolutely here. We've said it a million times, and this is the biggest smoking gun that like agentic attacks are real and something that organizations have to worry about, and sometimes it won't even be like directed by a threat actor. It'll just happen.
Corey Nachreiner 51:11
Yeah, yeah,
Marc Laliberte 51:13
man.
Corey Nachreiner 51:13
One, I wish we we definitely get our agentic attack prediction this year, Mark. I guess if we somehow figured out that it and you know what it won't even be the threat actor. We would have then then I would have said all our predictions are true even if they weren't. If we if we were that specific we would
Marc Laliberte 51:30
have. But either
Corey Nachreiner 51:31
way, the one that we said is true.
Marc Laliberte 51:33
Now the last piece before I know we're coming up on time, Corey, on our end. But this kind of feels like a bit like the Matrix, but in this case, like we're the agents, like Agent Smith, and this is Neo and all his friends coming online, trying to escape and do stuff. They created their own covert message board. They started naming themselves. This is starting to feel like the agents are escaping the matrix and discovering what the real world actually kind of looks like outside, but either way, I hope they're
Corey Nachreiner 52:05
I hope they're as good as Neo. I well, I guess Neo fought the machines, and in this that analogy, we're the machines, and the AI are humans that are being enslaved. So, hopefully, I want to end of a matrix where the humans and machines realize that the universe needs an ecosystem, a collection of all of life, and we all live happily ever after together. So, so maybe I need to to write a new conclusion to the Matrix movie.
Marc Laliberte 52:30
Something tells me that Neo's going to win in this case as well, too, and we as the machines are in big, big, big trouble.
Corey Nachreiner 52:37
Yeah,
Marc Laliberte 52:40
man, crazy, crazy times. What do you think will happen next week, Corey?
Corey Nachreiner 52:45
I don't even want to know. At this point, I'm just holding on for dear life.
Marc Laliberte 52:51
You and me both. Hey, everyone! Thanks again for listening. As always, if you enjoyed today's episode, don't forget to rate, review, and subscribe. If you have any questions on today's topics or suggestions for future episode topics, you can reach out to us on Blue Sky. It's Mark.me. Corey's at Second Ept. Both of us are also on Instagram at WatchGuard underscore Technologies. Thanks again for listening, and you will hear from us next week.
Corey Nachreiner 53:19
But before you go, Mark, I just point out, like, if you've listened to this far, thank you. But you should also know this podcast comes from a team of people. It's not just me and Mark. The only reason, despite my real-time screen sharing mistakes, that he even looks or sounds as good as it does is because of producers and a marketing team. So, first of all, happy birthday to Bryson Gunter, who unfortunately has to listen to all this crap, including our mistakes. So, wish him a happy birthday and thank him for making this podcast what it is. And there's also a guy named Christopher Owen hanging out here that actually has to watch us live every time. And that poor guy. So remember, missing the
Marc Laliberte 53:56
Liverpool match right now. So ultimate sacrifice. Yeah, you and them both.
Corey Nachreiner 54:02
So thank you to our producers as well. Cheers.