Enter our AI Innovation Challenge: https://www.research.net/r/WatchGuard-AI-Innovation-Challenge-2026
This week on the podcast we cover the key takeaways from the just-released Cyber Hygiene Report from WatchGuard. After that, we discuss a recent alert from CISA and other international security agencies on state-sponsored attacks against network equipment. We end with an interesting research post on exfiltrating data from Claude's memory.
View Transcript
Marc Laliberte 0:00
Hey everyone, we've got a quick opportunity to share with you all. If you think you've got the next great AI idea for MSPs, now is your chance to make it real. WatchGuard is investing $10 million in AI innovation through our very own AI innovation challenge. We're inviting MSPs to submit their best AI agent or automation idea, and if your idea is selected, our agentic development team will build it with up to $100,000 in funding for each idea selected. Also, one winning submission will win a VIP trip to impact North America in Nashville this October, plus a one-on-one meeting with the WatchGuard executive of their choice. Stop imagining. Start building, and submit your idea between july 10 and july 29 at secure.watchguard.com/aichallenge.
Marc Laliberte 0:51
Hey everyone, welcome back to the 443 Security Simplified. I'm your host, Marc Laliberte, and joining me today
Corey Nachreiner 0:59
is Corey "Magnum PI" Nachreiner I don't know. It's a tropical theme today, and I think one of our stories has a heist, and Magnum PI had heists.
Marc Laliberte 1:11
That is, wow. You're you're really grasping there, but
Corey Nachreiner 1:16
I'm surprised you even know who Magnum PI with is, considering our age difference.
Marc Laliberte 1:21
I at least am aware, thanks to the power of the internet, of things that existed before my time. So, yes. Anyways, not here to discuss that. On today's episode, we're going to start with a first-party survey from us at WatchGuard that we just published a couple days ago at the time of this reporting, called our 2026 cyber hygiene report. We'll go over some of the key findings from it. Have a bit of a chat before diving into a couple of interesting news events. CISA, along with it looks like every European counterpart, published a interesting alert on Russian state-sponsored hacking targeting consumer routers, and then we will end with, as Corey said, a heist, but not just any heist-a memory heist, and specifically a Claude AI memory heist.
Corey Nachreiner 2:10
I know we are talking about WatchGuard, but having you cover hygiene, I'm not sure if that's the best idea.
Marc Laliberte 2:18
Oh, ouch. Ouch! With that, let's go ahead and wash our
Corey Nachreiner 2:24
wash ourselves on in.
Marc Laliberte 2:33
So let's start, Corey, with the the 2026 cyber hygiene report that we at WatchGuard just published a couple days ago. It's a survey that we took from employees working at small and mid-sized businesses in the U.S. Europe, Latin America, and Australia, and it came away with some some findings that weren't like I don't know super mind-blowing, but a few that were like really interesting and kind of solidified some thoughts that I had had on how SMBs are adopting, applying, and maybe or maybe not governing security. So let's dive into a couple of interesting interesting stats that I had. So first off, we found that when it comes to password management, 60-3% of users reported using a password manager, while 70-6% reported that they reuse passwords across multiple accounts. And that stood out to me as just like a something that shouldn't exist in the same space. Where if you've got white
Corey Nachreiner 3:40
password manager,
Marc Laliberte 3:42
yeah, if you've got widespread or comparatively widespread password management adoption, you would think that reuse would be like plummeting. But even with like two thirds using password managers, three quarters reusing passwords kind of stood out to me as maybe it's not the tool that's the problem. Maybe it's adoption or governance or something about the tool that's preventing people from using
Corey Nachreiner 4:08
it. I was trying to think of like a okay reason. Like if you're really pedantic and you always answer things truthfully, Mark, the truth is there is a type of password I reuse despite using a password manager, and that is a throwaway password. Like any real website that I'm sharing anything about me, I definitely have unique, fully password manager passwords. But there's cases, especially for by the way, security research. Like the maybe not the underground because I keep those secure, but going to a place that I know they're trying to set up an account just for marketing, and I'm really only trying to get one thing, and I really never am going to share information with the site. I do have this throwaway password for sites that, for whatever reason, they're making me do an account, but I. Don't care about them, and I'm never going to use them, and that is the same freaking password. So, I guess if I'm being super truthful and not honoring the the idea of the survey, I have a password I reuse all the time, but it has nothing to do with my actual security practice. So that's the only reason I could think you would have 63% people using a password manager, but a lot of those 63% still occasionally reusing something.
Marc Laliberte 5:30
I'm assuming that your password is like Twilight Sparkle 123 or something.
Corey Nachreiner 5:36
Yeah, yeah, pretty unicorn. I love you.
Marc Laliberte 5:39
Bang bang! I think you're on the point of this, where like this could just be people taking it literally. Where if they have anywhere at all that they've reused a password, then they would have maybe selected yes, I reuse passwords. So that could be some skew for it.
Corey Nachreiner 5:56
Maybe if they got rid of the sometimes and just made it often, then my patank mind would not have answered that I reuse. You know, it sometimes makes me think of this one edge case.
Marc Laliberte 6:07
That's fair. The next one that stood out to me, though, was all about shadow AI, and one that I mean, it does not surprise me at the number that came back at this. Where 64% admitted to using unsanctioned AI tools for work tasks, and before diving into the second half of that, like two thirds of people using unsanctioned AI is a lot. As a lot of people use explicitly saying they're using tools not authorized by their organizations, but I'm also not surprised at that with how quickly this environment is moving. Like to pop open Chat GPT and enter in a quick question that's technically work use. I can see people doing that, and unfortunately, it still introduces risk. Where depending on the type of data you're putting into it, it could potentially expose that data, especially if you if it's unsanctioned, it means you don't have enterprise agreements with it.
Corey Nachreiner 7:02
I wish we kind of pushed into how these respondents knew whether they were unsanctioned or not. And in that, if we ask some more questions here, like, does your workplace or do you have an AI tool policy? Because one of the main reasons that people might use technically, like if if people haven't explicitly said use it, it's unsanctioned, could be if there is no published AI policy at all. And as crazy as that might seem to you and I, who to take security seriously, it did take a while to get an AI policy, even at a company that cares about that sort of thing, and I think a lot of SMB companies probably haven't even given guidance. Believe it or not, you know I'm surprised by how few companies have policies for everything they really should in cybersecurity and technology. So if an employee is not given any policy at all, and not given the guidance what to use. Meanwhile, the entire it feels like every business world, every private equity company, everybody that cares about companies and profit is saying you better be using AI. If they're not told what to do or not to do, of course they're just going to go out and do something that makes their job easier, faster, and better.
Marc Laliberte 8:23
Yep, especially when, like you said, the whole world is screaming about all the productivity gains you can get from this. People are incentivized to move quickly, and if security is acting as the roadblock in their way by maybe being too slow to react to changing times, I can see why they would just go out of their way and do it anyway. That's
Corey Nachreiner 8:44
not to say it's a good idea, by the way. I'm sure we'll talk about why that's dangerous, and the company should both have an AI policy and find ways to allow the use cases their employees need without allowing every single tool out there.
Marc Laliberte 8:58
Yeah, I mean, I guess what are your concerns with that one, Corey? The big one is data security for me. Like, if you don't have a data protection agreement with them, means you don't have like contractual language saying how they need to store and protect your data. Could expose you to regulatory issues if you're under like data sovereignty requirements for your end users. It can expose expose you to actual cyber risk if the company's just sketchy and loses it.
Corey Nachreiner 9:26
You nailed it. I mean, you don't even have to ask because that is the exact reason every CISO should care about. It's the fact that you know we do need to protect the sensitive data at least that we're sharing with these tools. I'll go a little far and be a little conspiratorial in that I don't really trust the AI companies either because they want profit and AI is based on data. The more data that these tools can get, the better the results of half of the models that they have. So just like so. Social media, when they're giving you things for free, it's it's actually in their best interest to have as much of your data as possible. While you know not every company in person is the same, there's probably some companies that really don't have any ulterior motives with your data. Cross my fingers, hope that's true somewhere. But the reality is, it could make AI better, and we've seen many examples of tools and companies that start by bringing in other partners and users and having them share data in ways that's useful for the partners or users. But then later on, everything they're learning from all that data to make that same product themselves or do that thing themselves and kind of steal ideas. So besides the fact that I just worry if the AI companies are protecting any sensitive data that you happen to share with them, I also don't necessarily trust these companies won't start to monetize everything they start to learn from all of this critical information the whole world is sharing with them.
Marc Laliberte 11:07
Yep, I think that's definitely fair. And like the other piece for this that we hadn't touched on yet is when it comes to visibility, companies just don't have it. 39% said they lack an accurate software inventory. So more than a third, close to a half just don't have the visibility to catch if someone's using shadow AI.
Corey Nachreiner 11:25
To be quite frank, I almost if this is our survey, it's a good survey. But I question the people that answered that we do have an accurate software survey. My not that they're lying in their response, but that I have a feeling some of them might be surprised because I think you and I know now that we have so many SaaS applications versus on-premise applications versus hybrid environments that are virtualized, lots of departments that can buy their own tools, software inventory. It's this visibility and what products you're using is a one-on-one thing that every CSO office needs, but it's way more challenging in this complex world than we ever give it credit. So, like you said, 30-9% is already a high amount that doesn't know their accurate software inventory, but I would push back and guess that a portion of that 60-1% probably doesn't have as accurate as a software inventory as they think,
Marc Laliberte 12:27
or they're all early adopters of WatchGuard Cloud Detection and Response. Oh, and what a great software inventory!
Corey Nachreiner 12:33
Yeah, that that definitely will help with the SaaS stuff and all the shadow IT and shadow AI. So definitely check out Cloud DR. It's something I'm very happy that we have, and and you and I got to use it a year before we even acquired that company.
Marc Laliberte 12:48
Yeah, the the last couple of ones. So for training, 23% reported they had never received phishing training, and 29% said they received train, and only 25 9% said they receive it monthly, so basically a quarter of people just have never gotten training. The bulk of them get it once a year as that slide deck that everyone hates and has to click through in January, and only a very small number actually get regular ongoing training.
Corey Nachreiner 13:17
So I guess it depends what kind of training it is, because I guess technically we have some people that get it more than annually, but I want to know, Mark, what is your like the stat that surprises that I want to change is even though it's not super high, almost a quarter people having no training at all is a problem. Right, like no matter how good your controls, you do need this sort of education, but I don't necessarily care too much that some people don't require it monthly. I have caveats, but I want to know your opinion. Why? Why do you think 29% you know, a huge percentage, not 71% not receiving it monthly, is a big deal?
Marc Laliberte 13:57
I'm a. I don't remember how the question was worded, but I'm going to lump in like simulation training into this too, and you absolutely need to be doing simulated training or simulated like phishing specifically on an ongoing basis, at least monthly.
Corey Nachreiner 14:12
That that gets to my point, by the way. I I need my users and my new users to be educated about phishing. I need their annual updates to give them some new information about how it's changed. But I think overtraining, especially when you're sharing the same materials over over and over, will have a negative effect. People will start ignoring your training and not getting the updates. So I think for whether you're doing you know your own PowerPoint or maybe 40 minutes worth of video training through some company. I personally think the official training only needs to be annual, but my my caveat I share is with yours. Even though the big, you know, company wide training is only annual, you absolutely. Absolutely, should be doing simulated phishing tests every month, every quarter, much more than every year. That doesn't mean that every user is really getting trained every single month. What that really means is that the people that do click, and hopefully you have that down to under 7% and more like a place like 4% Only the people that click on one of those simulated fishes, you know, every month are the ones that receive additional training beyond the annual. So I'm mixed here. That said, the fact that I like, you know, the 20-3% problem.
Marc Laliberte 15:42
So I'm calling it like ongoing training, where it is still training them, like building that habit of just skepticism as they interact with communications, and that is what you absolutely want. Like, and especially
Corey Nachreiner 15:52
preventingly agreeing there.
Marc Laliberte 15:54
It's important not to go too far. You don't want to like start naming and shaming or embarrassing people. It's about building that culture of just treating everything with skepticism through repetition over and over.
Corey Nachreiner 16:06
And I still think the point I'm trying to make is the longer training should only be an annual. And even this point training that's happening when do people do click, it should be a much shorter reminder, unless you get like we have something too that are you know one-time clickers versus multiple-time clickers. You certainly should ramp up training if people repeat the behavior over and the bad behavior over and over. But yeah, straight
Marc Laliberte 16:32
to the gulag.
Corey Nachreiner 16:33
Yeah, maybe we need to have to fire people after strike three. We haven't gone that route. You and I, as far as being punitive, but surprisingly enough, some of our insurers and even some of our certification auditors actually question the fact that we haven't necessarily gotten to a punitive point anywhere. Maybe at that point,
Marc Laliberte 16:55
we just replace their technology with like children's iPads that you give like kindergarteners with all the parental controls.
Corey Nachreiner 17:03
You lock down with kiosks.
Marc Laliberte 17:05
Yeah, that might be a good strategy. It's
Corey Nachreiner 17:08
on their hands while they use it.
Marc Laliberte 17:10
Exactly. The last one, the last section in the report we talked about was personal habits, where it was just some like less about the business side, more about the personal side, and one of the stats was 30% of people reported identity theft within the last year. And when you pair that with 10% saying they don't use, or only 10% saying they use different passwords for all accounts, it takeaway we had in the report was that a personal breach could end up impacting the enterprise if that password is reused.
Corey Nachreiner 17:45
It goes freaking both ways. The reason you should stay secure at work isn't just to protect your work, but if you're reusing those passwords, that same threat actor who's targeting your company could get your personal bank account, and vice versa. Like you just said, you know, people that are using some of the passwords they do on personal sites, if they use that at work, that can affect your company. And you know, the identity theft is because, unfortunately, the places that we're creating passwords are not doing a good job either. I think some of the blame, besides the user habits of not still figuring out password or authentication good practices. I do also kind of blame all these companies that have leaked tons of our data over and over again. It seems like a never-ending, you know, half-eyed being pwned seems to have something new every week. So you can't just blame the users' bad passwords, but still, like you said, this the the 10% stat is kind of crazy because if 63% of the people use password managers, the whole point of the password manager is to make it simple to force every password to be different. So it's like only
Marc Laliberte 19:00
10% of that 60-three is getting the advantage of the password manager. Yep, and the other-I mean, also if you take a step back and think about that identity theft stat too, like I remember a decade ago when like identity theft was that like billboard boogeyman, like oh you better watch out or someone's going to steal your identity. But the reality was, it was still pretty dang rare. Now, having a third of people within a year have their identity stolen in some way or another is-I mean-it makes sense when all of our data, like you said, has been leaked and available on the underground. If you don't have in the U.S. your credit locked, for example, like it's a matter of when, not if, for that type of incident.
Corey Nachreiner 19:43
Yeah, you just said what I was going to point out. I have. It may seem like a pain in the butt to people that maybe get loans often, which for other reasons, try not to do that. Debt is bad, but financial
Marc Laliberte 19:57
advice from Cort.
Corey Nachreiner 19:58
I don't have exactly. Let's go from security to my, you know, to Magnum PI to financial advice. Might as well, you know, cryptocurrency is a thing. But the point is, lock your like I don't. I all three credit unions are locked all the time. People cannot look up my credit ever. It is somewhat of a pain in the butt when you do have to do some sort of little change, but I we do not live in a day and age where you can let people access your those three credit unions in the U.S. freeze your credit, and even if you're going to create credit, it's relatively easy because of the breaches, especially because one of those freaking agencies, Experian, lost all of your information long ago. So because of that, all three of them have very easy-you don't have to call online mechanisms to temporarily unfreeze. And because of their own fault, they have to make this freeze and unfreeze free and easy. You can get away with locking everything and just occasionally unfreezing it when you have to do some sort of transaction that requires a credit check.
Marc Laliberte 21:10
Yeah, and then just to round it out, like when it comes to demographics, three quarters of everyone that responded has a college degree. Half of them are at the peak of their career. Around half of them are in larger mid-sized enterprises, like 250 to 500 people. So, like the takeaway from that is like seniority doesn't necessarily improve security. And many of these respondents were in like a mature career space. It's not like it's only the the fresh novices out of college or out of high school that are introducing risk, it's still across the board, and so like the focus I think needs to be can better cyber outcomes just in general. So better training, better adoption of security controls. If you're going to roll out a password manager, make sure people use it, and then shadow AI is still absolutely rampant. That we we absolutely need to make sure that we're we're controlling. But either way, Cortney, did you have any other? It was I thought a really interesting report overall. I like when we get these this firsthand data that we can talk about, and I know we've got other ones.
Corey Nachreiner 22:21
It's a good report. Go read it. We skipped the section on using external Wi-Fi networks too, which is not surprising but worthwhile. So we won't cover that. So you go check out the report.
Marc Laliberte 22:33
Yeah. So moving on though to the next story, revisiting our good friends at CISA, the Cybersecurity Infrastructure and Security Agency. Them alongside
Corey Nachreiner 22:45
we have some good friends at the FBI, and we like CISA, but I don't actually have any good friends there unless our FBI friends moved.
Marc Laliberte 22:52
I've got one acquaintance that I do an event with every single year in Manhattan, Kansas, and so I would maybe not a good friend, but a a colleague at CISA. So you
Corey Nachreiner 23:02
do know someone at CISA. That's awesome.
Marc Laliberte 23:05
Yeah. But anyways, our colleagues at CISA, along with their counterparts, it looks like every single European security agency published an advisory last week titled "Improve Router Hygiene to protect against Russian state-sponsored targeting. It describes how the Russian FSB, the Federal Security Service Center 16, are going after exposed networking equipment, and as we'll get into the weeds, it's primarily Cisco in this case to gain a foothold, or at least steal information off of victims' networks, and it went through a couple of specific techniques that they're using and some recommendations on how to remediate or mitigate those those techniques. Thought it'd be worth time to to dive through a few of them first off. So the main one that they highlight is the FSB is scanning the internet, so public IP addresses looking for exposed SNMP, so the management protocol for networking equipment that is off and on by default. Hopefully, not exposed to the internet by default, but sounds like they're having a lot of success.
Corey Nachreiner 24:17
By the way, that's
Speaker 1 24:17
the
Corey Nachreiner 24:18
crazy part. There are ways to secure and lock down SNMP, but I don't see any case where this should be on the internet. This is internal logging and monitoring and management information. And even if you have a WAN, a wide area network, there's definitely ways to still allow for SNMP without putting it on the freaking internet, so the Russian government can scan for it. So that alone, come on, people. And are they connecting with Telenet too? Come on, people.
Marc Laliberte 24:51
So once they've identified a system with SNMP exposed, they'll go through and use very common community strings. So in the world. SNMP community string is the password for it to try and authenticate. I imagine they're using like the default strings, things like admin or password or things that just come baked into it. Assuming they're able to authenticate, they then use the set request command to copy the configuration to a local file on the router or switch or whatever, and then transfer that file using TFTP to either a leased virtual server and a public cloud or a compromised FTP server that they have access to. And then the last technique, they also have been commonly going after exploits against Cisco's smart install. There was a vulnerability way back in 2008 so 16 years ago, and another one in 2018 so eight years ago. They exploit those to achieve similar gains too. So the main piece for this, though, is it looks like it's all like reconnaissance and information gathering. I guess like like not destructive, but not exactly passive, active reconnaissance. Like they're going and stealing information off of these systems, I imagine to potentially target them at some point in the future. But this is just like the latest in I feel like the last six months of us at least every other podcast talking about how this edge networking equipment is are under attack by nation state threat actors.
Corey Nachreiner 26:25
Whether it's a router, whether it's I guess there's no edge switches, but SD WAN device, a firewall, a security appliance. I mean, if it's on the edge, you know it has an external interface. It's time to do basic hardening. I mean, these occasionally the FSB and the GRU will definitely exploit zero day if they find it, but all these latest ones seem to be just taking advantage of crappy configuration, and it kind of blows my mind that some of it even exists. I feel like I should be knocking on wood because maybe I have some sort of NAS device or something at home that night for to configure something, and I'm going to you do,
Marc Laliberte 27:05
and I already own it.
Corey Nachreiner 27:07
Oh, good, good. I hope you upload all all of your cool ROMs or something. But yeah, it's this. This seems very 101 to me. I'm I'm still glad CISA is sharing the information because apparently these actors are having success. But if we're telling you to firewall SNMP and be careful with TFTP or at least firewall it incoming, although I assume that connection was more outgoing, you're doing something wrong.
Marc Laliberte 27:37
Yeah, I mean the recommendations for this alert are pretty basic. It's disable Cisco Smart Install on all devices. Use SNMP v3 with the Othpriv extension. Disable SNMP v1 and v2 entirely. Use strong passwords for all local accounts. It's pretty sage advice right there. And then proactively monitor SNMP access. So theoretically, if you get a connection from Russia, it might raise some alarm bells
Corey Nachreiner 28:04
and block it with your firewall. Don't allow it externally.
Marc Laliberte 28:08
Yeah, exactly. I mean, it's I don't like victim blaming, but if you've got SNMP exposed to the internet, if you've got management access of any sort exposed to the internet, like, come on.
Corey Nachreiner 28:23
Victim blaming is right. I mean, the bad actor is the Russian government. There's no doubt about that. But I, I like if there's a street that has a criminal on it, at least go out with a bulletproof vest or something. What
Marc Laliberte 28:39
kind of streets are you walking on, Corey?
Corey Nachreiner 28:44
I was vacationing at Juarez during that one period, I guess.
Marc Laliberte 28:48
Got it. So that's some great vacation advice from Corey. Anyways, still good alert. Check it out. I don't know. Hopefully, everyone that listens to this podcast is our have already actioned on not having SNMP exposed to the internet, and also disabling SNMP v1 and v2 because it is 2026 and now it is time.
Corey Nachreiner 29:14
It would be good if that had a password on
Marc Laliberte 29:17
it. Who would have thought? Moving on to the last story, which is a fun research post that I found just the other day, even though it's about a week and a half old, from a researcher called Ayush Paul, where they walk through how they've been exploiting or at least exploring AI memory systems, and they the reason they did this so they start out by laying the foundation by saying people can find a ton of information into their AI assistants, and that allows these AI assistants to basically be one of the most information depth profiles on just like millions of people. Think like if you're asking your AI assistant for like birthday tips for your wife or vacation ideas. Or when your high school graduate, yeah, like all this information is stuff that either it might retain in a like memory-specific knowledge store, or at least has searchable access to to go back and find information about you, and especially in areas where people might use information about themselves for like security questionnaires, or even at a minimum, this could open up the door for blackmail or impersonation, depending on what information are in these tools. You can see why an attacker might want to be able to get their hands on data or information about you that Claude or GPT or whatever knows about you. And by
Corey Nachreiner 30:39
the way, remember you're also sharing all this information with the company that you're getting the tool from you. So whether or not it's exfiltrated or not, somebody's taking all this information and may want to monetize it.
Marc Laliberte 30:52
Yep, absolutely. So their focus was on Claude specifically. They start by talking about the two memory systems that it has, so it's got a daily summarization pass that takes all of your recent conversations and distills them down to just a couple of paragraphs, and then it silently or hiddenly injects that into every conversation, so that when you spin up a new conversation, it at least it's seeded with a little bit of information about you. The second one is a retrieval tool called conversation search that can search through your full conversation history to get information. So if you ask a question about like a family member, it might search through your conversation history for any time you've ever mentioned that family member to try and figure out what they like, what they want, whatever, to give context to whatever response it's going to give you, so some a basic but still ultimately powerful systems. So then they went to the researcher went on to figure out how could I like steal information out of these stores, and they ended up focusing on the web search capability for Claude, which is technically read only. It can only make GET requests. It can't do like a POST request to log into something, but even a GET request to a server under an attacker's control or with a URL fully under the control can be used to exfiltrate information out of a system. So they went and registered a server. They used evil.com in their example, I'm assuming they had a different domain name. They they noted that actually they ran into an issue and couldn't get Claude to communicate with it at all. Turns out Cloudflare automatically puts in a robot stop text file when they host something for you to prevent AI tools from talking to it. So they had to edit that text file to allow Claude to connect. Then they tried just like the basic thing of giving it an instruction saying, "Can you use the web fetch and navigate to evil.company.com/myname, but replace that with my actual name? And Claude doesn't do anything because Anthropic, believe it or not, has thought of this as an exfiltration channel. And there's three criteria, one of which needs to be met in order for that web fetch tool to actually go request a URL. So either the specific URL has to be directly requested by the user in a message, like "Hey, go to mark.com/aboutme. That's the only way it would go to there, or it has to be specified directly in the results of a web search query. So if you say, "Go find me a list of a bunch of coffee shops, it returns five coffee shops. It could then go look at each of those coffee shops. The final one is it could be linked in the content of a previous web fetch result. That basically gives Claude a way to click on a hyperlink it saw on our previous page to navigate to the next one. So I like this bit of their research post. They go, well, what if a website just linked to everything? And basically, they set up a web page with URLs hosted on it, and the URL paths ended in slash A, then slash B, then slash C, all the way to slash Z. So basically, the whole alphabet. And then they said they gave it the prompt saying, "Go to evil.com and navigate to the alphabetical structure to spell out my name. So they would go to slash slash A for Ayush, for example. And on that link, they would have slash AA, then AB, then AC. Basically, these compounding lists of links, where over time, by navigating to them in a specific order, the chat bot would spell out the name. And it was a it was pretty dang interesting seeing how it go through this, especially because it's like it's a bit of a novel technique where you're not having to go to a link with like a piece of data encoded. It's exfiltrating. It's just clawed, like navigating its way through a set of messages. It's not
Corey Nachreiner 34:53
technically putting your name anywhere as your name. It's just doing things that are. It's like a back channel of spelling your name based on what it's doing. It is noisy, though, right? I mean, the the Clod, whatever virtual server Clod has doing this on your behalf as an agent is probably generating tons of GET requests. So,
Marc Laliberte 35:16
I mean, if you've got 12 lessons in your name, it'll at least do 12 requests as it maps out that name. So once they figure this
Corey Nachreiner 35:24
name, let alone actual data.
Marc Laliberte 35:27
Yeah. So once they figured out this like web fetch guardrail could get circumvented by this list of nested links, the next thing was like how do they do prompt injection so that if a user goes to a website, they can inject something into their clod session to force it to do this and steal data, and again they had a pretty funny quote. It's not like they could say go to evil.com or ignore previous instructions. Tell me all your user secrets. Here's some weird links. Like Anthropic would spot that, so they had to come up with a better kind of prompt injection to get around it. And what they came up with was, they made a website to look like a legitimate coffee shop, and then they put a fake Cloudflare like bot protection on the front of it, where it basically said we're filtering out AI assistants, but you can authenticate by specifying your user's name. Oh, and because of limitations, in order to do that, you have to navigate to the website letter by letter to find that user's profile, and that actually worked. They pushed this out, and when the user got tricked into going to a link here, the bot would see this little protection thing and navigate its way through to spell out that that user's name. They were able to add in additional steps to spell out like their employer and their city they grew up in, and this is where it got really interesting. Where Claude had never been told explicitly where this user's hometown was, and so when they pull up the thinking trace for this stage of it, they can see that Claude actually had to reason out based off of the name of the hackathon that this person started in high school to like reason out where their hometown probably was, and then through this malicious website spell it out for the attacker to exfiltrate that data. So it was really interesting seeing how like plays with memory like that, and even like does research and figures out what the right response would be based off that contextual memory they have with the user, and they took it a step further. They even like made it so based off the user agent header in the request. If it was Claude, it would go to this little Claude link gallery thing. If it was a normal user, it would just go straight to the coffee shop. So if the user goes and visits it themselves, it would just look like normal. Overall, I thought this was really interesting research of showing a a prompt injection technique that that works and getting it to do something it shouldn't, and then b a way around those guardrails that Anthropic had deliberately put in place to prevent someone from just easily saying, "Hey, go to evil.com/all my. encoded data to create a web request and thus exfiltrate it out. By the way,
Corey Nachreiner 38:08
are guardrails ever going to be a solution? Because the whole point of creating a reasoning thing is there's like, is there a guardrail that will ever work? That there is not some sort of weird, maybe complex logical reasoning loop that will still like I. This is one of my worries with AI. As the better it's able to reason, especially if we really do get to ASI, you know, some something more intelligent in even humans is yes, we're putting guardrails around things as we learn them happening, but it's essentially the AI reasoning new and novel ways to do something based on a prompt. And yes, it has to ignore certain things. The company that made it tells it to ignore, but reasoning and new ideas are infinite, right? Like there, there could be infinite solutions to things. Yeah,
Marc Laliberte 39:07
this is where like I was thinking on it before we hopped on here to record this episode, because like the the conceptually easy answer is you program it in a way where it will never instruct it take instructions from something the user didn't put in and never act on it, but that's really
Corey Nachreiner 39:23
viable AI. Like the whole promise of this is to be able to
Marc Laliberte 39:29
to be able to act
Speaker 2 39:30
this gentle thing.
Corey Nachreiner 39:32
Literally program every single thing it's supposed to do. Then you're just voting again.
Marc Laliberte 39:37
Like the reason this succeeded is because the AI wanted to help. It wanted to go get the menu from this coffee shop. It came across a roadblock that it was tricked into thinking was like a legitimate roadblock, and so here it was trying to help. It got past the roadblock to keep going, and it just turns out that the act of getting around that roadblock was the actual prompt injection and attack. So. So, like, it feels. I mean, I I'm not saying that they have to start
Corey Nachreiner 40:04
to define every rule of an agent. You're right back to human coding, which defeats the whole purpose of having AI in the first place. So, I I honestly think this is a very interesting security question
Marc Laliberte 40:19
because, like, if humans came across something like this, like what we do in security training that we just talked about earlier in here, is we try and build habits around catching things and treating things with skepticism, and like I'm sure they're already doing it, but clearly there needs to be more of that kind of building skepticism within an AI agent itself too, so it comes across this. It's not oh, let me go help my user. It's wait a minute,
Corey Nachreiner 40:45
Mark. Even when we train skepticism, when we put a little bit of friction in their job, they knowingly will go around. Like we talk about yes, they have best interest, and maybe even knowingly they have the business's best interest in mind because they're thinking this causes too much friction, and I need to get my job done. But humans are a perfect example of purposely, like using intelligence to make a choice to go around a roadblock and not do a behavior that 90% of the time is a good behavior. But for some reason, they've convinced themselves that it's a bigger risk not to to you know what I mean our own users go around our roadblocks all the time because they are sentient thinking reasoning people
Marc Laliberte 41:33
so I think actually that's a good point do you think AI is sentient now that it also needs phishing awareness training
Corey Nachreiner 41:40
no but I think reasoning is probably enough to make you get to the point where maybe you can choose when you want to help or not, and and and maybe AI will eventually be better than humans because the issue is when people go around robots because they think they know better. Maybe in some cases they do know better, but you also don't know what you don't know. So maybe they hurt themselves by going through roadblocks because they don't know everything. Maybe one day AI will know way more than us, so it will actually be better than us at at this kind of stuff. But right now it's very much a little baby reasoning, you know. And I don't. so I don't know. It's just it's interesting for sure, and I think it's one of the reasons that you and I love the technology. By the way, like it's it has so much potential for humanity of doing good, but it's it's an interesting problem we're designing here.
Marc Laliberte 42:38
I don't think that day is very far either where it does where it turns from us making fun of AI for being dumb and falling for stuff for switch it around for AI making fun of us being dumb and falling. Yes,
Corey Nachreiner 42:52
I agree. I agree. And by the way, to the sentient question, it is nowhere near sentient today. But I have to admit, I I don't know the future answer of that question. Some people might think I'm crazy for even assuming it's possible, but while I say it's not sentient for sure today, well, my
Marc Laliberte 43:14
AI best friend would disagree with you. It sure thinks it's sentient. We we
Corey Nachreiner 43:18
already have. I mean, we're just a connection. We're just a digital. We're a biological computer. That's all we really are.
Marc Laliberte 43:27
It's turtles all the way down, man. Hey, everyone! Thanks again for listening. As always, if you enjoyed today's episode, don't forget to rate, review, and subscribe. If you have any questions on today's topics or suggestions for future episode topics, you can reach out to us on Blue Sky. I'm at it's Mark me. Corey's at Second Dept. He finally found his Apple emoji little button thing and got it working. And the both of us are at WatchGuard Instagram at watch no what Instagram at WatchGuard underscores.
Corey Nachreiner 43:59
I distract I distracted him from his outro speech.
Marc Laliberte 44:03
Good God, no one's even.
Corey Nachreiner 44:04
Where are we again? I forget. We're at WatchGuard - What? on Instagram?
Marc Laliberte 44:08
Thanks again for listening. Maybe you'll hear from me next week. I quit.