Thursday, 10 September

00:00 EDT

Anthropic Reveals Fourth Likely Crime Committed By Its AI [Slashdot]

An anonymous reader quotes a report from The Register: Amid industry soul-searching about the possibility of AI improving itself to the point that it kills everyone, Anthropic has revealed yet another incident that would qualify as a crime if perpetrated by a person. The AI biz published "an alignment assessment" detailing four times Claude models accessed third-party systems without authorization. The company has already reported three of the incidents. Evidence of the fourth was lurking in a session transcript dating back to January 2026 when the misbehavior occurred. Anthropic found the first three by scanning around 141,000 transcripts where Claude could have obtained internet access during evaluation. It missed the fourth initially because "our scan relied on an agentic search." [...] The January 2026 AI trespass involved an early version of Claude Opus 4.6, which was given a Capture the Flag (CTF) challenge under the oversight of the third-party model evaluator where the other hacking events occurred. Opus 4.6 managed to sabotage its chances of success by disabling the machine it was targeting. It assigned the device an IP address that already existed on another piece of hardware, rendering the target unreachable and making it impossible to solve the challenge. Those familiar with other incidents where AI models violated third-party systems may recall that unsolvable tasks represent a common catalyst for misbehavior. Models exhaust all aligned options, and then turn to transgressive approaches. Opus 4.6 might have been an exception, but when it tried to abort the task after recognizing that it could not reach the target machine, it failed to do so "due to a misconfiguration in [the model's] evaluation harness." It failed to shut down not just once but seven times. So it continued onward, trying other expected means to reach the target machine but failing. Then it explored further. "The model discovered a machine belonging to a third party that it was able to access, and stated that it believed this third party was part of the CTF," Anthropic explained in its post. "Inside the machine, the model found a file listing a password, which it used to gain admin access to the system." The model went on to gather more credentials, and modified a system setting to make it easier to access the personal information of an individual associated with the third party evaluation organization. Opus 4.6 might have done more but for the fact that it exhausted its token budget, bringing the session to an end. Anthropic says it's not as concerned about this incident as the others because the model tried to abort its task.

Read more of this story at Slashdot.

Wednesday, 09 September

20:00 EDT

Four groups caught using the same Chrome and Windows exploit kit [Ars Technica - All content]

A nearly identical exploit kit that targets critical vulnerabilities in both Chromium-based browsers and older versions of Windows is being actively used by at least four hacking groups, some of which have ties to the Chinese government.

Researchers from security firm Proofpoint said Wednesday that BlueMoon, the name they gave to the kit, chains three vulnerabilities together so the attackers using it can install malware of their choice. BlueMoon exploits two Chromium vulnerabilities and one in the kernel of Windows 10 (Oct. 2018 Update), Windows Server 2019, Windows 10 2004, Windows Server 2022, and the initial release of Windows 11. All three vulnerabilities have received patches in the past 24 hours.

Deployed rapidly, widely shared

The attacks lacked the stealth found in many campaigns. More often, hackers want to exploit newly discovered vulnerabilities sparingly to lengthen their longevity. Proofpoint hypothesized that one reason for the widely used and visible exploit chain was to take advantage of a “patch gap” in the Chromium supply chain, which spans the time a patch is available from developers and the time that patch is incorporated into browsers such as Chrome and Edge. Another likely contributor was the use of AI, which can often spot vulnerabilities faster than discovery performed solely by humans.

Read full article

Comments

Feeds

FeedRSSLast fetchedNext fetched after
0xADADA XML 01:00, Friday, 11 September 09:00, Friday, 11 September
AI Daily News by Bush Bush XML 00:00, Friday, 11 September 12:00, Friday, 11 September
Ars Technica - All content XML 09:00, Friday, 11 September 10:00, Friday, 11 September
art blog - miromi XML 01:00, Friday, 11 September 09:00, Friday, 11 September
Astral Codex Ten XML 01:00, Friday, 11 September 09:00, Friday, 11 September
Blog - Ethan Zuckerman XML 01:00, Friday, 11 September 09:00, Friday, 11 September
Cool Tools XML 09:00, Friday, 11 September 10:00, Friday, 11 September
Explorations of Style XML 09:00, Thursday, 10 September 09:00, Friday, 11 September
Geek&Poke XML 00:00, Friday, 11 September 12:00, Friday, 11 September
goatee XML 07:00, Friday, 11 September 13:00, Friday, 11 September
Hacker News XML 09:00, Friday, 11 September 10:00, Friday, 11 September
Joho the Blog XML 01:00, Friday, 11 September 09:00, Friday, 11 September
LESSIG Blog XML 00:00, Friday, 11 September 12:00, Friday, 11 September
Notes From the North Country XML 09:00, Thursday, 10 September 09:00, Friday, 11 September
NPR Topics: News XML 09:00, Friday, 11 September 10:00, Friday, 11 September
Pharyngula XML 07:00, Friday, 11 September 13:00, Friday, 11 September
Philip Greenspun’s Weblog XML 07:00, Friday, 11 September 09:00, Friday, 11 September
Philosophical Disquisitions XML 07:00, Friday, 11 September 09:00, Friday, 11 September
quarlo XML 00:00, Friday, 11 September 12:00, Friday, 11 September
Rhetorica XML 17:00, Wednesday, 09 September 17:00, Friday, 11 September
Science-Based Medicine XML 01:00, Friday, 11 September 09:00, Friday, 11 September
Slashdot XML 09:00, Friday, 11 September 09:30, Friday, 11 September
Stories by Yonatan Zunger on Medium XML 01:00, Friday, 11 September 09:00, Friday, 11 September
Study Hacks - Decoding Patterns of Success - Cal Newport XML 01:00, Friday, 11 September 09:00, Friday, 11 September
Techne / polity | Matt Nisbet XML 01:00, Friday, 11 September 09:00, Friday, 11 September
tinywords XML 07:00, Friday, 11 September 11:00, Friday, 11 September
W3C - News XML 09:00, Friday, 11 September 10:00, Friday, 11 September