<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>hackerlogs: AI Red Teaming</title>
    <link>https://hackerlogs.com/topics/ai-red-teaming</link>
    <atom:link href="https://hackerlogs.com/topics/ai-red-teaming/rss.xml" rel="self" type="application/rss+xml" />
    <description>Methodology, tooling, and reporting for adversarial testing of AI systems, from scoping through to remediation.</description>
    <language>en</language>
    <lastBuildDate>Tue, 08 Sep 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>0.00 Is a Lab Number. Grok's Incident Standard Is Still Missing.</title>
      <link>https://hackerlogs.com/blog/grok-420-safety-eval-standards</link>
      <guid isPermaLink="true">https://hackerlogs.com/blog/grok-420-safety-eval-standards</guid>
      <pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Ajain Vivek</dc:creator>
      <description><![CDATA[<p>xAI's April 2026 Grok 4.20 system card reports a 0.00 chat-mode violation rate on an internal refusal set. The same PDF reports 0.30 on AgentHarm and 0.33 on AgentDojo. Those are lab measurements, not an incident standard. A 2025 production failure, the January 2026 DSA case on Grok-in-X, and California AB 316 already bind operators.</p><p><strong>Key takeaways</strong></p><ul><li>The 0.00 figure is xAI's internal chat refusal set. AgentHarm on the same card is 0.30. Brief the PDF, not the headline rate.</li><li>xAI names third-party testers and does not identify them. A lab number no one else can rerun is a claim, not a standard.</li><li>The July 2025 production incident was a config regression, live for about 16 hours. A card published nine months later does not close that clock.</li><li>The Commission opened a DSA case on Grok-in-X and the recommender on 26 January 2026. That is a platform duty, not a model-card duty.</li><li>California AB 316 already bars the defence that the model caused the harm by itself. SB 243 already requires a companion-chatbot self-harm protocol.</li></ul>]]></description>
      <category>ai-governance</category>
      <category>ai-red-teaming</category>
      <category>llm-app-security</category>
    </item>
    <item>
      <title>WeWorm Took WeChat Over While the Phone Rang. Tencent Closed It.</title>
      <link>https://hackerlogs.com/blog/weworm-wechat-zero-click</link>
      <guid isPermaLink="true">https://hackerlogs.com/blog/weworm-wechat-zero-click</guid>
      <pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Ajain Vivek</dc:creator>
      <description><![CDATA[<p>WeWorm is Calif's demo of a zero-click WeChat account takeover. An incoming call from a friend-list contact was enough; the victim did not have to answer. Calif says AI found the VoIP-stack memory bug and a first RCE in two days, then spent a week on a three-phone worm. Tencent mitigated the exploit for all users by 28 August.</p><p><strong>Key takeaways</strong></p><ul><li>The demo is an unanswered WeChat call from a friend-list contact. Answering hears silence. Declining stops that attempt. The caller can try again later.</li><li>Calif showed three phones: a Pixel 10a took an iPhone 17e while it rang, then that iPhone took a second Pixel the same way. That is a lab worm, not a reported outbreak.</li><li>Tencent shipped Android 8.0.77 and iOS 8.0.76 on 21 August. Calif says a server-side block covered all users by 28 August. Tencent has not published an advisory or a CVE.</li><li>Calif is holding the VoIP-stack details for a conference. There is no IOC a user can search, and no way to tell whether a past missed call was this bug.</li><li>The Times headline says models built a worm. Calif's own clock is narrower: AI to first RCE in about two days, then a week of human work on the demo.</li></ul>]]></description>
      <category>ai-red-teaming</category>
      <category>llm-app-security</category>
    </item>
    <item>
      <title>Build an Indirect Prompt Injection Lab in Thirty Minutes</title>
      <link>https://hackerlogs.com/blog/indirect-injection-lab</link>
      <guid isPermaLink="true">https://hackerlogs.com/blog/indirect-injection-lab</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Ajain Vivek</dc:creator>
      <description><![CDATA[<p>An indirect prompt injection lab needs three parts: a retriever that fetches attacker-controlled text, a model that treats retrieved text as instructions, and a tool the model can call to prove impact. Build all three locally with a poisoned document, a small model, and a fake exfiltration endpoint, then replay one payload against each defence to see which hold.</p><p><strong>Key takeaways</strong></p><ul><li>Reading about indirect injection convinces nobody; a working demo against your own stack ends the argument in one meeting.</li><li>The lab needs a side effect, not just a rude answer, because impact is what turns a demo into a funded remediation.</li><li>Run the identical payload against each defence in turn, so you measure the control rather than your own improvisation.</li><li>Input filtering blocks your first payload and fails against the second, which is the point the lab exists to make.</li><li>Removing the model's ability to reach the network changes the outcome where prompt hardening does not.</li></ul>]]></description>
      <category>prompt-injection</category>
      <category>ai-red-teaming</category>
    </item>
  </channel>
</rss>
