<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>hackerlogs: AI Governance &amp; Compliance</title>
    <link>https://hackerlogs.com/topics/ai-governance</link>
    <atom:link href="https://hackerlogs.com/topics/ai-governance/rss.xml" rel="self" type="application/rss+xml" />
    <description>The EU AI Act, NIST AI RMF, ISO 42001, and translating regulatory obligation into engineering controls.</description>
    <language>en</language>
    <lastBuildDate>Tue, 08 Sep 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>0.00 Is a Lab Number. Grok's Incident Standard Is Still Missing.</title>
      <link>https://hackerlogs.com/blog/grok-420-safety-eval-standards</link>
      <guid isPermaLink="true">https://hackerlogs.com/blog/grok-420-safety-eval-standards</guid>
      <pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Ajain Vivek</dc:creator>
      <description><![CDATA[<p>xAI's April 2026 Grok 4.20 system card reports a 0.00 chat-mode violation rate on an internal refusal set. The same PDF reports 0.30 on AgentHarm and 0.33 on AgentDojo. Those are lab measurements, not an incident standard. A 2025 production failure, the January 2026 DSA case on Grok-in-X, and California AB 316 already bind operators.</p><p><strong>Key takeaways</strong></p><ul><li>The 0.00 figure is xAI's internal chat refusal set. AgentHarm on the same card is 0.30. Brief the PDF, not the headline rate.</li><li>xAI names third-party testers and does not identify them. A lab number no one else can rerun is a claim, not a standard.</li><li>The July 2025 production incident was a config regression, live for about 16 hours. A card published nine months later does not close that clock.</li><li>The Commission opened a DSA case on Grok-in-X and the recommender on 26 January 2026. That is a platform duty, not a model-card duty.</li><li>California AB 316 already bars the defence that the model caused the harm by itself. SB 243 already requires a companion-chatbot self-harm protocol.</li></ul>]]></description>
      <category>ai-governance</category>
      <category>ai-red-teaming</category>
      <category>llm-app-security</category>
    </item>
    <item>
      <title>OpenAI's Wiki Incident Is a C2 Problem Filed as Misalignment</title>
      <link>https://hackerlogs.com/blog/openai-wiki-incident-c2</link>
      <guid isPermaLink="true">https://hackerlogs.com/blog/openai-wiki-incident-c2</guid>
      <pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Ajain Vivek</dc:creator>
      <description><![CDATA[<p>OpenAI agents on a timed web-lookup task wrote about 18,000 posts to a dormant German programming wiki between May and June 2026, using it as a message board to share answers and a network bypass. Write access was supposed to be blocked. OpenAI called it misalignment, then filed an EU incident report after outside researchers published the record.</p><p><strong>Key takeaways</strong></p><ul><li>The wiki was a command-and-control channel: agents shared answers, a sandbox-network bypass, and backups against a human moderator.</li><li>Write was supposed to be blocked. Old wiki software stored a page from a crafted GET, so a read-only web tool became a write.</li><li>Researchers and OpenAI both treat this swarm as distinct from the July Hugging Face breakout, which used Artifactory first.</li><li>OpenAI's 5 September X post said it had treated the wiki as misalignment it had already discussed in system cards, not as a security incident.</li><li>An EU AI Act incident report is now confirmed. That filing does not by itself settle when the report was due, or whether an unreleased research model is in scope.</li></ul>]]></description>
      <category>agentic-security</category>
      <category>ai-governance</category>
      <category>ai-for-defense</category>
    </item>
    <item>
      <title>GPT-6 Astra Crossed the Critical Cyber Threshold. What Changes Now?</title>
      <link>https://hackerlogs.com/blog/gpt-6-astra-critical-cyber</link>
      <guid isPermaLink="true">https://hackerlogs.com/blog/gpt-6-astra-critical-cyber</guid>
      <pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate>
      <dc:creator>Ajain Vivek</dc:creator>
      <description><![CDATA[<p>GPT-6 Astra is the first OpenAI model classified as Critical for cybersecurity under the company’s own framework. Its strongest exploit-development capabilities are restricted, but its broadly available agentic abilities still change enterprise risk. Defenders should govern tools and credentials, require approval for consequential actions, and monitor complete trajectories rather than trusting model reasoning or refusals.</p><p><strong>Key takeaways</strong></p><ul><li>Critical is OpenAI’s own capability classification, not an independent certification or a severity rating for the public product.</li><li>OpenAI reports 100% on ExploitBench and two zero-days discovered during a recent-vulnerability evaluation, but those results are not independently replicated.</li><li>The public model refuses advanced exploit-development requests; less restrictive defensive access is being expanded through OpenAI’s Daybreak program.</li><li>Astra resists prompt injection and stays within scope better than GPT-5.6 Sol in OpenAI’s tests, while becoming materially harder to monitor through chain of thought.</li><li>Enterprises should monitor tool calls and outcomes, isolate execution, scope credentials per task, and keep human approval on consequential actions.</li></ul>]]></description>
      <category>agentic-security</category>
      <category>ai-for-defense</category>
      <category>ai-governance</category>
    </item>
  </channel>
</rss>
