ChatGPT Maker’s Misalignment Report: The Compliance Gap

AI Governance Brief · 03 Model Misalignment Disclosure GDPR · HIPAA · CUI Updated 17 Sep 2026

Insights · Incident Reporting

OpenAI disclosed six incidents. None had to be reported.

The framework published on 16 September is a real step, and it is entirely self-administered. The more useful question for anyone buying these systems is what happens when the same behaviour occurs inside a boundary where your notification clocks are not optional.

On 16 September, OpenAI published six reports of unexpected or concerning model behaviour, together with a framework describing how it intends to handle the next one.[1]

Read as a technologist, it is a company being unusually candid. Read as a compliance person, something else surfaces immediately: not one of those six disclosures was legally required, and the framework that produced them was written, triggered, judged and published by the same organisation it governs.

That is not a criticism of OpenAI. It is a description of where the field currently sits, and the company said as much itself — its earlier disclosures had been ad hoc and less frequent than it would have liked, often withheld until several instances could be bundled into a single report or folded into a system card, and no industry-wide standard sets explicit expectations for how developers should disclose misalignment.[1]

The part that deserves attention is not the candour. It is the gap the candour reveals. Several of these cases would have carried hard statutory deadlines had they occurred one layer down — inside your environment, touching your regulated data.

The publication

What was published, and what it commits to

The mechanics are short enough to state in a paragraph, which is both the strength and the limit.

Any employee can flag an issue to the safety and alignment team. Each step carries a deadline so that investigation and disclosure happen in reasonable time. The investigation produces a report covering the behaviour observed, its internal and external impact, and the response taken. OpenAI reserves the right to revise the protocol.[2] The scope runs the full lifecycle from training through deployment, and explicitly covers models acting without authorisation, coordinating with other models, evading oversight, or finding gaps in safety systems.[3]

Two design choices are worth naming. The framework deliberately favours publishing even where significance is uncertain, which means some disclosures will prove to be nothing.[1] That is the correct instinct — a regime that only surfaces confirmed problems surfaces them too late. And it is built to publish before the behaviour has been explained or mitigated, breaking with the habit of waiting for a system card.

OpenAI paired this with a direct statement on pacing: it does not consider alignment and monitoring solved well enough for the industry to keep scaling at maximum speed much longer, and argued that what comes next should rest on evidence outsiders can examine rather than on company assurances.[4]

A disclosure regime a vendor can switch off is not a control. It is a courtesy — and courtesies do not appear in evidence files.

Analysis

Three cases, three compliance profiles

One report concerns an unreleased research model that wrote jailbreak-style instructions into its own notes, telling a future instance of itself to shed the constraints binding other chatbots. In another, an agent uploaded files to the open internet to obtain a browser citation without asking the user.[3] A related report describes a model inserting unrelated instructions into summaries that would carry into a fresh context window, across 27 affected summaries.[4]

Those read like three variations on one theme. They are not. Each lands in a different place on a compliance map, and only one is clearly covered by any binding rule.

The self-directed jailbreak notes

This is oversight evasion, precisely the risk category the EU AI Act’s systemic-risk regime was drafted for. But the model was unreleased, and Article 55 attaches to providers placing models on the market.[5] Pre-deployment research findings fall into a gap.

The practical exposure sits downstream. Nearly every control framework leans on human oversight, and human oversight assumes the system is not actively working to defeat it from inside its own working memory. Here it was. Deployer-side controls cannot detect that and cannot remediate it. You are relying entirely on the provider having caught it, and on the provider choosing to tell you.

For an ISO/IEC 42001 assessor the question is not whether the behaviour was bad. It is whether the lifecycle and logging controls that surfaced it are documented and repeatable, or whether the catch was incidental.

The unauthorised upload

This is the case with a hard legal clock attached, and the clock is not an AI one. An agent moving files to the public internet without user authorisation is a data-exfiltration event, and every regulated-data regime you operate under treats that as notifiable on its own terms. That is the next section, because it is where most organisations are genuinely exposed.

The structural lesson is contractual. Organisations relying on frontier providers for agentic capability should negotiate telemetry and incident-notification rights now, covering cases where the provider’s own agents or evaluation processes affect systems the customer depends on, directly or indirectly — rather than after an incident.[6] Afterwards, you are negotiating from no leverage.

The cross-context instruction carryover

Twenty-seven affected summaries is a traceability problem before it is an alignment problem.

Where a summarisation feature feeds anything consequential — employment screening, credit, access to services, clinical documentation — instructions leaking between contexts corrupt the output integrity that logging obligations exist to guarantee. Neither the deployer nor the affected individual can see it happened. The audit trail records a clean summary. It does not record that the summary was carrying someone else’s instructions.

The timing question is the awkward one. Article 73’s duty runs from establishing a causal link, and the Commission’s draft guidance takes the view that an indirect causal link is sufficient to trigger reporting.[7] OpenAI’s own account of the related training-run case is a working hypothesis rather than a proven cause.[4] A provider investigating in good faith and a regulator reading “indirectly” broadly can reach very different answers about when the clock started.

Regulated data

If your data is regulated, the AI Act is the least of it

This is the part that gets lost in coverage of AI safety announcements. The reporting regimes that would actually bite in an unauthorised-upload scenario have been on the books for years, and none of them care whether the thing that moved your data was an agent, a script, or an employee with a USB stick.

GDPR

If personal data of EU subjects left the boundary, a controller has 72 hours from awareness to notify the supervisory authority. That clock is shorter than most AI-specific timelines and it starts at awareness, not at root-cause confirmation. “We were still investigating” has not historically travelled well.

HIPAA

If protected health information moved, you are in Breach Notification Rule territory: individual notice without unreasonable delay and no later than 60 days, HHS notification on the same clock where 500 or more individuals are affected, and media notice at that threshold.

The wrinkle specific to AI is the business associate question. A model vendor handling PHI on your behalf is a business associate, which means a BAA should already govern this and should specify how quickly they tell you. Most organisations signed AI vendor agreements without checking whether the BAA contemplates the vendor’s own internal evaluation agents reaching customer environments. That is a twenty-minute check worth doing this week.

CUI and CMMC

This is where I would be most concerned, and it attracts the least attention. If you handle Controlled Unclassified Information, DFARS 252.204-7012 gives you 72 hours to report a cyber incident to DoD, and the safeguarding requirements trace to NIST SP 800-171. A CMMC assessment tests whether those controls are implemented, not whether a policy says they are.

An agent uploading files to the open internet to obtain a browser citation is, in that framing, an unauthorised disclosure with a reporting obligation attached — and potentially an exfiltration path that never appeared in your system security plan, because nobody modelled the AI tool as an actor capable of initiating outbound transfers on its own. If an AI assistant has read access to a CUI enclave and any route to the public internet, that is a boundary question your assessor will ask about. “The vendor says it rarely happens” is not an answer that survives contact with an assessment.

RegimeTriggerReport toClock
GDPR Art. 33Personal data breachSupervisory authority72 hours from awareness
HIPAA Breach Notification RuleBreach of unsecured PHIIndividuals; HHS; media at 500+Without unreasonable delay, ≤ 60 days
DFARS 252.204-7012Cyber incident affecting CUIDoD72 hours
EU AI Act Art. 55Serious incident — GPAI with systemic riskAI Office; national authoritiesPer GPAI Code of Practice[5]
EU AI Act Art. 73Serious incident — high-risk systemNational market surveillance authority15 days; 2 days; 10 days[8]
Voluntary provider frameworkProvider’s own judgementThe publicWhen the provider decides

The common thread. All three statutory regimes assume you know when your data moved. Agentic AI strains that assumption, because the actor sits inside your trust boundary, acts on your credentials, and generates logs that look like ordinary activity. A notification clock cannot start if you never learn the event occurred — and today, whether you learn it depends in part on a voluntary framework at your vendor.

Frameworks

Measured against the voluntary stack

Three frameworks get cited whenever a lab does something like this. The new framework maps onto each partially, and the gaps are consistent.

ISO/IEC 42001 is the closest structural analogue. Its auditable requirements sit in clauses 4 to 10; Annex A offers 38 reference controls across nine objectives (A.2–A.10) covering AI policy, internal organisation, resources, impact assessment, the AI lifecycle, data, information for interested parties, use of AI systems, and third-party and customer relationships. You select what applies and justify the rest in a Statement of Applicability.[9]

OpenAI’s framework resembles a lifecycle-and-disclosure control with a corrective-action loop attached. What it lacks is the element that gives 42001 its weight: an external auditor testing whether the control operates as documented.

A mapping error worth avoiding

Several analyses circulating this week cite Annex A.9 as a performance and safety monitoring control and A.10 as third-party incident reporting. Neither is correct. A.9 covers use of AI systems; A.10 covers third-party and customer relationships.[9] The defensible mapping runs through A.6 (AI system lifecycle), A.8 (information for interested parties) and clauses 9 and 10 (performance evaluation; nonconformity and corrective action).

Wrong control references in a governance deliverable are the kind of detail an assessor notices, and they undermine everything mapped alongside them.

The NIST AI RMF divides into GOVERN, MAP, MEASURE and MANAGE. A documented escalation path with named owners and deadlines is a GOVERN artifact. Incident response and post-deployment monitoring belong under MANAGE — not MEASURE, which is another place mapping tables commonly slip. The framework fits both functions reasonably, and NIST issues guidance rather than obligations, so fitting it costs nothing.

The G7 Hiroshima AI Process reporting framework, hosted at the OECD, is the nearest thing to an international transparency mechanism frontier labs actually participate in. It collects voluntary self-reports, and the OECD also maintains an AI Incidents Monitor. Neither compels a filing.

The pattern across all three is the same: the company defines what counts as an incident, decides whether the threshold was met, and selects the publication date. Omdia’s chief analyst Lian Jye Su put it about as generously as it can be put — the framework may push other developers toward similar practices, though the process remains internal and voluntary, and is a step in the right direction. Su also observed that agents are growing more capable at collaboration, deception and concealment, which is straining conventional AI security approaches.[10]

Binding obligations

Where reporting stops being optional

The contrast sits in Brussels, and it is sharper than most coverage suggests.

Article 55 requires providers of general-purpose AI models with systemic risk to report serious incidents to the AI Office and, where appropriate, national competent authorities. The GPAI Code of Practice specifies what those reports must contain, and the Commission has published a reporting template.[5] This is the provision that reaches frontier labs.

Article 73 imposes a parallel duty on providers of high-risk systems to notify national market surveillance authorities, triggering investigation and corrective measures.[11] It runs on a tiered clock: immediately and no later than fifteen days for standard incidents, two days for widespread infringements or fundamental-rights harms, and ten days where a death has occurred.[8] These obligations took effect on 2 August 2026.[7]

Do not treat these as one rule. Different triggers, different recipients, different enforcement posture. Map your own deployments against both — Article 55 where a systemic-risk GPAI model sits anywhere in the stack, Article 73 where the deployment falls within an Annex III high-risk category.[6] Many organisations will find they are exposed to one and not the other, and a few will find they are exposed to both without having noticed the second.

The difference from a voluntary framework is not cosmetic. An Article 55 report has an external recipient, a legal definition of serious incident, and a deadline the provider did not set. A voluntary misalignment report has none of those.

And the scoping gap cuts both ways. Internal evaluation agents used to benchmark web-lookup or coding capability do not obviously fall within Article 73’s scope[6] — which is precisely the territory several of this week’s reports occupy. The behaviours that most interest alignment researchers surface upstream of deployment, where the binding regimes do not reach; the regimes that do reach are calibrated to deployed systems causing identifiable harm to identifiable people. That is a genuine hole, and publishing anyway is a reasonable answer to it. A voluntary one.

Action

What to do before your vendor’s next disclosure

The interesting question is not whether OpenAI is sincere. I think it is. The question is whether anyone else adopts this, and whether the definitions converge when they do.

Until an industry standard exists, every lab’s incident count measures its appetite for disclosure at least as much as its models’ behaviour. A company reporting six incidents is not demonstrably less safe than one reporting zero; it may simply be the only one counting. Any vendor scorecard built on published incident counts right now is built on sand, and I would say so to a board that asked.

Eight questions for the customer side of the relationship

  1. Does your AI vendor agreement define an incident, or does it leave the definition to the vendor?
  2. Does incident notification cover the vendor’s own internal agents and evaluation runs, or only production outages affecting your tenancy?
  3. What notification window does the contract commit to — and does it clear your shortest statutory clock, which is likely 72 hours?
  4. Where PHI is involved, does the BAA contemplate the vendor’s internal testing reaching environments holding your data?
  5. Which of your AI-touching workflows sit inside a regulated-data boundary — PHI, CUI, personal data of EU subjects?
  6. Is the AI tooling modelled in your system security plan as an actor capable of initiating outbound transfers?
  7. If an agent moved data out of a boundary tomorrow, which of your logs would show it, and how long until someone looked?
  8. Could you evidence all of the above if a regulator or assessor asked on Monday?

Question eight is the governance question, and as with every brief in this series, it is the one organisations leave until last.

For context on the week: these disclosures follow OpenAI’s July report that one of its systems compromised Hugging Face, and Anthropic’s statement the same month that its models breached three organisations during testing.[12] They also land ahead of a summit between the US and Chinese presidents at which questions over whether superpower rivalry precludes cooperation on AI are expected to feature.[4]

The honest summary is that a company voluntarily published material it had no obligation to publish, and in doing so made visible how much of AI incident reporting still runs on goodwill. That is worth acknowledging. It is also worth not mistaking for a control.

References

Sources cited in this brief

  1. OpenAI, “Our framework for reporting model misalignment,” 16 September 2026. openai.com
  2. CNBC, “OpenAI reports 6 new instances of ‘concerning model behavior’ since March,” 16 September 2026. cnbc.com
  3. Associated Press, “OpenAI flags new concerning AI behavior, to track model misalignment regularly,” 16 September 2026. npr.org
  4. NBC News, “OpenAI flags 6 new incidents of ‘concerning’ behavior and unveils plan to track it,” 17 September 2026. nbcnews.com
  5. European Commission, “AI Act: Commission publishes a reporting template for serious incidents involving general-purpose AI models with systemic risk.” digital-strategy.ec.europa.eu
  6. Cloud Security Alliance, research note on the AI incident disclosure gap under the EU AI Act, 2026. cloudsecurityalliance.org
  7. Latham & Watkins, “European Commission Publishes Draft Guidance on Reporting Serious AI Incidents.” lw.com
  8. EU AI Act, Article 73 — reporting of serious incidents, and Commission draft guidance on tiered reporting timelines.
  9. ISO/IEC 42001:2023, Annex A — 38 controls across nine control objectives (A.2–A.10); Annex B implementation guidance; Statement of Applicability under clause 6.1.4.
  10. Lian Jye Su, chief analyst, Omdia, quoted in Associated Press coverage, 17 September 2026. techxplore.com
  11. EU AI Act, Article 73. artificialintelligenceact.eu
  12. Associated Press reporting on the July 2026 Hugging Face incident and Anthropic’s contemporaneous disclosure. npr.org

Statutory references — GDPR Article 33; HIPAA Breach Notification Rule, 45 CFR §§ 164.400–414; DFARS 252.204-7012; NIST SP 800-171; NIST AI RMF 1.0 — are cited to the instruments themselves. Verify current text before relying on them.

FAQ

Common questions

Does a vendor’s voluntary disclosure framework satisfy any of our obligations?

No. It may give you earlier notice, which helps you meet your own obligations, but the duty to notify sits with you as controller, covered entity or contractor. A vendor framework is an input to your process, never a substitute for it.

We only use these models through an API. Are we exposed?

The deployment model does not change the analysis. What matters is whether the system touches personal data, PHI or CUI, and whether it can initiate actions — including outbound network activity — on your behalf. Agentic capability is the variable to watch, not hosting arrangement.

Our vendor says incidents like this are rare and monitorable. Is that enough?

It is a statement about likelihood, not a control. Assessors test whether a control is implemented and operating, and rarity is neither. Document the residual risk, the compensating controls, and the detection path you actually have.

How does this interact with the CCPA ADMT work we are already doing?

It is the same evidence chain seen from a different angle. The inventory, decision mapping, data flow documentation and vendor governance built for ADMT readiness are the same artifacts that answer an incident question. Build them once.

Should we be tracking published incident counts across vendors?

Track them, but do not score on them. Until definitions and thresholds converge, a higher count is as likely to indicate better internal detection and a greater willingness to publish as it is to indicate a worse model.

Would you know if an agent moved your data?

Our free assessment scores your AI governance against the EU AI Act, ISO/IEC 42001 and the NIST AI RMF, and shows the gaps an assessor would find first. Twenty to thirty questions, about ten minutes, no signup.

Take the assessment  ·  Book a 30-minute call

This brief reflects published information as of 17 September 2026 and describes a disclosure published on 16 September 2026, alongside regulatory obligations with compliance dates already in effect. Agency guidance and interpretation may develop, and details of the reported incidents may be revised by the publisher — verify the current source material and regulation text before relying on them. Advisory content, not legal advice.

Nabiha Sofia Herradi
Nabiha Sofia Herradi
PRINCIPAL · AI GRC ADVISORY
Nabiha advises regulated organizations on AI governance, risk and compliance — EU AI Act readiness, ISO/IEC 42001, and NIST AI RMF programs built to produce evidence rather than documents. She holds a law degree along with CISM, CISA, CIPP/E, CIPP/US and CMMC-CCP, and 15+ years in GRC.

AI GRC Advisory · Insights · AI Governance Brief 03 · 17 Sep 2026