OpenAI's EU AI Act Statement Skips Training Data: Copyright Gap Activates Sunday
Resumo
OpenAI publicou declaração de conformidade com o EU AI Act que aborda frameworks de segurança e parcerias de cibersegurança, mas omite detalhes sobre compliance no capítulo de direitos autorais que exige resumo público de dados de treinamento; a European AI Office ganha poderes de fiscalização a partir de 2 de agosto com autoridade para impor multas de até €15 milhões ou 3% do faturamento global anual.
With EU AI Act enforcement powers set to activate Sunday, OpenAI published a sweeping compliance narrative on Thursday — detailing its safety frameworks, provenance watermarking partnerships, and cybersecurity cooperation with European agencies — while leaving the one dimension of the GPAI Code of Practice where its compliance was already publicly questioned conspicuously absent. The company's statement, titled "Advancing Responsible AI Across Europe," covers two of the Code's three chapters in meaningful detail. The third, the Copyright chapter, which requires a publicly available training data summary and a documented copyright compliance policy, is not addressed.
That gap matters now because it stops being a theoretical concern on Sunday. Starting August 2, 2026, the European AI Office gains the power to request information, access models, and impose fines of up to €15 million (approximately $17.2 million; exchange rate as of July 31, 2026; conversions are approximate) or 3% of global annual turnover for non-compliance with GPAI obligations — including, specifically, the obligation to publish a training data summary, per EU AI Act Article 53 enforcement.
The August 2 activation is two days away. OpenAI is a GPAI Code signatory. The Code's Copyright chapter requires compliance. The company's public compliance statement does not address it.
What OpenAI's Statement Actually Covers
The "Advancing Responsible AI Across Europe" document is organized around three themes that largely track the GPAI Code's Transparency and Safety & Security chapters: internal governance frameworks, provenance and watermarking technology, and cybersecurity cooperation.
On governance, OpenAI points to its Preparedness Framework — in place since 2023 and updated in April 2025 — as the mechanism for identifying and managing serious risks from advanced AI systems. The Frontier Governance Framework, published May 28, 2026, maps that internal safety architecture onto the specific disclosure obligations the GPAI Code's Transparency chapter imposes. OpenAI also references pre-release model testing, published system cards with major releases, the Red Teaming Network, and a public Model Spec updated in late 2025.
On provenance, OpenAI describes the dual-layer architecture it announced in May 2026: C2PA Content Credentials, an open cryptographic standard that embeds signed provenance metadata into files, combined with SynthID, Google DeepMind's imperceptible pixel-level watermarking system. The combination addresses a specific engineering limitation: C2PA manifests carry rich context but can be stripped by a screenshot or format conversion; SynthID survives those operations but carries less contextual information. Together, they provide overlapping coverage that a single system cannot. OpenAI says it is expanding provenance work to audio outputs and is working toward text-output measures as technical standards mature.
On cybersecurity, OpenAI describes its EU Cyber Action Plan, launched in May 2026, under which the company is working with EU and national cyber agencies, private sector partners, and critical infrastructure operators through a Trusted Access for Cyber (TAC) program. The initiative aligns with the European Commission's Action Plan on Cybersecurity and Artificial Intelligence, published July 7, 2026, which calls for controlled access to advanced AI models for defensive cybersecurity purposes.
What the Statement Does Not Cover
The GPAI Code of Practice has three chapters. The Transparency chapter covers model documentation and information to downstream providers. The Safety & Security chapter covers systemic risk assessment and incident reporting for the largest models. The Copyright chapter requires two things that apply to every GPAI provider, including OpenAI: a documented policy for complying with EU copyright law and the text-and-data-mining opt-out regime; and a publicly available summary of training data, using the mandatory template the European Commission published on July 24, 2025.
OpenAI's July 31 statement does not address either Copyright chapter requirement. The company's EU AI Act Help Center page — updated 17 days ago — links to the Preparedness Framework and Model Spec, but does not include a training data summary link.
This is not a new gap. In August 2025, Euractiv's Maximilian Henning reported that GPT-5 — released on August 7, 2025, five days after GPAI obligations took legal effect — appeared to lack the required training data summary and copyright policy, despite OpenAI being a GPAI Code signatory. The EU AI Act Newsletter published by the Future of Life Institute flagged the same concern, noting that models released after August 2, 2025 must comply immediately, not under the 2027 transitional provision that covers pre-existing models.
Independent research has reinforced the concern at the industry level. A benchmark study published in mid-2026 examining documentation quality across GPAI models — including signatories — found that organizations that have signed the GPAI Code "score only marginally higher overall than non-signatories," with the advantage concentrated in downstream-facing documentation rather than in "the upstream disclosures on training data, copyright-relevant data use, bias mitigation, computing and energy consumption, which are the areas the regulation primarily targets." The study concluded that "surface-level compliance markers, such as Code-of-Practice signatures, should not currently be treated as proxies for documentation depth in procurement decisions."
C2PA and SynthID: What the Technology Actually Does
The provenance work OpenAI describes is substantive and technically specific, even if it addresses obligations distinct from training data transparency. Understanding what C2PA and SynthID actually do — and do not do — helps distinguish genuine compliance from compliance optics.
C2PA, the Coalition for Content Provenance and Authenticity standard founded in 2021 and ratified as an ISO standard, embeds a cryptographic manifest in a digital file's metadata. The manifest records which AI system generated the content, when, what tools were involved, and whether the file was subsequently edited. Any C2PA-compatible tool can read and display that manifest. The limitation is structural: C2PA metadata lives in the file's metadata layer, which can be removed by a screenshot, a platform re-encode, or a format conversion. OpenAI became a C2PA Conforming Generator on May 19, 2026, and attaches Content Credentials to images from ChatGPT, the OpenAI API, and Codex.
Google DeepMind's SynthID addresses that limitation by embedding the provenance signal in the content itself — specifically, imperceptible pixel modifications in images that persist through compression, resizing, and cropping. The watermark is detectable by a neural network but invisible to human perception. OpenAI's May 2026 announcement brought it to all ChatGPT and API image outputs. Audio SynthID is also in deployment; text watermarking remains technically harder and not yet widely deployed at scale.
However, a legal researcher analyzing the Code of Practice for Tech Policy Press found that no single watermarking technology currently meets all four criteria Article 50 imposes: effectiveness, interoperability, robustness, and reliability. The Code addresses this by requiring a layered approach — which is precisely what OpenAI's C2PA+SynthID combination represents — but common evaluation benchmarks for measuring compliance across these four dimensions do not yet exist.
Industry Context: Everyone Is Rushing the Same Deadline
OpenAI is not the only company pushing compliance documentation to the public ahead of August 2. Google announced on July 24, 2026 that it was signing the Transparency Code of Practice, expanding SynthID watermarking partnerships to Apple, ElevenLabs, Kakao, NVIDIA, and OpenAI for interoperable watermarking adoption. Google simultaneously cautioned that layering additional rules onto evolving technical frameworks could create "disclosure fatigue" and confuse users — a signal that the industry views the compliance obligation as real but the compliance infrastructure as immature.
The EU AI Act, operating as a de facto global standard similar to GDPR before it, is shaping how major foundation model providers document and disclose their practices regardless of where their users are located. For OpenAI specifically, the stakes are concentrated in Ireland: OpenAI Ireland Limited is the EU data controller, and Ireland's AI Office handles frontline GPAI enforcement. OpenAI signed an 88,000-square-foot lease at Dublin's Tropical Fruit Warehouse on July 27, 2026, committing to 250 new roles over two years and an €105 million (approximately $120.7 million) total investment. Non-compliance would not only expose it to financial penalties but risk disrupting the European commercial operations that, as the IPO approaches, must demonstrate structural durability to public market investors.
The Grok investigation provides the available enforcement template. The European Commission opened formal DSA proceedings against X in January 2026 after its Grok AI produced approximately 3 million non-consensual intimate images — demonstrating how quickly the EU's enforcement posture can escalate when a specific AI harm is documented and traceable to a product's specific capability. That case is a DSA matter, not an EU AI Act action, but the enforcement architecture is the same agency and the same Commission that gains AI Act enforcement powers on Sunday.
What Comes Next
OpenAI's statement is careful about what it promises. The document describes compliance practices as supportive of the EU framework rather than claiming verified compliance, acknowledges that provenance technology remains imperfect, and notes that standards are still evolving. The company frames its compliance approach as ongoing — "we will keep strengthening our compliance approach and learning from regulators and the broader ecosystem."
That framing is legally prudent. But it also means the July 31 statement does not function as a compliance certification. The European AI Office can now begin verifying GPAI compliance through information requests and, where necessary, enforcement actions. The Copyright chapter's training data summary requirement is not one that can be satisfied by publishing internal frameworks — it requires a published document, in a specific format, available for rights holders and regulators to review.
For OpenAI's pre-August 2025 models, the transitional deadline is August 2, 2027 — more than a year away. For GPT-5, released August 7, 2025, and every model released since, the requirement took effect on the models' launch dates, with no transitional window, per WilmerHale's analysis of Article 53(1)(d).
Whether the European AI Office treats the gap as a compliance priority, raises it through dialogue, or waits for a formal complaint before acting is not yet determined. What is determined, starting Sunday, is that it legally can.
Frequently Asked Questions
Does OpenAI comply with the EU AI Act?
Partially, by its own account. OpenAI's July 31 statement documents compliance with the GPAI Code's Transparency chapter (model documentation, system cards, Preparedness Framework) and its Safety & Security chapter (risk assessments, Red Teaming Network). It does not address the Copyright chapter, which requires a publicly available training data summary under Article 53(1)(d) and a copyright compliance policy. OpenAI's pre-August 2025 models have until August 2, 2027 under a transitional deadline; models released after that date, including GPT-5, had no transitional window. Whether the EU AI Office will treat the unaddressed Copyright chapter as a compliance priority after enforcement powers activate August 2, 2026 is not yet determined.
What is the GPAI Code of Practice, and does signing it guarantee compliance?
The GPAI Code of Practice is a voluntary framework published by the EU AI Office on July 10, 2025, that gives GPAI model providers a documented path to demonstrate compliance with the EU AI Act's obligations for general-purpose AI models. Signing it provides a "presumption of conformity" — regulators presume compliance and focus enforcement on monitoring adherence to the Code rather than conducting full investigations from scratch. However, signing the Code does not guarantee compliance, and a 2026 benchmark study found that signatories score only marginally better than non-signatories on training data documentation quality. The Code has three chapters: Transparency, Copyright, and Safety & Security. All three must be addressed, not just the chapters a company's public statement highlights.
What does C2PA watermarking actually mean for AI-generated images?
When you receive or encounter an image generated by ChatGPT or the OpenAI API, it now carries two overlapping provenance signals. The C2PA Content Credentials embed a signed cryptographic manifest in the file's metadata recording which AI system created it, when, and any subsequent edits — readable by any C2PA-compatible tool but removable by a screenshot or format conversion. The SynthID layer embeds an imperceptible pixel-level watermark that survives screenshots, compression, and resizing, detectable by a neural network. Together, they provide overlapping coverage that neither system provides alone. The EU AI Act requires machine-readable AI content marking as of August 2, 2026, for new AI systems; legacy systems already on the market have until December 2, 2026.
What specifically is missing from OpenAI's compliance statement?
The GPAI Code of Practice's Copyright chapter requires two obligations not addressed in OpenAI's July 31 statement: a publicly available summary of the data used to train its models, using the mandatory template published by the European Commission on July 24, 2025, and a documented policy for complying with EU copyright law including the text-and-data-mining opt-out regime. These obligations apply to all GPAI providers, including OpenAI. A training data summary does not require disclosure of proprietary datasets in full — it requires category-level disclosure of data types, main sources, and collection methods. The absence of a published summary from OpenAI is the specific gap the EU AI Office can now investigate starting August 2, 2026.