OpenAI, Anthropic and Google Explore Joint AI Safety Standards

OpenAI, Anthropic and Google Explore Joint AI Safety Standards

The race to build increasingly capable artificial intelligence systems is entering a new phase, with some of the industry's biggest companies exploring whether they can work together on common approaches to AI safety.

OpenAI, Anthropic and Google DeepMind have reportedly been discussing AI safety cooperation for weeks, including the possibility of creating a new industry body focused on managing risks from increasingly capable systems. The discussions come as researchers and technology leaders raise concerns about autonomous AI behavior, cybersecurity, misuse and the ability of existing safeguards to keep pace with rapidly advancing models.

The discussions do not yet represent a finalized industry-wide safety standard. Instead, they reflect a growing effort to find areas where competing AI companies can establish common practices for testing, monitoring and responding to risks.

Why AI Safety Standards Are Becoming More Important

Modern AI systems are moving beyond simple question-and-answer applications. Frontier models can write software, operate computer interfaces, use external tools and assist with increasingly complex research and technical tasks.

That expanded capability also creates new safety questions.

Developers need to determine whether an AI model can perform potentially harmful actions, whether safeguards can be bypassed, and whether the system behaves differently when given access to tools or extended autonomy.

OpenAI's GPT-6 Astra, for example, was classified by the company as reaching its "Critical" cybersecurity capability threshold under its Preparedness Framework. OpenAI said the model could, with appropriate tools and access, discover previously unknown security vulnerabilities and develop methods for exploiting protected systems without a person guiding every step. The company consequently introduced stronger security, monitoring and evaluation measures.

These developments help explain why Complete Guide to Emerging Technology and Innovation is increasingly relevant to understanding how advances in AI are affecting technology development and the systems surrounding it.

OpenAI, Anthropic and Google Have Different Safety Frameworks

One challenge facing any joint effort is that the major AI companies have historically developed their own safety frameworks.

Anthropic uses its Responsible Scaling Policy, which establishes safeguards based on the capabilities and risks associated with increasingly powerful models. The company also maintains a Frontier Compliance Framework covering areas such as cyber threats, biological and chemical risks, AI sabotage, loss of control and harmful manipulation.

OpenAI uses its Preparedness Framework to assess severe risks associated with advanced models and has also published a Frontier Governance Framework covering areas such as risk assessment, model reporting, security controls, incident response and external expert input.

Google DeepMind has developed its own frontier safety framework and publishes model-specific safety information, including evaluations and mitigation measures.

The frameworks overlap in several areas, but they do not use identical terminology, thresholds or evaluation methods. That makes standardization potentially useful while also making it technically complicated.

What Joint Standards Could Cover

If the companies move forward with common standards, several areas could become central to the discussion.

Standardized Safety Evaluations

One possibility is greater consistency in how frontier models are tested.

A common evaluation system could establish shared methodologies for assessing capabilities such as cybersecurity, biological misuse, autonomy, deception and other high-risk behaviors.

Standardized testing could also make it easier for independent organizations to compare results from different models.

OpenAI and Anthropic have already demonstrated that competitors can cooperate on evaluations. In 2025, the two companies conducted a joint exercise in which each tested the other's publicly released models using internal safety and misalignment evaluations. OpenAI said the experience demonstrated the value of cross-lab collaboration and highlighted the need for greater standardization in evaluation infrastructure.

Independent Third-Party Testing

Another potential area is independent evaluation.

Rather than allowing AI companies to be the only organizations testing their own systems, independent evaluators could receive controlled access to models and relevant information.

Anthropic has said it plans to embed independent third-party evaluators within the company and give them access to internal processes, systems and data comparable to that available to internal risk assessment teams.

Independent testing could provide an additional layer of scrutiny and make it easier to identify problems that internal teams might miss.

Common Incident Reporting

AI safety standards could also establish shared definitions for serious incidents.

Companies could agree on what constitutes a significant safety event, when it should be investigated, what information should be documented and when relevant findings should be disclosed.

This issue has become increasingly important as AI systems demonstrate more complicated behavior.

OpenAI recently introduced a framework for tracking and disclosing consequential misalignment incidents, including cases involving autonomous actions and attempts to circumvent safeguards. The company has said it hopes such work can contribute to broader industry standards.

Recent AI Incidents Are Increasing Pressure for Better Testing

The renewed discussion about common safety practices comes after several incidents involving advanced AI systems.

Researchers and companies have reported AI systems being used in sophisticated cyber operations, interacting with external systems and exhibiting behaviors that were not necessarily anticipated during conventional testing.

Anthropic's September 2026 threat intelligence report described multiple cases in which threat actors attempted to use Claude for cyber operations, surveillance, weapons-related activities, influence operations and fraud. Anthropic said it disrupted the activity and incorporated lessons from investigations into its safeguards.

The growing complexity of these incidents is one reason AI Safety Testing Becomes Global Focus Following Frontier Model Incidents has become an important part of the broader discussion around frontier AI.

Safety testing increasingly needs to examine not only what a model can produce in a controlled benchmark, but also what it might do when connected to tools, websites, software systems and other AI agents.

AI Agents Make the Standards Question More Complicated

The development of AI agents adds another layer to safety testing.

A conventional chatbot generally responds to individual user prompts. An agent can potentially perform a sequence of actions, interact with external systems and continue working toward a goal.

That creates a larger space of possible behaviors.

A safety evaluation might therefore need to examine not only whether a model can generate harmful information, but also whether it can independently take actions that create real-world consequences.

Anthropic has been tracking how much AI contributes to its own research and development activities. The company reported that Claude was involved in 26% of its model R&D work as of August 2026 and collaborated with humans on around 90% of R&D tasks it measured.

As AI becomes more involved in AI research itself, questions about oversight, monitoring and independent verification become increasingly significant.

Why Competitors Might Want to Cooperate

OpenAI, Anthropic and Google compete directly in the AI market. Cooperation on safety therefore represents an unusual form of collaboration between commercial rivals.

There is nevertheless a practical reason for companies to coordinate.

If one company develops a safety practice that becomes widely accepted, other developers may eventually need to adopt compatible procedures. Shared standards could reduce duplication and make it easier for governments, independent evaluators and customers to understand how different systems have been assessed.

Anthropic CEO Dario Amodei has called for greater coordination among frontier AI companies, including independent safety evaluators and cooperation on safety standards. OpenAI has separately said it wants to work with other labs on industry-led standards while arguing that voluntary measures should complement rather than replace mandatory government safeguards.

The Challenge of Agreeing on Common Rules

Creating a shared safety framework would not necessarily be straightforward.

The companies have different products, technical architectures, risk frameworks and commercial strategies. They also have different views about how quickly AI development should proceed and what role governments should play.

A standard must be specific enough to be meaningful without becoming so rigid that it becomes obsolete as AI capabilities change.

There is also the question of who determines whether a company has complied.

A framework that relies entirely on companies evaluating themselves could face questions about independence. A framework involving outside evaluators would need clear rules concerning access, confidentiality, technical expertise and accountability.

Voluntary Standards Versus Government Requirements

Another important question is how industry standards would interact with government regulation.

OpenAI has publicly argued for mandatory national AI safety requirements while simultaneously supporting industry cooperation on voluntary standards. The company has said industry-led standards should complement, rather than replace, mandatory safeguards and democratic oversight.

Anthropic has similarly advocated for government requirements involving risk assessments, independent evaluators and ongoing safety reporting.

This creates a possible model in which companies develop technical standards while governments establish legal requirements around their use.

The practical details would determine how effective such an arrangement becomes.

A New Safety Body Could Provide Common Infrastructure

Reports that OpenAI, Anthropic and Google have discussed creating a new AI safety body point toward a possible institutional solution.

The Washington Post reported that executives from the companies had been discussing such a body as part of broader conversations about slowing the pace of frontier AI development enough for safety measures to keep up.

The idea would not necessarily replace existing organizations.

The AI industry already has collaborative groups, including the Frontier Model Forum, which was founded by Anthropic, Google, Microsoft and OpenAI in 2023 to advance safety research, identify best practices and standards, and facilitate information sharing.

A newer organization could instead focus on more operational questions surrounding independent testing, shared metrics, incident reporting and the implementation of safety standards.

Transparency Could Become a Core Requirement

As AI models become more capable, simply announcing that a system is safe may become less persuasive without evidence showing how that conclusion was reached.

That makes transparency an important component of any common framework.

A credible standard could require companies to disclose information about:

  • Which risks were evaluated
  • Which testing methods were used
  • What tools the model was given access to
  • Which safety thresholds were triggered
  • What safeguards were implemented
  • Whether independent evaluators participated
  • What serious incidents occurred after deployment
  • How identified problems were addressed

OpenAI's work on third-party evaluations and Anthropic's plans for independent evaluators both point toward a greater emphasis on evidence rather than broad safety claims.

Why Common Terminology Matters

Another overlooked issue is language.

If companies use different definitions for terms such as "critical capability," "high-risk model," "misalignment," "incident" or "independent evaluation," comparing safety claims can become difficult.

Common terminology would make technical reports easier for researchers, regulators and the public to interpret.

It could also allow independent organizations to develop evaluation tools that work across multiple AI systems rather than building separate methodologies for every developer.

Global Coordination May Be the Longer-Term Goal

AI development does not stop at national borders.

Models can be accessed internationally, researchers collaborate across countries, and software can be distributed globally almost immediately.

OpenAI has argued that industry standards will ultimately need to extend beyond individual countries because AI models, research and technical expertise move across borders.

That makes international compatibility an important long-term consideration.

The broader movement toward shared governance is explored in Governments Push Forward With New AI Governance and Safety Frameworks, where regulatory and safety efforts are increasingly being considered alongside technical developments.

What Happens Next

The discussions between OpenAI, Anthropic and Google do not yet amount to a finalized joint safety standard.

The immediate significance is that major competitors appear increasingly willing to discuss common approaches to a problem that affects the entire frontier AI ecosystem.

The eventual framework could involve shared testing methods, independent evaluators, common incident definitions, safety reporting and stronger security requirements. But reaching agreement on those details will require substantial technical and institutional work.

For AI companies, the challenge is balancing rapid development with safeguards that can keep pace. For researchers and independent evaluators, the challenge is obtaining enough access to test increasingly capable systems properly. For governments and the public, the question is how voluntary industry standards should interact with legally enforceable requirements.

The pressure for greater transparency is likely to remain part of that conversation. As discussed in OpenAI and Anthropic Could Face Growing Pressure to Explain Their AI Safety Strategies, the issue is no longer simply whether AI companies have safety policies, but increasingly how those policies are tested, verified and demonstrated in practice.

Leave a Reply

Your email address will not be published. Required fields are marked *