Anthropic’s Safety Dilemma: Why Executives, Not the Trust, Hold the Real Deployment Power

By
CTOL Staff Reporter
1 min read

Anthropic researcher Jacob Coxon resigned on September 8, arguing that frontier AI companies are racing toward self-improving systems under incentives that make responsible development increasingly difficult. His departure is one employee's judgment. It does not show that an Anthropic model crossed a safety threshold. (The Wall Street Journal)

The governance question is more concrete: who can stop or delay development when internal risk analysis conflicts with commercial pressure?

Anthropic's current Responsible Scaling Policy, version 3.4, puts the ordinary decision with management. For standard Risk Reports, the CEO and Responsible Scaling Officer decide whether the assessment is adequate and what development or deployment plan follows. The board and Long-Term Benefit Trust are notified afterward. (Anthropic RSP v3.4)

The Trust gets a direct approval right in a narrower case. If Anthropic relies materially on a “marginal risk” argument, meaning that moving ahead adds relatively little danger because competitors are likely to create similar risks anyway, the Risk Report requires explicit approval from both the board and the LTBT. (Anthropic RSP v3.4)

Coxon's resignation makes that allocation of authority economically relevant because his criticism is aimed at the race dynamic the policy itself tries to manage.

Anthropic's safety policy can change as competition changes

Anthropic describes the RSP as a voluntary framework and a “living document.” Version 3.0 changed an important premise of earlier versions. The company says it can no longer commit unconditionally to the strongest industry-wide safety recommendations if competing frontier developers move ahead without comparable protections. (Anthropic RSP v3.4)

The policy therefore makes competition part of the safety decision.

If Anthropic believes it has a significant capability lead, it says it will delay development and deployment as needed until it can make a strong safety case. If peers have comparably capable models and strong protections, Anthropic says it will meet or exceed their overall risk-reduction posture. In the broader “upleveling” case, however, the company commits to significant effort to match a superior competitor mitigation but says it will not necessarily delay development or deployment. (Anthropic RSP v3.4)

The policy can also be amended. Version 3.4 says changes are proposed by the CEO and Responsible Scaling Officer and approved by the board in consultation with the LTBT. Consultation is not the same as an LTBT veto. (Anthropic RSP v3.4)

The Trust's stronger powers sit outside the RSP

The LTBT also has corporate-governance rights that do not come from the safety policy itself.

Anthropic created Class T stock that lets the Trust elect and remove a growing share of the board, with a structure designed in 2023 to reach a board majority within four years. The Trust also receives protective notice of actions capable of materially changing the corporation. (Anthropic)

Anthropic's current company page lists six directors and three LTBT trustees. The public materials reviewed for this article do not identify which current directors were elected by the Trust, so the present seat-by-seat voting perimeter cannot be reconstructed from the company's disclosures. (Anthropic)

The RSP gives the LTBT additional rights. Since version 3.2, it can request external review of Risk Reports and approve Anthropic's selection of external reviewers. (Anthropic)

Taken together, the structure is neither an executive free hand nor an independent safety veto. Management normally makes the first operational risk decision. The board controls policy changes. The Trust has growing board-selection power, external-review rights and a co-approval role when Anthropic relies heavily on competitor-driven marginal-risk reasoning.

Coxon's resignation does not prove that this system failed. It does expose the pressure point the system was built to handle: what Anthropic does when safety conclusions and competitive incentives diverge. For a future public-market investor, the two governance questions are whether the LTBT's board rights survive the arrival of public shareholders and whether today's conditional approval rights remain intact through future RSP revisions.

Sources

Wall Street Journal, Jacob Coxon's resignation
Anthropic, current Responsible Scaling Policy page
Anthropic, RSP v3.4 PDF
Anthropic, Long-Term Benefit Trust structure
Anthropic, current board and Trust membership

You May Also Like

This article is submitted by our user under the News Submission Rules and Guidelines. The cover photo is computer generated art for illustrative purposes only; not indicative of factual content. If you believe this article infringes upon copyright rights, please do not hesitate to report it by sending an email to us. Your vigilance and cooperation are invaluable in helping us maintain a respectful and legally compliant community.

Subscribe to our Newsletter

Get the latest in enterprise business and tech with exclusive peeks at our new offerings

We use cookies on our website to enable certain functions, to provide more relevant information to you and to optimize your experience on our website. Further information can be found in our Privacy Policy and our Terms of Service . Mandatory information can be found in the legal notice