If AI makes more of what people need easier to produce, what would it take for people to actually have enough? We should ask that before treating greater capability as a sufficient definition of progress. A useful service can exist without reaching the people who need it. A system can work exactly as intended while leaving those people with less say over their lives.
Common Obligations begins with a distinction: what technology makes possible and what our institutions make available are not the same thing. Ownership, income, public provision and political power help determine the distance between them. These arrangements should be open to argument, not treated as fixed scenery around a changing technology.
What is greater productivity for?
Imagine that a service becomes cheaper to provide with AI. The saving might become a lower price, better quality, shorter working hours or a larger return to its owners. A publicly funded service might reach more people, or its budget might be reduced. These are possible choices, not forecasts. The technical achievement does not tell us which outcome to prefer or who should choose it.
This is more than a question about replacing jobs. If substantially less paid human work were needed, an arrangement that distributes income mainly through employment would face a difficult question: how should people claim a share of what can be produced? New work and lower prices might offset some disruption. Shared ownership, income support and public services could play different roles. None should be assumed to appear automatically, and each needs an account of funding, power and who could still be excluded.
Money need not disappear for that question to matter. Making some services cheaper would not by itself remove limits on housing, energy, materials, human attention or the infrastructure needed to deliver them. Nor does a need become unimportant because meeting it offers little commercial return. We should distinguish the ability to produce, the incentive to provide and a person’s practical ability to obtain what they need.
Who owns the possibility?
Cheap, reproducible intelligence could widen participation. It could also increase dependence on whoever controls essential infrastructure or access. An open model does not supply everyone with computing resources; a low subscription price does not guarantee that its terms will remain acceptable. The question is not only whether a tool is available today, but whether people can build on it, leave it and contest decisions made by its owner.
Capitalism’s existing arrangements are neither guaranteed to survive unchanged nor destined to disappear because AI becomes more capable. A change in ownership would not, by itself, end destructive competition or unaccountable power. Private firms, public institutions and cooperatives should all face concrete questions: who receives the benefit, who carries the cost, and what can people do when the arrangement fails them?
My starting position is that people deserve a voice beyond their value as workers, customers or investors. That is a political commitment, not a conclusion supplied by a model. It calls for public choices about access, ownership and provision, with affected people able to challenge those choices. It does not prescribe one ownership model for every activity.
When the race defeats its purpose
A firm might reduce its costs by reducing paid work. If many firms did so without other ways for people to receive income, who would be able to buy their output? Lower prices, new demand and different institutions could change the outcome. The point is not that collapse follows, but that an individually sensible decision does not settle the collective result.
The same distinction matters in the race to build AI. Each participant can feel compelled to move faster while preferring a world in which everyone could proceed more carefully. Such competition might undermine the trust, security or livelihoods on which its participants depend. It need not produce its own remedy. We should ask which rules make restraint viable, which rewards favour useful provision, and who has the authority to change those rules.
A practical starting point, not a complete economic programme
The six scenarios examine parts of this larger problem. Surveillance asks who may exercise power. Release asks what authority a system should receive. Defense and discovery ask whether capability becomes protection and usable knowledge. Incident response and coordination ask who must answer, even when doing so is costly. They cannot tell us how to distribute all the benefits of AI. They can make particular responsibilities harder to evade.
The organisations building the most powerful AI systems have a difficult conflict of interest. They are trying to establish whether their technology is safe while competing to make it more capable. The governments overseeing them have a version of the same problem: protecting the public sits alongside the ambition to lead an industry that could reshape economic and military power.
Both can sincerely want a good outcome. Neither can settle, on everyone else's behalf, how much risk everyone else should accept.
This is where the practical obligations begin: what AI builders and operators should owe the people affected by their systems, wherever those people live. I want a framework that a person can understand, an engineer can implement, and an independent body can examine. Its commitments should apply to foundation model developers, agent companies, data suppliers, and organisations putting AI to work, with requirements that reflect what each actually controls.
The central claim is simple: the freedom to build powerful AI should come with responsibilities to people who never chose to use it.
A race nobody can settle alone
In An Alien Mind, published on 6 September 2026, OpenAI's chief scientist Jakub Pachocki describes a tension directly relevant to this problem. He argues that pursuing recursive self-improvement is necessary to remain at the research frontier, while questioning whether accelerating that process is the right collective choice. He calls for shared safety requirements, external enforcement, and international coordination.
His assessment is a participant's view, partly grounded in internal results that readers cannot independently inspect. It deserves scrutiny alongside the proposal itself. But the conflict he describes does not depend on accepting his forecasts: a company can believe collective restraint would help and still fear the consequences of practising it alone.
Imagine two competitors evaluating a new capability. Both would prefer the other to spend another month testing. Neither wants to be the only one that does. An appeal to responsibility asks each to absorb a cost without knowing whether its competitor will reciprocate. Shared requirements could change that calculation, provided a broken commitment is detectable and has consequences.
The same problem survives at national scale. A government might support caution in principle while treating a rival's progress as a reason to accelerate. Some countries have much more influence over this decision than others. A country importing AI services can still bear the costs of failures originating elsewhere, with little access to the evidence behind deployment decisions.
We should be precise about what coordination would accomplish. It would make certain obligations harder to escape by switching providers or jurisdictions. It would also give cautious organisations a better answer to investors and customers asking why a competitor is moving faster.
The institution needs a job description
The distinction is not simply between promises and meaningful checks. Amodei’s September 2026 pacing proposalincludes concrete provisions for outside review. The coordination scenariodescribes those proposed terms and their limits. We should credit such steps without treating an announcement as proof that an arrangement is operating.
An international AI regulator is an appealing answer. The nuclear comparison helps, provided we understand what makes it work. The International Atomic Energy Agency's safeguards use technical verification under agreements accepted by states. They have defined objects of inspection and a basis for access. The existence of an international organisation, by itself, is insufficient.
AI presents a different verification problem. A model can be copied, adapted, and connected to tools after its original assessment. A system's power depends on its deployment conditions as well as its training. Inspecting the original model will not tell us everything about a service built around it, particularly when that service can take actions or delegate work.
There is already substantial work to build on. The OECD AI Principles address human rights, transparency, robustness, and accountability. NIST's AI Risk Management Framework offers a voluntary approach to managing risks throughout a system's life. The UN established an Independent International Scientific Panel on AI and a Global Dialogue on AI Governance in August 2025, providing mechanisms for assessment and discussion.
My proposal borrows from that work. Its contribution would be to connect a short public commitment to evidence and an accountable decision. Before designing an organisation's headquarters or voting structure, we should be able to say what it would ask a builder to demonstrate, and what would happen when the demonstration fails.
I would start with six obligations.
1. Keep responsibility traceable
Every consequential AI system should have an identifiable organisation responsible for its operation, and a clear account of the responsibilities held by its suppliers. The person affected should have somewhere to go when something fails.
Consider a hiring service assembled from a foundation model, a candidate database, and an agent that ranks applicants. The employer controls the decision to use the ranking. The agent provider controls its workflow. The data supplier controls the information it supplies and what it says about that information. The model developer controls the underlying release and its documented limitations. Those responsibilities overlap, but they need not disappear into the overlap.
Each participant should document the decisions it controls, preserve the evidence needed to investigate failures, and identify the organisation receiving that responsibility at the next handoff. For the data supplier, this includes origin, permission to use the data, known gaps, and a process for corrections. A claim that a dataset is representative needs an explanation of whom it represents and for which purpose.
This obligation does not make every supplier responsible for every downstream act. It does require a usable record of who knew what, who could change what, and who decided to proceed. An organisation deploying a system should remain answerable for that deployment even when its investigation later identifies a supplier's fault.
2. Match power with evidence
The more consequential a system's capabilities and permissions, the stronger the evidence required before expanding them. That evidence should identify the version tested, the conditions of the test, the failures observed, and the limits of what the results establish.
A model that suggests a database query and an agent that executes it against a hospital's records have different opportunities to cause harm. Adding credentials, persistent memory, external communication, or the ability to delegate can change the risk even if the model remains identical. The assessment should follow those changes.
A useful safety case is an argument someone else can inspect: here is what could go wrong, here is why we believe our controls address it, and here is the evidence that might prove us wrong. Passing a benchmark is one input. It cannot establish that an entire deployment is safe in every environment.
For the most powerful systems, this obligation should reach decisions about further development when development itself creates material risk. The relevant thresholds would need public justification and technical revision. I do not think we can responsibly invent a universal number in an essay and call everything below it safe.
The requirement must also be affordable in proportion to the risk. A small tool with narrow permissions should have a straightforward way to demonstrate compliance. A large company should face demanding requirements when the capabilities warrant them. Revenue and headcount are poor substitutes for examining what a system can do.
3. Make scrutiny independent
Builders should enable qualified outside scrutiny, with access proportionate to the claims and risks being examined. For higher-risk systems, the builder should not be able to make an unfavourable assessment disappear by choosing a different evaluator.
Independence needs an operating model. Who chooses the examiner? Who pays? Can the examiner run its own tests? Can it report a serious concern to an authorised oversight body without the client's permission? A company-funded audit can still be useful, but these questions determine how much confidence it deserves.
Access to evidence and the freedom to publish findings are meaningful safeguards in their own right. They are not the same as authority to require a response. An assessment should make a decision better informed; it should not give its author unchecked power to stop work. The responsible authority, its duties and the route to challenge its decision must be specified too.
Public accountability does not require publishing private training records, security vulnerabilities, or every model artefact. Sensitive evidence can be examined through controlled access, with public summaries explaining the scope, findings, and unresolved concerns. Restrictions should have stated reasons and a route to challenge them, because commercial confidentiality can otherwise become a permanent answer to every difficult question.
Evaluators should also face scrutiny. Their methods, conflicts of interest, and material errors need examination. Independent assessment improves the basis for a decision; it cannot turn uncertainty into a guarantee.
4. Bound autonomy and prepare for failure
An AI system that acts should have explicit limits on its authority and tested ways to contain failure. Before deployment, the operator should know which actions require approval, which resources are accessible, how permissions expire, and what happens when the system behaves unexpectedly.
Suppose an agent can issue refunds. Its spending limit should be enforced by the payment service. A sentence asking the model to stay within budget is insufficient. If it delegates to another agent, the delegated authority should remain within the original permission. Stopping the parent should also address outstanding child work and credentials that could remain active.
Recovery deserves equal attention. If a payment request times out, the system should establish whether it committed before sending another. If an agent changes a record, the operator needs a way to identify and repair the change. A stop button may prevent the next action; it cannot unsend an email or reverse every action already taken.
Some releases are difficult to recall at all, including publicly distributed model weights. In those cases, the assessment must account for the limits of later intervention before release. Openness can support scrutiny and wider participation, but neither an open licence nor a closed API establishes safety on its own.
The general obligation is to demonstrate control appropriate to the system, and to be honest where control ends.
5. Report serious failures so others can learn
Builders and operators should report material incidents and significant near misses through a shared process. A finding that affects other systems should reach the people who can act on it, even when disclosure is commercially uncomfortable.
If an agent finds a way around an authorisation boundary, quietly patching one product may leave the same weakness in another. A useful report would distinguish what was observed from what is suspected, identify affected versions and conditions, and describe the containment measures. It should be possible to update the report as the investigation improves.
This requires agreed definitions of severity, reporting deadlines, and recipients. It also requires protection for personal information and security-sensitive details. A practical approach could combine prompt confidential notification to a competent body with a later public account of what happened and what changed.
The incentives matter. A process that punishes every honest report as if it were concealment will encourage silence. Accountability should distinguish responsible disclosure from repeated negligence, misleading claims, and deliberate suppression. Employees need a protected route to raise serious concerns when internal reporting fails.
6. Give affected people a way to challenge
People should be able to discover when AI materially influences a consequential decision about them, understand enough of its basis to contest it, and reach someone with the authority to correct it.
In the hiring example, an applicant needs a route to correct an erroneous record or challenge an unsuitable assessment. Sending them a technical explanation of a model does little if nobody can reconsider the decision. A meaningful appeal needs an owner, a timescale, and the power to provide a remedy.
This also applies to people whose data enters a system. Providers should explain the uses they make of it and offer workable processes for access, correction, and deletion where applicable. They should describe technical limitations honestly, including the difference between removing a source record and undoing its influence on an already trained model.
Human agency extends beyond individual appeals. Workers, affected communities, and countries using systems developed elsewhere should have a role in shaping the standards. Access to that discussion should not depend on owning a training cluster. Participation will take funding, translation, and technical support if it is to mean anything beyond an invitation to comment.
These six obligations form a floor, not a complete account of justice or shared prosperity. They fit together: identify who is responsible, require evidence, enable scrutiny, contain failure, share what goes wrong, and give people recourse. Their implementation will vary. Their purpose should remain recognisable across the supply chain.
When the boundary fails during testing
Research and evaluation can themselves expose outsiders to risk. The incident chapter brings together OpenAI’s account and reporting on Anthropic, then separates those documented accounts from an illustrative response. It asks how far containment should extend, how investigators establish the coverage of a review, and what evidence should support resumption.
Preserving logs is only a beginning. Investigators should reconcile the activity that occurred with the records they examined, disclose missing evidence, and explain why their search was wide enough. A review that starts only from known failures can reproduce the blind spots that hid them.
A repair should lead to a new, bounded claim: this activity may resume, with these permissions, under these conditions. That claim needs an accountable decision-maker, scrutiny proportionate to the risk, continuing monitoring, and explicit reasons to stop again. Broader pauses and independent examinations have costs; their scope and duration need justification and review.
What happens when a commitment fails?
Outside review can change behaviour by revealing a problem, prompting a repair or informing people who rely on a system. Shared obligations go further: they establish who must respond to a serious finding, under what rules, and with what consequences if the finding is ignored. Publication, a required response and enforcement are related, but distinct.
I would begin with a public record for a specific system and version. It would state which obligations apply, what evidence supports them, who examined that evidence, what remains unresolved, and when reassessment is due. Companies could adopt this structure before an international agreement exists. Buyers could then require the record in procurement, and contracts could specify access for review and the consequences of misleading claims.
For higher-risk systems, my preferred direction is enforceable requirements through public authorities, supported by independent technical evaluation. International agreements could establish common minimums and recognise assessments across participating jurisdictions. Domestic institutions would retain responsibilities for enforcement and remedies. That distributes the work while giving the shared requirements a basis beyond company promises.
An adverse finding should trigger a response proportionate to the danger: correction of a misleading claim, restriction of particular permissions, withdrawal of a deployment, or a pause on a defined development activity. The decision should explain the evidence, the scope of the restriction, and the conditions for resuming. Emergency powers need time limits and independent review, so a precautionary intervention cannot drift into an indefinite ban without justification.
Verification will remain incomplete. Some actors will refuse access, and some states may not participate. An international body would face those limits too. Recognising them helps define an honest initial goal: make compliance inspectable and consequential among participants, then expand participation and improve the ability to detect evasion.
That would be progress even before universal agreement. It would not justify claiming that the world's AI risks had been brought under control.
Who keeps the rule-makers accountable?
The strongest objection to this proposal is that it could help concentrate the power it is meant to oversee. Expensive assessments, proprietary tests, and licensing rules shaped by incumbents could make independent development harder while leaving the largest companies comfortable.
That possibility should shape the design from the beginning. Assessment methods should be open to examination, with multiple qualified evaluators and support for small organisations facing legitimate testing costs. Requirements should be justified against capabilities and deployment conditions. New evidence should be able to overturn them. A provider's business model, including whether it releases open models, should not automatically settle the assessment.
The institutions would need published funding, conflict disclosures, transparent appointments, and representation beyond the countries and companies with the most compute. People subject to their decisions should have an appeal route. Industry expertise is necessary, but a company's technical knowledge should not entitle it to decide the acceptable risk for everyone else.
There are limits to agreement as well. A shared framework will not settle every country's views on speech, every dispute over data rights, or every question about distributing AI's economic benefits. It should state its scope clearly while protecting a meaningful minimum. Cooperation that depends on ignoring the rights of people with the least political influence would fail its own purpose.
I am deliberately leaving the final institutional shape open. A treaty body, coordinated national authorities, and a network of accredited evaluators have different strengths and failure modes. We should compare them against the same practical questions: can they obtain the evidence, act on it, withstand pressure, and be held accountable themselves?
When progress creates new obligations
Accountability also matters when a capability works. The defensive-access scenario asks who should receive tools that can find and exploit weaknesses. Restrictions can reduce misuse while excluding under-resourced defenders. Access programmes should explain their criteria, permit challenges, and connect findings to repair capacity. A count of vulnerabilities found is a poor substitute for evidence of protection delivered.
The same responsibility extends through software supply chains. A model’s assessment cannot vouch for every gateway, dependency, or tool server around it. Operators should know which artefacts actually ran and what authority they inherited. After compromise, a clean update may be only one part of recovery: exposed credentials, persistence, and downstream installations need investigation too.
The discovery scenario tests a different definition of success. Generating a result, verifying its argument, and developing shared understanding are distinct contributions. Our reading of Tao’s discussion of related fluid-equation work is that the community’s ability to explain and extend an idea matters beyond the achievement of a solved problem. His post precedes OpenAI’s announcement and should not be treated as a review of that proof.
A builder claiming a research breakthrough should provide precise claims, usable evidence, contribution records, and a plan for explanation and review. Where the original authors cannot develop the work further, an explicit, supported handoff is preferable to silently leaving that work to others. This is a proposed application of traceable responsibility and inspectable evidence, not a requirement that every discovery receive a company or regulator’s permission.
These cases also constrain oversight itself. Costly reviews can entrench incumbents; restrictive access can leave defenders exposed; compulsory approval by established experts can exclude unfamiliar ideas. The framework should make those effects visible and contestable. Its purpose includes preserving the conditions for widely shared progress.
A proposal we can put to work
Common Obligations is an initial proposal. These principles still need sharper definitions, examples from actual deployments, and people willing to argue with the difficult parts. Several draw directly on existing governance work; the test is whether expressing them together makes responsibilities easier to understand and enforce.
A useful next step would be to apply them to three different systems: a foundation model release, an agent with authority to act, and a dataset supplied for consequential decisions. For each, we should be able to name the responsible parties, the evidence required, the independent checks, and the remedy when a commitment fails. Where we cannot, the framework needs more work.
We should also examine a service people depend on: who owns its productive capacity, how its gains are shared, who cannot obtain it, and what voice those people have. A safe system serving only those already able to pay would not settle the question with which this essay began.
I want powerful AI to help more people have enough, with more freedom to shape their lives. That requires both accountable systems and public choices about the economy they enter. We need not predict the end of capitalism, or assume its present form is permanent, to start making those choices explicit.