COMMON
OBLIGATIONS
An independent proposal001 / A future we can question

The future is a
series of decisions.

Who gets to make them?
Who lives with the consequences?

Explore what could change
Conceptual monochrome aerial view of a city at night, its roads illuminated by moving traffic
PROGRESS IS COLLECTIVE.
POWER SHOULD BE ACCOUNTABLE.
COMMON OBLIGATIONS
FIELD NOTES / 01
Millions of lives. A handful of decisions.AI-generated conceptual illustration

We don’t all build AI.
We all live with it.

A model leaves a lab. An agent gains permission to act. A camera becomes part of a searchable network. Each step can be useful. Each changes who has power over someone else.

Common Obligations asks what AI builders owe the people affected by their systems. Explore three decisions, examine the tradeoffs, and trace the checks that could make a difference.

How to read thisThese are conditional scenarios, not forecasts. Change one condition at a time. The outcomes are arguments to examine, not calculated predictions.
01 The watching city02 The release decision03 The race to lead04 Common ground

A camera sees a car.
A network sees a life.

A useful investigation tool can become a much broader source of power. The question changes as the system grows.

Conceptual illustration of a generic security camera overlooking a lone pedestrian in a city at night
ILLUSTRATIVE SCENARIO01 / COLLECTION

A narrowly defined purpose

Find a
stolen car.

A camera records a vehicle passing a street. Investigators look for a car connected to a reported theft.

01 / COLLECTION

Who decides what gets collected?

A camera records a vehicle passing a street. Investigators look for a car connected to a reported theft.

OPERATOR DISCRETION

A faster path to use.

The operator can deploy quickly and adapt collection to investigative needs.

WHAT IT ASKS US TO ACCEPT

The same organisation defines the purpose and decides whether its own collection is proportionate.

WHO FEELS THE DIFFERENCE

Investigators gain flexibility. Residents have fewer opportunities to shape collection before it begins.

02 / CONNECTION

Who can search across the network?

Records from many cameras become searchable together. A local tool can reveal movement across a much larger area.

OPERATOR DISCRETION

A wider field of view.

Broad authorised access can help investigators connect sightings across locations and organisations.

WHAT IT ASKS US TO ACCEPT

More access expands the opportunity for misuse and for searches beyond the purpose residents originally understood.

WHO FEELS THE DIFFERENCE

Investigators gain reach. People captured by the cameras can face scrutiny far beyond a single location.

03 / CONSEQUENCE

What turns a result into evidence?

A search result enters an investigation. Someone must decide how much weight it deserves before taking action.

OPERATOR DISCRETION

Judgment stays with the operator.

Investigators can respond quickly using the result alongside their existing practices and expertise.

WHAT IT ASKS US TO ACCEPT

An uncertain match can acquire more authority than the evidence supports, especially under pressure.

WHO FEELS THE DIFFERENCE

Investigators retain discretion. A wrongly identified person may bear the cost of discovering the mistake.

Ground this in the real worldFlock, facial recognition & the limits of accuracy+

Useful does not settle legitimate.

Flock offers automated licence plate recognition and a national LPR network. Its trust materials describe its approach to privacy and customer controls. Those are the provider’s account of its safeguards, rather than independent proof of how every deployment operates.

The scenario above is a general policy comparison, not a claim that every Flock customer lacks these checks. Plate recognition and facial recognition are different technologies. Both raise questions about who can search, for which purposes, and with what recourse.

A documented failure.

Robert Williams was wrongfully arrested in Detroit in 2020 after police relied on an incorrect facial-recognition result. The 2024 settlement introduced requirements for independent supporting evidence and an audit of past cases. The ACLU’s account and settlement links come from his legal representatives.

Our inference: a safeguard has to govern how a result becomes an action. Better matching alone cannot decide whether a use is justified.

THE CHECK THAT MATTERS

Ask whether a use is justified
before it becomes ordinary.

Obligation 06 / Give people recourse ↗
Conceptual monochrome data center with rows of server cabinets and a distant doorway admitting daylight
Behind every capability, a physical system.AI-generated conceptual illustration

More capable.
Ready for what?

A model answering a question and an agent authorised to move money are different systems. Evidence needs to follow the power we give them.

MODEL

Can suggest
an action

+ tools & permissions
AGENT

Can change
the world

Same underlying model. Different opportunities for harm.

What should happen before wider release?

Choose a condition to examine its strongest argument and its costs.

Benefits arrive sooner.

The strongest argument

Developers know their systems intimately. Letting them assess readiness can bring useful capabilities to people sooner and avoid an expensive external bottleneck.

The unresolved cost

The organisation making the case for release also benefits from it. External parties may be unable to examine the evidence or challenge the decision.

Depends on

The quality of internal testing, honest disclosure, and credible consequences for misleading claims.

↳ Evidence before deploymentThe FTC’s Rite Aid case+

Testing is an obligation
with human stakes.

In December 2023, the FTC alleged that Rite Aid’s facial-recognition system generated thousands of false-positive matches and that the retailer failed to adequately test accuracy before deployment. The agency described customers being wrongly accused and proposed a five-year surveillance facial-recognition ban as part of a settlement.

This is an account of the agency’s allegations and announced proposed settlement, not a claim that every allegation was adjudicated. Read the FTC announcement ↗

Our inference: testing the product includes examining the actions people take because of its output.

Open or closed?
Neither is an answer.

Openness and deployment control are separate questions. Open weights are not automatically open source; a closed API is not automatically safe.

Openly released weights

Can enable outside research, adaptation, and wider participation.

Copies can persist after release. Later restrictions may be difficult to enforce.

Provider-controlled access

Can support access limits, updates, and withdrawal of a hosted service.

Scrutiny and access depend more heavily on the provider’s discretion.

Conceptual international meeting chamber with an empty circular table and many chairs in soft daylight
A shared table. Terms still to be agreed.AI-generated conceptual illustration

“What if the other
side doesn’t stop?”

The strongest objection deserves a real answer. A US–China race framing puts reciprocal restraint under pressure. The consequences reach far beyond either country.

A government may see leadership as a source of security, prosperity, and influence. Slowing down alone could surrender those advantages without reducing a rival’s risks.

Racing can impose risks that cross borders. Shared requirements could make restraint more credible if participants can verify compliance and respond to evasion.

Trust carries the weight.

Shared language can establish intent, but participants have limited reassurance that a rival will absorb the same costs.

THE LIMIT

A declaration alone does not provide access to evidence or consequences for evasion.

A conceptual comparison, not a forecast of US or Chinese policy. Verification is incomplete; non-participation remains possible. Other countries and affected communities need a voice in setting the terms.

↳ What international oversight can build onExisting institutions & their limits+

Give the institution
a job description.

The IAEA’s safeguards rely on technical verification under agreements accepted by states. An AI institution would likewise need a basis for access, a defined remit, and legitimate authority.

The UN established an AI scientific panel and global dialogue in August 2025. Assessment and discussion are valuable functions; they do not themselves confer global enforcement powers.

This proposal leaves the institutional form open. The test is whether it can obtain evidence, act on it, resist pressure, and be held accountable itself.

Different systems.
Common obligations.

Six commitments for model developers, agent companies, data suppliers, and deployers. The requirements scale with their power and what they control.

01

Keep responsibility traceable.

+

Identify who operates the system, what each supplier controls, and where an affected person can seek a remedy. Delegation should preserve responsibility.

What evidence looks like

A named operator, records of material decisions, data provenance and permissions, and documented responsibilities at every handoff.

Where it can fail

Records become paperwork if no one has the authority or resources to investigate and repair harm.

02

Match power with evidence.

+

Require stronger evidence as capabilities, permissions, and potential consequences grow. Reassess material changes to the system and its environment.

What evidence looks like

A version-specific safety case, realistic tests, documented limitations, and explicit conditions that would invalidate the assessment.

Where it can fail

A benchmark can become a target that substitutes for evaluating the actual deployment. Testing costs can exclude smaller builders.

03

Make scrutiny independent.

+

Enable qualified outside evaluation with sufficient access to challenge material claims. Sensitive evidence can be inspected under controlled conditions.

What evidence looks like

Published assessment scope, disclosed conflicts, a route to report adverse findings, and public explanations of unresolved concerns.

Where it can fail

Evaluators can be captured, mistaken, or selected for favourable conclusions. Their independence needs scrutiny too.

04

Bound autonomy. Prepare for failure.

+

Enforce limits outside the model. Understand what can be interrupted, what can be repaired, and what cannot be recalled.

What evidence looks like

Enforced spending and access limits, expiring credentials, bounded delegation, recovery exercises, and release-specific containment plans.

Where it can fail

A stop button cannot undo every completed action. Publicly released artefacts may remain accessible after withdrawal.

05

Share what goes wrong.

+

Report serious incidents and meaningful near misses so others can act. Distinguish observed facts from hypotheses as investigations evolve.

What evidence looks like

Defined reporting thresholds, confidential disclosure channels, incident records, and later public accounts that protect sensitive information.

Where it can fail

Punishing honest reporting as harshly as concealment creates an incentive to stay silent.

06

Give affected people recourse.

+

People need to challenge consequential decisions and the legitimacy of a system’s use. A technically accurate system can still serve an unacceptable purpose.

What evidence looks like

Notice, understandable reasons, an appeal with an owner and deadline, authority to correct harm, and public participation before new uses are authorised.

Where it can fail

A review process is hollow if the reviewer cannot change the outcome. Governments must face scrutiny of their own deployments.

Where can we still
change what
happens next?

We do not need to agree on a single future to make responsibility concrete. We need to know who decides, what evidence they owe, and who can challenge them.

Go deeper into the proposal

An argument you
can inspect.

The original essay develops the obligations, the enforcement question, and the danger of putting the largest labs in charge. The visual chapters above extend it with real cases and policy comparisons.

Read the full essayAbout 14 minutes +

The organisations building the most powerful AI systems have a difficult conflict of interest. They are trying to establish whether their technology is safe while competing to make it more capable. The governments overseeing them have a version of the same problem: protecting the public sits alongside the ambition to lead an industry that could reshape economic and military power.

Both can sincerely want a good outcome. Neither can settle, on everyone else's behalf, how much risk everyone else should accept.

That is the starting point for Common Obligations: a proposal for what AI builders should owe the people affected by their systems, wherever those people live. I want a framework that a person can understand, an engineer can implement, and an independent body can examine. Its commitments should apply to foundation model developers, agent companies, data suppliers, and organisations putting AI to work, with requirements that reflect what each actually controls.

The central claim is simple: the freedom to build powerful AI should come with responsibilities to people who never chose to use it.

A race nobody can settle alone

In An Alien Mind, published on 6 September 2026, OpenAI's chief scientist Jakub Pachocki describes a tension directly relevant to this problem. He argues that pursuing recursive self-improvement is necessary to remain at the research frontier, while questioning whether accelerating that process is the right collective choice. He calls for shared safety requirements, external enforcement, and international coordination.

His assessment is a participant's view, partly grounded in internal results that readers cannot independently inspect. It deserves scrutiny alongside the proposal itself. But the conflict he describes does not depend on accepting his forecasts: a company can believe collective restraint would help and still fear the consequences of practising it alone.

Imagine two competitors evaluating a new capability. Both would prefer the other to spend another month testing. Neither wants to be the only one that does. An appeal to responsibility asks each to absorb a cost without knowing whether its competitor will reciprocate. Shared requirements could change that calculation, provided a broken commitment is detectable and has consequences.

The same problem survives at national scale. A government might support caution in principle while treating a rival's progress as a reason to accelerate. Some countries have much more influence over this decision than others. A country importing AI services can still bear the costs of failures originating elsewhere, with little access to the evidence behind deployment decisions.

We should be precise about what coordination would accomplish. It would make certain obligations harder to escape by switching providers or jurisdictions. It would also give cautious organisations a better answer to investors and customers asking why a competitor is moving faster.

The institution needs a job description

An international AI regulator is an appealing answer. The nuclear comparison helps, provided we understand what makes it work. The International Atomic Energy Agency's safeguards use technical verification under agreements accepted by states. They have defined objects of inspection and a basis for access. The existence of an international organisation, by itself, is insufficient.

AI presents a different verification problem. A model can be copied, adapted, and connected to tools after its original assessment. A system's power depends on its deployment conditions as well as its training. Inspecting the original model will not tell us everything about a service built around it, particularly when that service can take actions or delegate work.

There is already substantial work to build on. The OECD AI Principles address human rights, transparency, robustness, and accountability. NIST's AI Risk Management Framework offers a voluntary approach to managing risks throughout a system's life. The UN established an Independent International Scientific Panel on AI and a Global Dialogue on AI Governance in August 2025, providing mechanisms for assessment and discussion.

My proposal borrows from that work. Its contribution would be to connect a short public commitment to evidence and an accountable decision. Before designing an organisation's headquarters or voting structure, we should be able to say what it would ask a builder to demonstrate, and what would happen when the demonstration fails.

I would start with six obligations.

1. Keep responsibility traceable

Every consequential AI system should have an identifiable organisation responsible for its operation, and a clear account of the responsibilities held by its suppliers. The person affected should have somewhere to go when something fails.

Consider a hiring service assembled from a foundation model, a candidate database, and an agent that ranks applicants. The employer controls the decision to use the ranking. The agent provider controls its workflow. The data supplier controls the information it supplies and what it says about that information. The model developer controls the underlying release and its documented limitations. Those responsibilities overlap, but they need not disappear into the overlap.

Each participant should document the decisions it controls, preserve the evidence needed to investigate failures, and identify the organisation receiving that responsibility at the next handoff. For the data supplier, this includes origin, permission to use the data, known gaps, and a process for corrections. A claim that a dataset is representative needs an explanation of whom it represents and for which purpose.

This obligation does not make every supplier responsible for every downstream act. It does require a usable record of who knew what, who could change what, and who decided to proceed. An organisation deploying a system should remain answerable for that deployment even when its investigation later identifies a supplier's fault.

2. Match power with evidence

The more consequential a system's capabilities and permissions, the stronger the evidence required before expanding them. That evidence should identify the version tested, the conditions of the test, the failures observed, and the limits of what the results establish.

A model that suggests a database query and an agent that executes it against a hospital's records have different opportunities to cause harm. Adding credentials, persistent memory, external communication, or the ability to delegate can change the risk even if the model remains identical. The assessment should follow those changes.

A useful safety case is an argument someone else can inspect: here is what could go wrong, here is why we believe our controls address it, and here is the evidence that might prove us wrong. Passing a benchmark is one input. It cannot establish that an entire deployment is safe in every environment.

For the most powerful systems, this obligation should reach decisions about further development when development itself creates material risk. The relevant thresholds would need public justification and technical revision. I do not think we can responsibly invent a universal number in an essay and call everything below it safe.

The requirement must also be affordable in proportion to the risk. A small tool with narrow permissions should have a straightforward way to demonstrate compliance. A large company should face demanding requirements when the capabilities warrant them. Revenue and headcount are poor substitutes for examining what a system can do.

3. Make scrutiny independent

Builders should enable qualified outside scrutiny, with access proportionate to the claims and risks being examined. For higher-risk systems, the builder should not be able to make an unfavourable assessment disappear by choosing a different evaluator.

Independence needs an operating model. Who chooses the examiner? Who pays? Can the examiner run its own tests? Can it report a serious concern to an authorised oversight body without the client's permission? A company-funded audit can still be useful, but these questions determine how much confidence it deserves.

Public accountability does not require publishing private training records, security vulnerabilities, or every model artefact. Sensitive evidence can be examined through controlled access, with public summaries explaining the scope, findings, and unresolved concerns. Restrictions should have stated reasons and a route to challenge them, because commercial confidentiality can otherwise become a permanent answer to every difficult question.

Evaluators should also face scrutiny. Their methods, conflicts of interest, and material errors need examination. Independent assessment improves the basis for a decision; it cannot turn uncertainty into a guarantee.

4. Bound autonomy and prepare for failure

An AI system that acts should have explicit limits on its authority and tested ways to contain failure. Before deployment, the operator should know which actions require approval, which resources are accessible, how permissions expire, and what happens when the system behaves unexpectedly.

Suppose an agent can issue refunds. Its spending limit should be enforced by the payment service. A sentence asking the model to stay within budget is insufficient. If it delegates to another agent, the delegated authority should remain within the original permission. Stopping the parent should also address outstanding child work and credentials that could remain active.

Recovery deserves equal attention. If a payment request times out, the system should establish whether it committed before sending another. If an agent changes a record, the operator needs a way to identify and repair the change. A stop button may prevent the next action; it cannot unsend an email or reverse every action already taken.

Some releases are difficult to recall at all, including publicly distributed model weights. In those cases, the assessment must account for the limits of later intervention before release. Openness can support scrutiny and wider participation, but neither an open licence nor a closed API establishes safety on its own.

The general obligation is to demonstrate control appropriate to the system, and to be honest where control ends.

5. Report serious failures so others can learn

Builders and operators should report material incidents and significant near misses through a shared process. A finding that affects other systems should reach the people who can act on it, even when disclosure is commercially uncomfortable.

If an agent finds a way around an authorisation boundary, quietly patching one product may leave the same weakness in another. A useful report would distinguish what was observed from what is suspected, identify affected versions and conditions, and describe the containment measures. It should be possible to update the report as the investigation improves.

This requires agreed definitions of severity, reporting deadlines, and recipients. It also requires protection for personal information and security-sensitive details. A practical approach could combine prompt confidential notification to a competent body with a later public account of what happened and what changed.

The incentives matter. A process that punishes every honest report as if it were concealment will encourage silence. Accountability should distinguish responsible disclosure from repeated negligence, misleading claims, and deliberate suppression. Employees need a protected route to raise serious concerns when internal reporting fails.

6. Give affected people a way to challenge

People should be able to discover when AI materially influences a consequential decision about them, understand enough of its basis to contest it, and reach someone with the authority to correct it.

In the hiring example, an applicant needs a route to correct an erroneous record or challenge an unsuitable assessment. Sending them a technical explanation of a model does little if nobody can reconsider the decision. A meaningful appeal needs an owner, a timescale, and the power to provide a remedy.

This also applies to people whose data enters a system. Providers should explain the uses they make of it and offer workable processes for access, correction, and deletion where applicable. They should describe technical limitations honestly, including the difference between removing a source record and undoing its influence on an already trained model.

Human agency extends beyond individual appeals. Workers, affected communities, and countries using systems developed elsewhere should have a role in shaping the standards. Access to that discussion should not depend on owning a training cluster. Participation will take funding, translation, and technical support if it is to mean anything beyond an invitation to comment.

These six obligations fit together: identify who is responsible, require evidence, enable scrutiny, contain failure, share what goes wrong, and give people recourse. Their implementation will vary. Their purpose should remain recognisable across the supply chain.

What happens when a commitment fails?

A voluntary declaration can make a position visible. To change behaviour under competitive pressure, it needs consequences beyond embarrassment.

I would begin with a public record for a specific system and version. It would state which obligations apply, what evidence supports them, who examined that evidence, what remains unresolved, and when reassessment is due. Companies could adopt this structure before an international agreement exists. Buyers could then require the record in procurement, and contracts could specify access for review and the consequences of misleading claims.

For higher-risk systems, my preferred direction is enforceable requirements through public authorities, supported by independent technical evaluation. International agreements could establish common minimums and recognise assessments across participating jurisdictions. Domestic institutions would retain responsibilities for enforcement and remedies. That distributes the work while giving the shared requirements a basis beyond company promises.

An adverse finding should trigger a response proportionate to the danger: correction of a misleading claim, restriction of particular permissions, withdrawal of a deployment, or a pause on a defined development activity. The decision should explain the evidence, the scope of the restriction, and the conditions for resuming. Emergency powers need time limits and independent review, so a precautionary intervention cannot drift into an indefinite ban without justification.

Verification will remain incomplete. Some actors will refuse access, and some states may not participate. An international body would face those limits too. Recognising them helps define an honest initial goal: make compliance inspectable and consequential among participants, then expand participation and improve the ability to detect evasion.

That would be progress even before universal agreement. It would not justify claiming that the world's AI risks had been brought under control.

Who keeps the rule-makers accountable?

The strongest objection to this proposal is that it could help concentrate the power it is meant to oversee. Expensive assessments, proprietary tests, and licensing rules shaped by incumbents could make independent development harder while leaving the largest companies comfortable.

That possibility should shape the design from the beginning. Assessment methods should be open to examination, with multiple qualified evaluators and support for small organisations facing legitimate testing costs. Requirements should be justified against capabilities and deployment conditions. New evidence should be able to overturn them. A provider's business model, including whether it releases open models, should not automatically settle the assessment.

The institutions would need published funding, conflict disclosures, transparent appointments, and representation beyond the countries and companies with the most compute. People subject to their decisions should have an appeal route. Industry expertise is necessary, but a company's technical knowledge should not entitle it to decide the acceptable risk for everyone else.

There are limits to agreement as well. A shared framework will not settle every country's views on speech, every dispute over data rights, or every question about distributing AI's economic benefits. It should state its scope clearly while protecting a meaningful minimum. Cooperation that depends on ignoring the rights of people with the least political influence would fail its own purpose.

I am deliberately leaving the final institutional shape open. A treaty body, coordinated national authorities, and a network of accredited evaluators have different strengths and failure modes. We should compare them against the same practical questions: can they obtain the evidence, act on it, withstand pressure, and be held accountable themselves?

A proposal we can put to work

Common Obligations is an initial proposal. These principles still need sharper definitions, examples from actual deployments, and people willing to argue with the difficult parts. Several draw directly on existing governance work; the test is whether expressing them together makes responsibilities easier to understand and enforce.

A useful next step would be to apply them to three different systems: a foundation model release, an agent with authority to act, and a dataset supplied for consequential decisions. For each, we should be able to name the responsible parties, the evidence required, the independent checks, and the remedy when a commitment fails. Where we cannot, the framework needs more work.

I want powerful AI to be useful, widely available, and worth trusting. That requires a public say in the conditions under which it is built and deployed. We can start making those conditions concrete while the argument about who enforces them continues.

Sources & editorial approach +

Common Obligations is a proposal by Prassanna Ravishankar, not an established standard or certification. Scenarios illustrate possible mechanisms and tradeoffs; their effects are not simulated estimates. Illustrations are AI-generated and do not depict documented incidents.

Primary sources are linked at the claims they support. Provider statements and advocates’ accounts are identified as such. The full essay was drafted on 8 September 2026; this visual edition adds surveillance, release, and coordination comparisons.

Further foundations: OECD AI Principles · NIST AI Risk Management Framework · An Alien Mind · Open Source AI Definition.