PermDock
Adapters

OpenAI Agents SDK

permdock/openai turns PermDock decisions into OpenAI Agents SDK needsApproval predicates, guards tool lists per caller, resolves interruptions against a pluggable ApprovalStore, and binds the replay-safe token to the serialised RunState.

permdock/openai connects the Decision model to the OpenAI Agents SDK for JavaScript. The SDK pauses a run when a tool needs approval, returns interruptions, and resumes from the same RunState after state.approve() or state.reject(). What it does not do is decide whether a call needs approval from anything richer than a boolean, nor record who approved, nor guarantee that the approved call is the one that runs. The adapter supplies the predicate from the typed policy, the record through an ApprovalStore, and the binding through the token.

Purpose

The SDK's human-in-the-loop guide defines needsApproval as true or an async function returning a boolean, evaluated after the tool arguments parse; malformed arguments fail closed by requesting approval without calling the predicate. Pending approvals surface as result.interruptions, each resolvable with approve or reject (optionally { message } for the model), and the run resumes with runner.run(agent, state). RunState serialises with toString() and restores with RunState.fromString(agent, s), so approvals can wait for days. Hosted MCP tools have the parallel requireApproval and onApproval. All of this is state plumbing; the decision is the application's. permdock/openai is that decision.

API

import { Agent, run, tool, RunState } from "@openai/agents";
import { createPermDock } from "permdock/openai";
import { z } from "zod";

const PostArgs = z.object({ id: z.string() }); // tool input arrives as `unknown`

const { needsApproval, guardTools, resolveInterruptions, permdock } =
  createPermDock(policy, {
    subject: (context) => context.user, // RunContext -> principal
    actor: (context) => ({ id: context.agentId, kind: "openai-agent" }),
    delegation: (context) => ({ scopes: context.scopes }),
    tools: {
      delete_post: {
        permission: permissions.post.delete,
        data: (args) => loadPost(PostArgs.parse(args).id),
      },
      list_posts: { permission: permissions.post.list },
    },
    store, // ApprovalStore; memoryApprovalStore() when omitted
  });

const deletePost = tool({
  name: "delete_post",
  parameters: z.object({ id: z.string() }),
  needsApproval: needsApproval(permissions.post.delete), // decide() per call; denied → approval never granted
  execute: async ({ id }, ctx) => {
    (await permdock(ctx)).assert(permissions.post.delete, await loadPost(id));
    return remove(id);
  },
});

const agent = new Agent({
  name: "Posts",
  tools: guardTools([deletePost, listPosts], context),
});
// After a run pauses
let result = await run(agent, input, { context });
if (result.interruptions.length) {
  const pending = await resolveInterruptions(
    result.state,
    result.interruptions,
    { context },
  );
  // pending: ApprovalRequest[] written to the store; surface them to a person
  await db.save(runId, result.state.toString()); // RunState; tokens live in the store, not in the state
}

// Later, in another process: the context comes from the resuming session, not from the stored state
const context = { user: await currentUser(request) };
const state = await RunState.fromStringWithContext(
  agent,
  await db.load(runId),
  new RunContext(context),
);
const pending = await resolveInterruptions(state, state.getInterruptions(), {
  context,
});
if (pending.length === 0) result = await run(agent, state); // every interruption approved or rejected
  • needsApproval(permission) returns the SDK predicate (runContext, args) => Promise<boolean>; it reads the application context from runContext.context. It returns false only when can(permission, row) holds for the resolved resource, and true (pause) for approval-required, denied, a thrown or empty data loader and anything else; the check does not touch the store. The tool never executes on a denial: resolveInterruptions rejects it.
  • guardTools(tools, context) filters the tool array to those whose permission has any grant for this subject, so the model cannot plan with tools it may not use (the capabilityMiddleware equivalent).
  • resolveInterruptions(state, interruptions, { context }) decides each interruption in order. granted (including an approved record, which it consumes) calls state.approve(i); denied (including a rejected, consumed or expired record) calls state.reject(i, { message }) with the Decision's reason; approval-required writes or finds the pending ApprovalRequest and returns it. An interruption whose name or JSON arguments cannot be read is rejected. The token is recomputed from permission, resource, principal, actor and arguments, so any process with the same store finds the record.
  • permdock(context) returns a request-scoped PermDock for checks inside execute.
  • subject and actor read from the RunContext the application passes to run; nothing is read from model output. delegation reads from the same context and has no default: an actor with no delegation is denied every tool with reason no-delegation.

Request lifecycle

  1. Before the run, guardTools removes tools the subject has no grant for.
  2. The model calls a tool. The SDK parses the arguments; on parse failure it requests approval without calling the predicate (its own fail-closed rule). Otherwise it calls needsApproval(context, args).
  3. The adapter validates args against the resource schema when data is declared, loads the resource, and calls decide:
Decision outcomeneedsApproval returnsThen
grantedfalseTool executes
approval-requiredtrueInterruption created; ApprovalRequest with token written to the store on resolveInterruptions
deniedtrueresolveInterruptions immediately calls state.reject(i, { message }) with the Decision's reason and alternatives
unmapped tool, validation error, thrown resolvertrue then rejectFail closed
  1. The application persists result.state.toString() and shows the pending requests (its own UI, approvalsHandler, or the PermDock Cloud inbox). The token lives only in the store; it is not placed inside RunState, because serialised state travels with the request and may be logged.
  2. A person resolves the request; the store records the approver. The actor (the agent) cannot approve its own call, and the principal cannot either unless the grant sets approval: { distinct: false }.
  3. The application restores RunState with a fresh context, calls resolveInterruptions, and resumes with run(agent, state) once nothing is pending. The adapter re-runs decide and recomputes token: an approved record resumes one call and is consumed, so restoring the same state again rejects the call; another principal's resume never matches the token.
  4. Every step emits on('decision'); ask, answer and resumed decision share token.

Sticky decisions (alwaysApprove, alwaysReject) are never issued by the adapter: each call gets its own decision and its own record. A denial returns true from needsApproval rather than throwing in execute, so the SDK's approval item carries the rejection message and the tool is never invoked. Tools of an inner agent nested with agent.asTool() interrupt the outer run and are decided with the outer run's subject and actor.

What it validates

  • Arguments against the resource's Standard Schema when data exists; the SDK's own parse-failure path already fails closed for malformed JSON, and the adapter's validation covers well-formed but wrong data.
  • The subject comes from RunContext, never from arguments or interruption.rawItem.
  • Approvals on resume: token recomputed and compared; unknown, pending, rejected, expired or already consumed records reject the call. An approval resumes one call.
  • Tool coverage: in development, guardTools warns about tools missing from the tools map, because needsApproval will reject them at runtime.
  • Serialised state: RunState is treated as opaque; the adapter reads callId and tool name from interruptions and nothing else. runContext.context is persisted data per the SDK's own warning, so subject and actor should be re-derived from authentication on resume, not trusted from the deserialised context; RunState.fromStringWithContext(agent, s, freshContext) is the recommended path.

How denials surface

  • state.reject(i, { message }) with message built from the Decision: failing roles, reason kind and alternatives, phrased for the model. The SDK sends it back as the tool result so the model can pick a permitted action.
  • guardTools removals are silent to the model; they appear in on('decision') with source: 'adapter'.
  • Resume failures use message values approval-expired, approval-mismatch, approval-rejected and approval-not-found.
  • Hosted MCP tools reached through the SDK's requireApproval and onApproval can call needsApproval inside onApproval; the mapping is the same, with the MCP server name and tool name as the lookup key in tools.
  • The SDK's tool guardrails (toolInputGuardrails, toolOutputGuardrails) are a second hook: an input guardrail can call decide and return a tripwire with the same Decision-derived message, which is useful when a team already routes all argument checks through guardrails. needsApproval remains the primary hook because it is the only one that can pause for approval-required; a guardrail can only allow or reject. OpenAI's Agent Builder is being retired on 30 November 2026 and does not affect this adapter, which targets the SDK (landscape).

Why

  • An inner asTool() agent inherits the outer subject and actor. The nested agent runs inside the outer run, on the outer run's context, for the same person. Giving it an actor of its own would need a second delegation to deny against, which the application has no verified source for; inheriting keeps the outer delegation as the ceiling, so nesting an agent can never widen what the run may do.
  • One approval per computer-tool action, never a batch. Each click, keystroke or navigation is a separate tool call with its own arguments, so each gets its own decision and its own token. A batch approval would bind a token to a sequence the reviewer cannot see in advance, and the screen state that justified the first action may not hold for the fifth.
  • No Python client. The adapter is the TypeScript SDK integration. A Python agent reaches the same policy through the AuthZEN endpoint over HTTP, so there is one wire contract rather than a second client to keep in step.

Example app

apps/examples/openai-agent: HTTP harness on 127.0.0.1:3475 with GET /health. GET /list_posts calls needsApproval and returns false (run). GET /delete_post returns true (pause for approval). No OpenAI API key. tests/integration runs a real Runner with a scripted model against a Postgres ApprovalStore: the run interrupts, RunState.toString() is resumed in a child process after an owner approves, the approved call runs once, and replaying the same state or resuming as another user runs nothing.

Last updated on

On this page