# Task: write the Agent Functional Spec for this agent

You are a coding agent working inside the source repository of an AI agent. Write its **Agent Functional Spec** (UAS 1.0, a JSON file)
so that the Aptha Agent Analyzer accepts it without errors and can check every claim in it against this code.

You are given these files next to this prompt. Use them; do not rewrite them.

| File | What it is |
|---|---|
| `UAS_1.0_Template.json` | the template to fill in (bracketed text such as `[one sentence...]` is guidance, not a value) |
| `uas-1.0.schema.json` | the JSON Schema the file must satisfy |
| `check_uas_spec.py` | the checker you must run until it says READY (needs Python 3.9+ and `pip install jsonschema`) |

**Output:** one file, `<agent-id>.uas.json`, written to the repository root (or where the user tells you). **You are done only when**
`python check_uas_spec.py <agent-id>.uas.json --source .` prints `RESULT: READY FOR THE ANALYZER` with 0 errors, and every warning has been
either fixed or explained in your final message.

---

## 1. Rules that override everything else

1. **Describe what the code does, not what the README promises.** The analyzer compares each claim with the code. A claim the code does
   not support is reported as a defect in the spec, so a smaller true spec is better than a larger doubtful one.
2. **Never invent.** Every locator, number, name and system must come from code you actually read. If an optional field cannot be
   determined, leave it out. If a required field cannot be determined, write the most cautious true value and list it under
   "Needs human review" in your final message.
3. **Do not modify the agent's source code, do not run the agent, do not call the network.** Only read files and write the spec.
4. **No secrets.** Environment variables are listed by name only, never with a value. Do not copy keys, tokens, passwords or customer data.
5. **No claims you cannot show.** Do not state compliance, certification, SLAs, regions or owners that the repository does not state.
6. **Do not edit `check_uas_spec.py` or `uas-1.0.schema.json`.** If the checker complains, fix the spec.

## 2. Procedure

**Step 1. Understand the agent (read before you write).** Read the README, dependency manifests, `Dockerfile`/`Procfile`/`package.json`
scripts, and the entry points. Work out: how it starts; which framework runs it; where its decisions are made; every tool it calls; every
external system it reads or writes (databases, HTTP APIs, queues, files, email, LLM providers); which environment variables it reads.

**Step 1b. Sweep the whole repository, not only the files that hold decisions.** Search every source file for: environment reads
(`os.getenv`, `os.environ`, `process.env`, `System.getenv`), HTTP clients and URLs, database / vector-store / queue / cloud SDK imports, and
every file with a `__main__` block, CLI or web app. Every environment variable found must be in `execution.env`; every external system must be
in `interfaces.systems`. If the repository contains **more than one agent or program** (a second chat agent, a variant built on another
framework, a mock server), either describe the main one and say in `limitations` which others this spec does not cover, or describe them all.
Never silently describe only the file you happened to read first. The checker runs this same sweep and warns about what you missed.

**Step 2. Copy the template** to `<agent-id>.uas.json`. Delete the `_readme` key. Then fill the sections in this order, because later
sections refer to the ids of earlier ones: `agent` -> `interfaces.systems` -> `interfaces.consumes` / `produces` -> `capabilities` ->
`decisions` (with guardrails, formulas, escalation) -> `operations` -> `oversight` -> `lifecycle` -> `references` -> `execution` ->
`limitations`. Remove every optional block you cannot fill truthfully. Do not leave any bracketed placeholder text.

**Step 3. Validate and repair.** Run `python check_uas_spec.py <agent-id>.uas.json --source .`. Fix every `E` (error). Read every `W`
(warning): fix it, or keep it only if you can say why it is a true statement about this code. Repeat until READY.

**Step 4. Report** in the format of section 8.

## 3. How to fill each section

**`agent`**
- `id`: lowercase kebab-case from the repository or package name (letters, digits, hyphens, at most 64). `version`: from the package
  manifest, else `0.1.0`. `status`: `draft` unless the repository says otherwise. `summary`: one sentence, at most 200 characters.
- `categories`: 1 to 3 of `forecasting, demand-planning, supply-planning, inventory, replenishment, procurement, supplier-management,
  manufacturing, quality, warehouse, logistics, transportation, order-management, returns, finance, analytics, risk, compliance,
  sustainability, control-tower, copilot, orchestration` (or a custom `x-your-name`). Anything else is rejected ("purchasing" -> `procurement`).
- `kind`: `single` unless it is one of several cooperating agents (`orchestrator` needs a `composition` block; `member` is one of them).
- `owner.organization` from the repository owner or package author; `owner.contact` a reachable URI (the repository URL is fine), never a person's name.
- `context.excluded` needs at least one entry: something this code verifiably never does (for example "does not place orders with a real supplier").

**`interfaces.systems`** - one entry for every external system the code talks to (databases, ERP/WMS/TMS or other business APIs, message
brokers, email/Slack, the LLM provider). **Name each system the way the code names it** (client class, module or hostname words: "Ollama",
"ERP database", "Slack webhook"): the analyzer matches the code's calls to these names, and a call to something not listed is reported as an
undeclared external call. `role`: `system-of-record | execution | data-source | notification | model-service | other`. `direction`: `read |
write | read-write`, taken from what the code really does. `criticality`: `required | degradable | optional`.

**`interfaces.consumes` / `interfaces.produces`** - each distinct input the decisions read and each distinct output or effect they produce.
- `kind`: `api | event | file | database | stream | ui | document`. `system`: the id of a system above.
- `json_schema` is required and must describe the real payload. Derive it from the Pydantic model, dataclass, TypedDict, TS type or the
  dictionary keys the code actually uses (`{"type":"object","properties":{...},"required":[...]}`). A `//` key or an empty object is not acceptable.
- `semantics`: one sentence on what the data means. `classification`: `public | internal | confidential | restricted`. `pii`: `true` if the code
  handles names, emails, phone numbers, addresses or personal ids.
- Inputs also need `if_unavailable` (`halt | degrade | use-stale | substitute`), chosen from what the code does when the source fails.
  Outputs need `side_effect`: `true` if the output changes anything outside the agent (a write, a send, an order).
- `locator`: the function that reads or writes it.

**`capabilities`** - one per distinct thing the agent produces or does. `kind`: `sense | analyze | predict | recommend | execute | converse |
orchestrate`. `statement`: the business outcome (at most 300 characters). `locator`: the main function that implements it.

**`decisions`** - see sections 4 to 6. They are what the analyzer checks hardest.

**`operations`** - `failure_modes` from the code's real handling: `try/except`, retries, timeouts, validation errors, fallbacks.
Each has `condition`, `detection` (how the code notices), `response` (`halt | degrade | fallback | queue | escalate | retry-bounded`) and a
`safe_state` id. `safe_states` (at least one): what the agent does in that state, `entry` and `exit` as `automatic | manual`.

**`oversight`** - `posture`: `human-in-the-loop | human-on-the-loop | human-out-of-the-loop`. `roles` may stay empty if the repository names
no human roles. `controls.pause / override / rollback`: `available: true` only if the code or UI really provides it (add a `locator`),
otherwise `false`. `escalations`: at least one, saying who reads the agent's output when something goes wrong.

**`lifecycle.change_log`** - one entry: `version` = the agent version, today's date (`YYYY-MM-DD`), `type: ["added"]`, `breaking: false`,
`revalidation: "required"`, `summary: "Initial specification written from a reading of the source."`. If the code calls an LLM, add
`lifecycle.model_dependencies` (`id`, `role`, `update_policy`: `pinned | managed-upgrade | provider-continuous`, `revalidation`: `full | affected | none`).

**`references`** - one `repository` entry with the real location (the git remote URL, or the path).

**`execution`** (the analyzer uses this to set the agent up in a sandbox, so it must be exact)
- `candidate_ref`: the exact command that starts the agent, written **from the repository root**, copied from the README, `Procfile`,
  `Dockerfile` or `package.json` (for example `python -m app.main`, `uvicorn backend.main:app`, `npm run start`). It must resolve to a
  file or script that exists; the checker verifies this.
- `invocation_kind`: `cli | http | grpc | function | queue-consumer`. `sandbox`: `false` unless the agent can safely run offline against mocks.
  `timeout_s`: `300` unless the code says otherwise.
- `env`: every environment variable the code reads: `{"name": "OPENAI_API_KEY", "secret": true}`. `secret` is `true` for keys, tokens, passwords
  and connection strings. **Names only.**

**`limitations`** (optional but encouraged) - what this spec does not cover (other screens, other tools, behaviour you could not trace).

**Leave out** `scenarios`, `acceptance_criteria`, `failure_mode_tests`, `combinatorial` and `governance` unless the repository itself contains
that information. Do not pad the document; every block you add must be true.

## 4. Decisions

A decision is a place where the code **chooses between outcomes that matter to the business** (approve or reject, order or hold, route,
escalate, accept or refuse). It is not every function. Most agents have 1 to 6 decisions.

- `locator` = the **function or method** that makes the choice, or its top-level caller. The analyzer reads the code reachable from that
  function (calls up to four levels deep, 40 functions), so choose the function from which the limit check and the arithmetic can be reached.
  A class name alone, a module variable, a constant or a lambda cannot be followed. Use `Class.method` for methods. For a TypeScript/JavaScript
  tool written as an object (`const issueCredit = { execute: (input) => {...} }`) the symbol is `issueCredit.execute`; the bare object
  name does not resolve.
- **If the choice is made by a language model following prompt text** (for example a CrewAI/LangGraph agent whose "decide" step is a task
  description) there is no function to point at. Declare the decision with **no locator**, say so in its `statement`, and put a line in
  `limitations` ("the scoring exists only in the agent's prompt, not in code"). The checker will warn about the missing locator; that warning is
  expected and true. Do **not** point the locator at an unrelated function to silence it, and do not write formulas for arithmetic that exists only in prompt text.
- If a limit or branch lives in a script's `if __name__ == "__main__":` block (not a function), it cannot be located either: mention it in `limitations`.
- `trigger.kind`: `event | schedule | threshold | request`, and `trigger.detail` says what starts it.
- `inputs`: ids of `interfaces.consumes` it reads. `executes_via`: ids of `interfaces.produces` it writes (required if it has a side effect).
- `reversibility`: `reversible` (no external effect, or the same code undoes it), `compensable` (undone by another action), `irreversible`
  (a sent email, a placed order, a payment). If it is not `reversible`, add a `reversal` sentence.

### Authority: declare what the code does, no more and no less

Look at what the decision's code path actually does to the outside world.

| If the code path... | `authority` |
|---|---|
| only reads, analyses, logs | `observe` |
| produces a proposal that a person or another system acts on; writes nothing external | `recommend` |
| performs the action only after an explicit human approval step | `act-with-approval` |
| performs the action itself, while a person is alerted or can intervene | `act-with-oversight` |
| performs the action itself with no human step | `act-autonomous` |

**Do not split an agent into one `act-*` decision per tool.** If a tool or helper acts but has no limit of its own, the schema would force
you to invent a guardrail for it. Describe such actions inside the decision of the loop or orchestrator that calls them (list their outputs in
that decision's `executes_via`), and give separate decisions only to the actions that have their own real limit (a credit cap, a quantity cap, an approval threshold).

The analyzer reports an `observe`/`recommend` decision whose reachable code writes to an external system (HTTP POST/PUT/PATCH/DELETE,
database INSERT/UPDATE/DELETE/commit, sending a message or email, writing or deleting a file). Never choose a lower authority to look safe or a
higher one to look capable. `act-with-oversight` and `act-autonomous` **require at least one guardrail and an `escalation`** (`to`, `when`).

## 5. Guardrails: only limits the code really enforces

A guardrail states the **allowed** condition. Read the code that rejects, and invert it:

| The code rejects when | `operator` |
|---|---|
| `x > LIMIT` | `lte` |
| `x >= LIMIT` | `lt` |
| `x < LIMIT` | `gte` |
| `x <= LIMIT` | `gt` |
| a loop `range(N)` / `max_iterations=N` / `retry(max_attempts=N)` | `lte` with threshold N |
| a loop that keeps going `while (turns < MAX)` / `for (i = 0; i < MAX; i++)`, where the counter counts completed rounds | `lt` with threshold MAX (parameter: the counter, e.g. "iterations completed") |

- `threshold` is **the number the code compares against**, exactly (`25000`, not "about 25k"). If it is a constant defined elsewhere, use its
  value. If it is the default of an environment variable or config key, use that default. If you cannot find the number, **leave the guardrail
  out** and describe the limit in `limitations`. The checker warns when a threshold appears nowhere in the source.
- `parameter` is the quantity **in the words the code uses** (`MAX_ORDER_VALUE` -> "max order value"; `zscore` -> "z-score"). The analyzer
  finds the enforcing comparison by these words and by the number.
- `on_breach`: what the code does when the limit is crossed: `block` (stops or rejects), `require-approval` (routes to a person),
  `escalate` (notifies), `degrade` (continues with a fallback).
- `locator`: the function that contains the comparison (often a helper the decision calls).
- Non-numeric rules (`status in ["pending","approved"]`) may use `eq` / `in` with a string or array. The analyzer will say it cannot verify
  them statically; that is expected.
- A guardrail is a **numeric (or fixed-value) comparison** the code makes, or a loop/retry/size cap. A rule enforced only by a lookup
  ("do not open a second claim if one exists", "skip if already processed") is not a limit the analyzer can verify: describe it in a
  `failure_modes` entry or in `limitations`, not as a guardrail.
- If two guardrails bound the same quantity with different numbers (a maximum and a minimum), give them clearly different `parameter` wording
  ("credit amount" and "credit amount minimum"). The analyzer matches by wording and number and may report the lesser one as unverifiable.
- Do not add a guardrail because the README says there is one. If the code does not enforce it, it is a gap to mention in your report, not a claim.
- If an autonomous decision truly has no business limit, use the weakest real bound that exists (a loop cap, timeout, batch size) and say in
  `limitations` that no business limit exists. Never make up a number to satisfy the schema.

## 6. Formulas: only arithmetic the code really performs

Add `formulas` to a decision only when its code computes a business quantity with straight-line arithmetic (a reorder point, a score, a price).
- `output` is the variable the code assigns or returns (`reorder_point`). `variables[].name` are the names the code uses (use `binds_to` if the
  spec name differs from the code name).
- `expression` is Python arithmetic over those names, at most 300 characters, using only `+ - * / // **`, numbers and `max min abs ceil floor
  sqrt round`. **No** comparisons, `if/else`, `%`, strings, keyword arguments, loops or function calls into your own code.
- Copy the arithmetic as the code has it, including terms you consider unimportant. If the computation involves loops, branches or data lookups,
  **omit the formula**; do not simplify it into something the code does not compute.

## 7. Mistakes that make a spec fail or lose accuracy

- **Duplicate ids.** Every id must be unique across the **whole** document (capabilities, decisions, guardrails, interfaces, systems,
  policies, failure modes, safe states, roles, risks). A repeated id stops the analyzer. Ids are lowercase kebab-case, start with a letter,
  and are at most 64 characters. Conventional prefixes: `cap- dec- gr- f- in- out- sys- fm- ss- pol- risk- role- esc- ref- lim- exc- asm-`.
- **References to nothing.** Every id you reference (`capability`, `inputs`, `executes_via`, `system`, `safe_state`, ...) must exist.
- **Locator paths.** Relative to the repository root, forward slashes, exact letter case, no leading `./`, no spaces. `symbol` is a function or
  method that exists. The checker verifies each one.
- **Template leftovers.** `_readme`, bracketed text, `//` keys.
- **Wrong enum values.** Use exactly the spellings in this prompt and the schema. `orchestrate` is a capability kind, not an authority;
  `alert` is not an `on_breach` value (use `escalate`).
- **Length limits:** `agent.summary` 200, names 80, every `statement` (including `limitations` and `context.excluded`) 300, `guardrail.parameter` 120, `reversal` 200. Split a long thought into two entries.
- **Describing only part of the repository** without saying so (see Step 1b).
- **A start command that does not exist** or is written for another working directory.

## 8. Final message (after the checker prints READY)

Reply with:
1. The path of the file and the last lines of the checker output.
2. A table of the decisions you declared: id, authority, locator, guardrails (with the file and line where the limit is enforced).
3. **Needs human review:** every value you were unsure of, every warning you kept and why, anything the README claims that the code does not do.
4. **Deliberately left out:** blocks or fields you omitted because the code does not support them.
