Agents and LLM backends¶
TriageSim separates who is speaking (the agent, which owns a persona and a prompt) from what generates the text (the LLM backend). This page covers both layers and the schemas that connect them.
The two-layer design¶
flowchart LR
P[Persona] --> A[NurseAgent / PatientAgent]
H[Transcript] --> A
A -->|prompt| B[BaseLLM]
B -->|validated Pydantic model| A
A -->|NurseOutput / PatientOutput| E[TriageEnv]An agent builds a prompt from its persona, the transcript, and the triage protocol. The backend turns that prompt into a schema-validated object, not a string. Nothing downstream ever parses free text.
Output schemas¶
Both schemas inherit from BaseAgentOutput, which forbids extra fields. If the
model invents a key, validation fails rather than the key being quietly
accepted; the comment in the source calls this out as critical, and it is the
main defence against a model smuggling unsanctioned state into the run.
PatientOutput¶
| Field | Type | Meaning |
|---|---|---|
utterance | str | What the patient says this turn. |
The patient has exactly one thing it can do: speak.
NurseOutput¶
| Field | Type | Meaning |
|---|---|---|
action | "utterance" \| "check_vital" \| "log_red_flag" \| "end" | The next action the nurse chooses. |
utterance | str \| None | The nurse's line, or the request naming a vital. Null for end and log_red_flag. |
triage | int, 1-5 | Current predicted triage level. Required on every output, not just at the end. |
confidence | "low" \| "medium" \| "high" | Confidence in that triage level. |
red_flags | list[str] | All red flags identified so far. Defaults to empty. |
explanation | str | Clinical reasoning grounded in the selected triage algorithm. |
Because triage, confidence and explanation are emitted at every step, the
trace records how the decision evolved, including the
turn at which the nurse first arrived at the correct level.
utterance carries the vital request
There is no separate "which vital" field. For a check_vital action, the
environment infers the vital from the wording of utterance, so
"Let me take your blood pressure" resolves to sbp. See
Vitals.
BaseLLM¶
The backend interface is a single abstract method:
from abc import ABC
from typing import List, Optional, Union
class BaseLLM(ABC):
def generate(
self,
prompt: str,
max_tokens: Optional[int] = None,
stop: Optional[List[str]] = None,
**kwargs,
) -> Union[str, BaseAgentOutput]:
...
Anything implementing generate can drive an agent.
OpenRouterLLM¶
The bundled backend routes through OpenRouter using
pydantic-ai:
from triagesim.agents import OpenRouterLLM
from triagesim.core import NurseOutput, PatientOutput
nurse_llm = OpenRouterLLM(model_name="anthropic/claude-sonnet-4-5",
output_type=NurseOutput)
patient_llm = OpenRouterLLM(model_name="google/gemini-3-pro-preview",
output_type=PatientOutput)
| Parameter | Type | Meaning |
|---|---|---|
model_name | str | OpenRouter model identifier. |
output_type | type[BaseAgentOutput] | Pydantic model the agent is forced to produce. |
**agent_kwargs | Forwarded to the underlying pydantic_ai.Agent. |
The API key is read from triagesim.config.OPENROUTER_API_KEY, not passed in.
See Configuration.
One backend per agent. The output schema is bound at construction, so a nurse backend physically cannot serve a patient agent. This turns a whole class of wiring mistakes into an obvious two-line setup instead of a confusing validation error twenty turns into a run.
Structured output is required
OpenRouterLLM requests a low reasoning effort and constrains generation to
the given schema. Models without tool-calling or JSON-mode support will fail
validation. Check the model's OpenRouter listing before using it.
PatientAgent¶
from triagesim.agents import PatientAgent
patient = PatientAgent(llm=patient_llm, persona=patient_persona)
| Parameter | Type |
|---|---|
llm | BaseLLM bound to PatientOutput |
persona | PatientPersona |
Its single method is:
output = patient.act(
history="Nurse: What brings you in today?",
chief_complaint="Syncope",
pain=7,
)
print(output.utterance)
The patient sees only the transcript, its chief complaint, and its pain score. It never sees vitals or the true acuity; it can only report what a person in that situation would plausibly know and choose to say.
NurseAgent¶
from triagesim.agents import NurseAgent
nurse = NurseAgent(llm=nurse_llm, persona=nurse_persona, algorithm="esi")
| Parameter | Type | Meaning |
|---|---|---|
llm | BaseLLM bound to NurseOutput | Generation backend. |
persona | NursePersona | Conditioning traits. |
algorithm | str | "esi" or "ats". Selects the protocol text embedded in the prompt. |
act¶
known_vitals is the set of vitals already released. The agent uses it to build
the list of actions still available, so the nurse is not offered a vital it has
already seen.
infer_belief_updates¶
After each patient utterance the runner asks the nurse to extract structured belief updates from what was just said:
updates = nurse.infer_belief_updates(
history=transcript,
last_utterance="I went dizzy and the next thing I knew I was on the floor.",
turn=3,
)
These are merged into the environment's belief graph. This is a separate call
from act; reasoning about what was learned is deliberately not entangled with
deciding what to do next.
Custom backends¶
To use a provider other than OpenRouter, subclass BaseLLM:
from typing import List, Optional, Union
from triagesim.agents import BaseLLM
from triagesim.core import BaseAgentOutput, NurseOutput
class MyLLM(BaseLLM):
def __init__(self, output_type: type[BaseAgentOutput]):
self.output_type = output_type
def generate(
self,
prompt: str,
max_tokens: Optional[int] = None,
stop: Optional[List[str]] = None,
**kwargs,
) -> Union[str, BaseAgentOutput]:
raw = my_provider_call(prompt, max_tokens=max_tokens, stop=stop)
return self.output_type.model_validate_json(raw)
nurse_llm = MyLLM(output_type=NurseOutput)
Two requirements:
generatemust return an instance of the bound output type, not a string, or the environment will reject it with aTypeError.- Your provider must be able to honour the schema. Prompting for JSON and
validating with
model_validate_jsonworks, but budget for retries: an unconstrained model will occasionally emit prose.
Next: Running a simulation.