ADR 0008 — FastMCP server mounted at /mcp, Azure AD token-validated, all-tools¶
- Status: Accepted
- Date: 2026-06-29
- Deciders: @TheurgicDuke771
- Related: ADR 0003 (unified suite/check/result model the tools read), 0010 / 0013 (generic
get_current_user; no Azure claim-reading in business logic), CLAUDE.md §10 (MCP tool descriptions are LLM-facing)
Context¶
Week 7 calls for a FastMCP server exposing 8 curated tools at /mcp, reachable from Claude Desktop / Claude.ai / Copilot / Cursor. Three design questions had to be settled against the installed library (fastmcp v3, not the v2 API the roadmap snippet assumed):
- How to mount into the existing FastAPI app.
- How to authenticate — reusing the same Azure AD bearer token the web UI already carries, not a second login.
- Tools vs resources for the 4 read operations the roadmap labelled "resource".
Decision¶
Mount — mcp.http_app(path="/") returns an ASGI app mounted at /mcp. Its streamable-http session manager needs its lifespan run, so the app's own startup is combined with it via combine_lifespans(lifespan, mcp_app.lifespan) (fastmcp's documented FastAPI pattern). The roadmap's get_asgi_app() is a stale v2 name.
Auth — a fastmcp JWTVerifier configured from the same tenant / audience / scope as core.auth: Azure JWKS (https://login.microsoftonline.com/{tenant}/discovery/v2.0/keys), issuer (…/v2.0), audience = the API app's client id, required scope = azure_api_scope. This validates the identical token the REST API accepts, without depending on fastapi-azure-auth internals (a Starlette-request-bound dependency that can't verify a raw token string). The full OAuth AzureProvider was rejected — it drives an authorization-code flow and needs a client secret; clients already hold a token. Inside a tool the validated claims (oid, preferred_username, name) resolve + upsert the User via the shared core.auth._upsert_user, so the row is identical to a web-UI login and queries scope by the generic user id (no Azure claim read in service code — ADR 0010/0013).
Two modes + fail-closed — real Azure mode uses the JWTVerifier; local dev-bypass (ENVIRONMENT=dev + AUTH_DEV_BYPASS=true, no Azure vars) mounts unauthenticated and resolves the fixed dev user, exactly like the REST API. If neither is configured the server is not mounted at all — /mcp never goes live without auth (CLAUDE.md §10 security note).
All 8 are MCP tools, not resources — despite the roadmap labelling 4 as "resource". An LLM client invokes tools from natural language; fastmcp resource-templates with required arguments aren't reliably auto-called. The acceptance bar is "Claude answers the canonical NL queries" (what failed today? / run the orders suite on DEV / why did the customer pipeline fail? / add a null check on email), which is best served by tools. Read tools (list_suites, get_suite_results, get_health_score, get_adf_pipeline_status) + action tools (trigger_suite_run, get_run_status, create_check, profile_column).
No logic duplication — each tool is a thin wrapper: open a session → resolve the caller → call the same service function with the same require_permission / accessible_suite_ids authz the REST routers use → return an LLM-shaped dict. get_suite_results reuses run_service.redact_sample_failures so failing-row PII is masked exactly as in the REST results path.
Consequences¶
- The MCP surface inherits per-suite sharing, existence-hiding, and sample redaction for free — there is one authz + redaction implementation, not two.
- The
JWTVerifieraudience assumes a v2 token whoseaudis the API client id (the single-tenant configcore.authuses); the deferred "test end-to-end with Claude Desktop" task validates this against a live token, since it needs the deployed tenant. - Tool bodies remain plain, directly-callable functions (the
@mcp.tooldecorator returns the function unchanged), so they're unit-tested by calling them with a test session + a stub user — no MCP transport needed. - The roadmap's resource/tool split is superseded here; the progress ledger's "Resource: X" items are delivered as tools.
Alternatives considered¶
AzureProvider(full OAuth) — rejected: needs a client secret and an auth-code/redirect flow; clients already present a token.- Bridge to
fastapi-azure-authby faking a StarletteRequest— rejected as brittle coupling to that library's request-bound internals;JWTVerifieris the clean, documented path and validates the same token. - Resources for the reads — rejected for LLM invocability (above); revisit if a client surfaces resources usefully.
Amendment — Tier 1 expansion to 19 tools (2026-08-17)¶
The original 8 tools were deliberately the smallest set that answered the roadmap's canonical
NL queries. context/post-v1-roadmap.md Theme 13 catalogued
the rest of the REST surface as MCP candidates, tiered by risk. This amendment ships Tier 1 —
high-value safe reads in full: list_checks, get_check, get_check_history, list_runs,
get_run_results, list_connections, list_schedules, list_trigger_bindings,
get_notification_config, get_suite_performance, export_suite — bringing the server to
19 tools: 17 read-only, 2 mutating (the original trigger_suite_run and create_check).
The decision above is unchanged, not revisited. Every new tool is the same thin wrapper this ADR already describes — open a session, resolve the caller, call the same service function, return an LLM-shaped dict. No new authz path was introduced; each tool reuses whichever gate its REST counterpart already applies, which is deliberately not uniform:
- Suite-scoped reads (
list_checks,get_check,get_check_history,get_run_results,get_notification_config,export_suite) gate onsuite_authz.require_permission(minimum="view"), so ADR 0027 per-suite sharing and the ADR 0033 Viewer clamp apply exactly as on the original 8. - Suite-scoped lists (
list_runs,list_schedules,list_trigger_bindings) scope throughsuite_service.accessible_suite_ids— with the workspace-admin view where their REST route has it — and additionally callrequire_permissionup front when the optionalsuite_idis given, so naming a suite you cannot see is an error rather than an empty list. get_suite_performanceis scoped by the dashboard's own accessible-suite subquery.list_connectionscalls neither, because connections are workspace-scoped rather than suite-scoped — the same rule its REST route follows. That is why what it returns is constrained instead (below); the ADR 0033 role axis gates connection mutations, and MCP deliberately has none.
Three standing exclusions carried forward, made explicit because Tier 1 sits right next to them:
list_connectionsis workspace-scoped and returns metadata + health only — id, name, type, env, whether a credential is stored, and health signals. It never returns a connection's configuration (account identifiers, hosts, paths) or a secret reference.get_notification_configreports channel presence, never webhook URLs. A webhook URL is itself a bearer credential; the tool answers "is Teams/Slack/email wired up, and from where" without ever resolving the secret.- Connection create/update/reauth remain excluded — a credential must never transit an LLM. This was true before Tier 1 and stays true after it; no mutating connection tool exists.
Tier 2 (mutating, edit-permission-gated: dryrun_check, update_check/delete_check,
snooze_check, cancel_run, schedule/trigger-binding CRUD, import_suite, test_connection,
etc. — see the roadmap's Theme 13 table) stays deferred; nothing in this amendment ships it.
Amendment — Tier 2 expansion to 30 tools¶
This amendment ships the Tier 2 set deferred above, in full: update_check, delete_check,
snooze_check, dryrun_check, cancel_run, create_schedule, delete_schedule,
create_trigger_binding, suggest_column_policy, test_connection, import_suite — 11 new
tools, bringing the server to 30 tools total. Theme 13 is now fully delivered.
The decision above is still unchanged. Every Tier 2 tool is the same thin wrapper — open a session, resolve the caller, call the same service function, return an LLM-shaped dict — reusing whichever gate its REST counterpart already applies.
The 30 tools split three ways, and the split is deliberately not "read vs mutate":
- 16 read-only — one fewer than the 17 the Tier 1 amendment above states.
profile_columnwas reclassified out of read-only when the gate table was built: it persists nothing, but it opens a live datasource with stored credentials and had always gated onedit. The behaviour did not change; the label was wrong, and a comment could not catch it. - 10 that change state —
create_check,update_check,delete_check,snooze_check,trigger_suite_run,cancel_run,create_schedule,delete_schedule,create_trigger_binding,import_suite. All exceptimport_suitegate onsuite_authz.require_permission(minimum="edit") against the suite they act on, so per-suite sharing (ADR 0027) and the ADR 0033 Viewer read-only clamp apply exactly as on every existing mutating tool.import_suiteis the exception: it creates a suite, so there is no existing resource whose ladder could gate it — it takes the coarserole:membergate below. - 4 that persist nothing but open a live datasource connection using stored credentials —
profile_column,dryrun_check,suggest_column_policy,test_connection. None of these write a row, but all four spend a real credential against a remote system, which is not a read-only action even though nothing is saved.profile_column,dryrun_checkandsuggest_column_policyare suite-scoped and gate onrequire_permission(minimum="edit")like a write.test_connectionhas no suite to gate on at all — a connection is workspace-scoped, not suite-scoped — so it gates on the coarse axis instead:server._require_role(user, "member"), the MCP-side twin ofcore.auth.require_role(ADR 0033), asserting the caller holds at least thememberworkspace role.import_suitegates the same way, for the symmetric reason: creating a suite has no existing suite to check permission against either.
MCP exposes no admin-only tool at all. Every Admin-only capability in ADR 0033's
authorization matrix is a connection mutation — create, edit, delete, re-auth — and none of
those are exposed here, before or after this amendment: a credential must never transit an LLM.
test_connection is the closest any tool comes to touching a connection, and it deliberately
stops at reporting whether the live probe succeeded; it never returns a credential or a secret
reference, matching the standing exclusions the Tier 1 amendment already recorded for
list_connections and get_notification_config.
The three standing exclusions from the Tier 1 amendment are unaffected and still hold in full.
Amendment — Tier 3A coherence tools, 30 → 33 (2026-08-17)¶
Three tools that close dead-ends in the existing surface rather than adding reach:
update_suite, get_column_policy, set_column_policy. The split becomes
33 tools: 17 read-only, 12 that change state, 4 live-probe.
They exist because two shipped tools could not finish their own job:
import_suitecreates a suite with no run target, andtrigger_suite_runfails fast without one — so an assistant could create a suite it had no way to make runnable.update_suitesets the target (validated through the sameSuiteTargetmodel the REST route uses) and reportsrunnableexplicitly. It also fires the samedispatch_auto_classifythe REST route does, so a suite made runnable here still derives a redaction policy.suggest_column_policycould propose a policy that nothing could read back or apply.
All three gate on suite_authz.require_permission against an existing suite, so no new authz
path is introduced and tests/support/mcp_gates.GATES picks them up in the four sweeps.
The exclusions are unchanged. DELETE /suites/{id} is now explicitly excluded as well: it
cascades every run and result the suite ever produced, and unlike delete_check there is no
lesser action to steer an assistant toward.
Amendment — Tier 3A batch 2, 33 → 38 (2026-08-17)¶
Five more coherence tools, closing the last three asymmetric-verb pairs in the surface:
update_schedule, update_trigger_binding, delete_trigger_binding, list_check_versions,
restore_check_version. The split becomes 38 tools: 18 read-only, 16 that change state,
4 live-probe.
Each pair was create-without-update, or create-without-delete, or a mutation with no way to inspect or undo it:
create_schedule+delete_schedulewith no update meant "pause the nightly run" had only one available answer — delete it — which discards the cron expression the user would need to restore it.update_schedulemakes pause a first-class, reversible action, and both tools now point at each other so the destructive one is not chosen by default.create_trigger_bindinghad neither a delete nor a disable, so an assistant could wire a trigger and then had no way to unwire it.update_check/delete_checksnapshot every edit intocheck_versions, and none of that was readable over MCP.list_check_versionsexposes the edit history — deliberately distinct fromget_check_history's result history, with both docstrings cross-referencing the other, since "did this start failing because the data moved or because someone changed the check?" needs both and the names are otherwise easy to confuse.restore_check_versionis also the only path that can clear a field back to empty:update_check's PATCH convention reads an omitted argument as "leave alone", so it structurally cannot. That correction was applied toupdate_check's own docstring, which had said recreating the check was the only option.
All five gate through require_permission on the owning suite (the schedule and binding tools
resolve it from the row), so again no new authz path. The gate rows in
tests/support/mcp_gates.GATES drive the four sweeps as before — with one trap worth recording:
restore_check_version takes a version_no, and a probe check inserted directly rather than
through check_service has no version rows, so the tool raised "check version not found"
before reaching authz and the sweep passed with the gate deleted. That is the same vacuous-pass
shape _REAL_RUN and _REAL_CHECK were each added to close, one level deeper; the fix inserts a
real CheckVersion alongside the check. Every gate here was mutation-verified by removing it and
confirming the sweep goes red.
The exclusions are unchanged.
Amendment — Tier 3B batch 1, 38 → 42 (2026-08-17)¶
Four read-only tools over the asset and incident surfaces — list_assets, get_asset,
list_incidents, get_incident. The split becomes 42 tools: 22 read-only, 16 that change
state, 4 live-probe.
Unlike Tier 3A, this is not coherence — it is a capability gap the earlier tiering pass never
evaluated, because that pass was written on 2026-07-04 and assets, lineage and incidents shipped
on 2026-07-10/11. It is also the grain users actually reason in: "is orders healthy?" is an
asset question, and answering it from suites alone requires knowing which suites target the
table, which is exactly what the asset view exists to remove.
Two scoping rules had to be carried into the docstrings, not just the code.
- The asset rollup is workspace-true (ADR 0037):
list_visible_assetstakes no user and aggregates over ALL composing suites. That is the established REST decision and is mirrored unchanged — but a model handed a health number with no caveat will describe it as covering what the caller can see.get_assettherefore reportsrestricted_suite_countbeside the grant-filteredsuiteslist, and both docstrings state the split explicitly. - Incidents stay behind suite grants (also ADR 0037, deliberately unchanged there): they are
itemized failure evidence, not identity.
get_incidentis 404-no-leak.
That second rule is why incident:view is a new gate value in mcp_gates.GATES rather than
suite:view: the tool takes an incident id and resolves the ladder through the incident's
suite, so the RBAC sweep needs a real incident materialised on the probe suite. A fabricated id
raises "incident not found" — an accepted denial word — and the sweep would have passed with the
gate deleted, the same vacuous-pass shape as _REAL_RUN / _REAL_CHECK / the Tier-3A
CheckVersion trap. Both new gates were mutation-verified by removing them and confirming the
sweep goes red.
The 404-no-leak rule itself moved out of the HTTP layer into
incident_service.load_visible_incident, so REST and MCP share one implementation. A second
hand-written copy of the only thing standing between an incident's evidence and an ungranted
caller is the "guard applied at one door and not its sibling" shape this track has hit
repeatedly, and the divergence would be invisible until it leaked.
Three honesty fields exist because the underlying value is true and misleading on its own:
monitored (an asset with no suite has worst_severity: null, which reads as a clean bill of
health for something nothing checks), truncated (computed against the real total, since
len(page) == limit is wrong on the exact-boundary page), and lineage.qualified_by (an empty
graph behind a failing poller must never read as "nothing feeds this table" — a known confident-wrong-answer failure mode).
The exclusions are unchanged. PATCH /assets/{id} is not exposed: it is workspace-Admin-only
(ADR 0034 §4), and MCP exposing no admin-only tool at all remains an asserted invariant.
Amendment — Tier 3B batch 2, 42 → 46 (2026-08-17)¶
ack_incident, resolve_incident, list_columns, get_near_misses. The split becomes
46 tools: 23 read-only, 18 that change state, 5 live-probe. Tier 3B is complete.
list_columns is classified as a live probe, not a read, for the profile_column reason
exactly: it persists nothing but opens a live datasource connection with the stored credential,
so it gates on edit. It is also the cheap authoring step the surface was missing —
profile_column reads data to compute statistics when the question was only "what are the column
names?", and a guessed column produces a check that runs and errors.
get_near_misses closes a loop create_trigger_binding opened: that tool warns about an
environment mismatch it has no way to investigate, and before this route existed the only way to
see one was a direct database query. Its docstring states the asymmetry explicitly — an empty result does
not prove a binding is firing, only that no mismatch was observed.
The two lifecycle verbs required the most docstring care in the batch, because both are
statements about the incident and neither touches the data. Acknowledging does not stop
alerting (the check still runs, still fails, still notifies) and points at snooze_check for
what a user asking to "silence it" usually means. Resolving re-runs nothing and fixes nothing: if
the underlying problem persists, the next failing run opens a new incident, since a resolved
incident is never reopened — so the tool steers toward trigger_suite_run to confirm a fix
rather than resolving on an assumption. Both refuse the closed-state transition (the service's
IncidentNotActiveError) instead of silently no-opping, which would let an assistant report an
action that never happened.
One harness finding is worth recording: SuiteForbiddenError from load_visible_incident names
the required level ("requires 'edit' on its suite") and matched none of the RBAC sweep's
accepted-denial vocabulary, so a correctly-denied tool failed the sweep for looking like the
wrong kind of failure. That exact phrase was added — deliberately not a loose word like
"requires", because anything that also matches an argument-validation message would let a
gateless tool pass by failing for the wrong reason. Every gate in this batch was mutation-
verified.
The exclusions are unchanged.
Amendment — get_adf_pipeline_status renamed to get_pipeline_status, 46 → 47¶
The name predated dbt (ADR 0029) and Airflow support and was never ADF-only — it always covered
ADF, Airflow and dbt. Renamed to get_pipeline_status, which states its actual scope. The old
name stays registered as a deprecated, behaviorally-identical alias (a plain delegate — same
arguments, same validation, same return shape) so a client with it pinned in a saved prompt or
static config does not break. The split becomes 47 tools: 24 read-only, 18 that change state, 5
live-probe — the extra tool is the alias, not new capability.
Amendment — get_doc added, 47 → 48¶
A read-only tool exposing DataQ's own published docs pages (docs/site/) to an AI client, so a
question like "how does DataQ handle X" can be answered from the curated docs rather than the
model guessing or an assistant grepping the repo. A hand-scanned allowlist (five top-level pages
plus every compliance/*.md page) is walked live at call time — the same "declare as data, not
as code comments" shape mcp_gates.GATES itself enforces — and deliberately excludes ADRs and
architecture.md (contributor design-rationale, not this tool's audience). Returns the page
verbatim, no summarization; an unrecognized page restates the current valid list, and the
parameter's own JSON-schema enum carries the live catalog, so a client sees the full valid set
on every turn with no companion list_doc_pages() tool needed. Workspace-agnostic — no per-suite
or per-workspace gate, since these pages carry no tenant data. The split becomes 48 tools: 25
read-only, 18 that change state, 5 live-probe.