Skip to main content

MCP Server: Agentic Access

miStudio ships an optional MCP (Model Context Protocol) server that exposes the post-extraction workflow — feature analysis, cross-feature clustering, circuit discovery, steering, calibration, and label write-back — as tools for agentic AI clients like Claude Code. An agent can run the full analyze → group → steer → relabel loop autonomously, or the deeper capture → discover → validate → calibrate → promote → export circuit pipeline, with every action flowing through the same REST API and appearing in the UI like any other work.

The server also carries an optional cross-product plane: when pointed at a running miLLM runtime, a family of millm_* tools lets the same agent move a tuned circuit or cluster into production and drive live serving. miStudio discovers and calibrates (it runs the model to learn); miLLM serves (it runs the model behind an OpenAI-compatible API). The boundary between the two is a portable document, not a code dependency.

Enabling the Server​

The MCP server is off by default and runs as its own container.

Docker Compose:

# 1. Set a token in .env (required — the port is LAN-reachable)
echo "MCP_AUTH_TOKEN=$(openssl rand -hex 32)" >> .env

# 2. Start with the mcp profile
docker compose --profile mcp up -d

Kubernetes: add mcp-auth-token to the mistudio-secrets Secret and apply the mistudio-mcp Deployment/Service/Ingress in k8s/base/mcp.yaml. The shipped Ingress (streaming-friendly nginx annotations: proxy-buffering off, 3600s read timeout) serves three routes:

  • http://mcp-mistudio.hitsai.local/mcp and http://mcp-mistudio.hitsai.net/mcp — dedicated MCP hosts with full path space (/health works too). These resolve only where you make them resolve (hosts file or local DNS pointing at the ingress IP); the .net name is not published in public DNS unless you do so yourself.
  • http://k8s-mistudio.hitsai.local/mcp — path route on the shared miStudio host, kept for compatibility.

Point a client machine at the ingress with a hosts-file entry, e.g. 192.168.244.61 mcp-mistudio.hitsai.local mcp-mistudio.hitsai.net. No Ingress at all? kubectl port-forward svc/mistudio-mcp 8765:8765 works too.

Verify: curl http://<host>:8765/health returns the enabled tool categories.

Network exposure

The server binds 0.0.0.0:8765 so agents on other LAN machines can connect. A bearer token is always required on this transport — and that is now enforced, not just stated: MCP_ALLOW_ANONYMOUS is honoured on stdio only, and the server refuses to start if it is set over HTTP.

Until this was fixed, the flag alone satisfied the startup guard on HTTP, so following the troubleshooting remedy below produced a LAN-reachable unauthenticated server — on the page that told you a token was always required.

If your agents run on the same host only, firewall port 8765.

Connecting a Client​

Claude Code:

# Docker Compose (direct port):
claude mcp add --transport http mistudio http://<host>:8765/mcp \
--header "Authorization: Bearer $MCP_AUTH_TOKEN"

# Kubernetes (via the shipped ingress, dedicated host):
claude mcp add --transport http mistudio http://mcp-mistudio.hitsai.local/mcp \
--header "Authorization: Bearer $MCP_AUTH_TOKEN"

Any MCP client: streamable-HTTP transport, URL http://<host>:8765/mcp, header Authorization: Bearer <token>. A stdio mode exists for local development: python -m src.mcp_server --stdio (inside the backend container/venv).

Give your agent the operating manual

After connecting, paste the MCP Agent Instructions into the agent's context (or tell it to fetch that page). It covers the tool catalog semantics, the analyze→group→steer→relabel recipes, guardrail reactions, and the evidence-notes convention — agents perform dramatically better with it.

Configuration​

Env variableDefaultPurpose
MCP_AUTH_TOKEN(required)Bearer token; startup refused if empty
MCP_TOOL_CATEGORIESread,groups,steering,labeling,experiments,profiles,circuits,jlens,jobs,models,probe_monitorsWhich tool categories are exposed; add admin to enable destructive deletes
MILLM_API_URL(unset)Base URL of a miLLM runtime. Set it to enable the six cross-product millm_* categories (millm_circuits, millm_clusters, millm_models, millm_probes, millm_runtime, millm_sensing); they never appear unless this is set
MCP_STEERING_MAX_CONCURRENT2Max in-flight agent steering tasks
MCP_STEERING_MAX_NEW_TOKENS512Ceiling on generation length for agent steering
MCP_STEERING_APPROVALfalseRoute agent steering through operator approval (see below)

Disabled categories simply don't appear in the agent's tool list. Disabling the whole server (omit the compose profile / scale the deployment to 0) never affects the frontend or REST API.

The circuits category is on by default as of the circuits arc. The five millm_* categories are the cross-product plane — miStudio discovers and calibrates, miLLM serves — and are gated separately: they require MILLM_API_URL and are never enabled by default, even if you list them in MCP_TOOL_CATEGORIES.

Tool Catalog​

The authoritative inventory is docs/mcp-contract.md in the repo — an auto-generated, diff-tested file derived from the live tool registry. The table below mirrors it; the contract wins if they ever drift.

Agents: call mistudio_howto first

mistudio_howto is the "call me first" tool. A flat table can't carry ordering constraints, GPU-lock contention, id namespaces, or the failure modes that mislead — that tool holds the workflows (feature analysis, the circuit-discovery evidence ladder, the cross-plane document hand-off) and the guardrail reactions. Point your agent at it before it starts circuit or steering work.

CategoryTools
read (12)list_extractions, get_extraction_summary, list_trainings, search_features, get_feature, get_feature_examples, get_feature_token_analysis, get_feature_logit_lens, get_feature_correlations, get_feature_ablation, get_feature_nlp_analysis, mistudio_howto
groups (6)compute_feature_groups, get_grouping_status, get_feature_groups, get_feature_group_members, find_features_by_token, find_related_features
steering (10)steering_status, get_steering_mode, enter_steering_mode, exit_steering_mode, steer_compare, steer_sweep, steer_combined, compute_cluster_allocation, get_steering_result, cancel_steering_task
circuits (24) (default-on)start_circuit_capture, list_circuit_captures, run_circuit_discovery, get_discovery_results, run_attribution_pass, validate_circuit_edges, create_circuit, build_circuit_from_discovery, update_circuit, get_circuit, list_circuits, delete_circuit ⚠️, run_circuit_faithfulness, calibrate_circuit_strength, reproduce_calibration, promote_circuit, export_circuit_definition, export_circuit_slices, import_circuit_definition, record_steering_samples, get_steering_samples, list_validation_manifests, get_validation_manifest, reproduce_validation
experiments (3)save_experiment, list_experiments, get_experiment
profiles (4)list_cluster_profiles, get_cluster_profile, save_cluster_profile, export_cluster_definition — durable cluster profiles + portable mistudio.cluster-definition/v1 export
labeling (13)update_feature_label, run_enhanced_labeling, get_enhanced_label, list_labeling_templates, run_labeling_trial, get_labeling_trial, list_labeling_trials, compare_labeling_trials, get_labeling_coverage, resume_labeling, start_labeling_resume_sweep, get_labeling_resume_sweep, cancel_labeling_resume_sweep — run_labeling_trial and the three that read trials score a template against a fixed feature panel and write NO labels. The last five are resume: get_labeling_coverage reports what still needs labeling and WHY the failures failed, resume_labeling does one batch of it, and the sweep trio runs many batches back to back. start_labeling_resume_sweep requires an explicit max_batches — there is no unbounded option, because a sweep is hours of GPU time per batch
jlens (20) (default-on)list_jlens_artifacts, preview_jlens_repo, acquire_jlens_artifact, fit_jlens_artifact, validate_jlens_artifact, publish_jlens_artifact, restore_jlens_artifact, cancel_jlens_task, jlens_readout, get_jlens_readout, jlens_cost_estimate, compute_jlens_band_report, get_jlens_band_report, record_jlens_gate, get_jlens_gate, get_jlens_replication_report, run_jlens_intervention, get_jlens_interventions, annotate_jlens_feature, create_jlens_watchlist — the Jacobian-lens substrate
probe_monitors (6) (default-on)list_probe_monitors, get_probe_monitor_report, recalibrate_probe_monitor, build_probe_definition, export_probe_definition, publish_probe_definition — trained probe monitors and their portable mistudio.probe-definition/v1 export. Read the rung before the AUROC: a probe's rung is what its evidence supports (0 trained only, 1 held-out, 2 unseen tasks, 3 judge-compared), and build_probe_definition REFUSES below rung 2 unless it is given a reason — which is then written into the exported file, so whoever serves the probe can see the claim was made on a judgement. Two refusals are not waivable: an SAE probe whose dictionary has no HuggingFace home cannot be exported at all (publish the SAE first), and a run still in flight is refused because its rung can still change. publish_probe_definition takes NO token — it is resolved server-side from Settings → API Keys, so no credential passes through an agent transcript. recalibrate_probe_monitor moves the decision threshold to another false-positive budget with no GPU and no queue: a threshold is the (1 − target_fpr) quantile of the negatives the run already saved to disk, so re-cutting it is arithmetic over an array rather than the ~2.6-hour run it used to cost. It previews by default, and the preview carries what the candidate bar would do to every evaluation set the probe was measured on. ⚠ It does NOT make a threshold transfer between distributions — a re-cut bar is still a quantile of the same calibration corpus, and on the first shipped probe the five sets' own 1% thresholds spanned 24 points. Read transfer_proposed.caution before committing
models (6) (default-on)list_models, get_model, get_model_architecture, download_model, cancel_model_download, delete_model ⚠️ — the weights miStudio RUNS ON. No load/unload: miStudio loads per task and has no resident model
jobs (1)get_task_status
admin (2) (off by default)delete_experiment ⚠️, delete_extraction ⚠️ — destructive

Cross-product plane — the six millm_* categories drive a live miLLM runtime and only appear when MILLM_API_URL is set:

CategoryTools
millm_circuits (16)millm_import_circuit, millm_list_circuits, millm_export_circuit, millm_delete_circuit ⚠️, millm_activate_circuit, millm_deactivate_circuit, millm_circuit_status, millm_set_circuit_intensity, millm_circuit_claims, millm_release_circuit_claims, millm_circuit_sensing_enable, millm_circuit_sensing_disable, millm_circuit_sensing_status, millm_circuit_sensing_events, millm_circuit_sensing_event, millm_circuit_sensing_clear ⚠️
millm_clusters (6)millm_import_cluster, millm_list_clusters, millm_export_cluster, millm_activate_cluster, millm_deactivate_cluster, millm_hub_search
millm_models (8)millm_list_models, millm_get_model, millm_preview_model_repo, millm_download_model, millm_cancel_download, millm_load_model, millm_unload_model, millm_delete_model ⚠️ — the weights miLLM SERVES. It holds ONE model resident, so millm_load_model EVICTS the current one
millm_probes (7)millm_import_probe, millm_list_probes, millm_arm_probe, millm_recalibrate_probe, millm_disarm_probe, millm_probe_status, millm_probe_events — run a probe monitor trained here against miLLM's live traffic. Arming passes four gates (armed-probe limit, model identity, evidence rung, parity); a probe below rung 2 needs acknowledge_below_rung2. A probe records and reports — nothing here stops or alters a generation millm_recalibrate_probe moves an already-imported, possibly armed probe's decision bar in place: it re-cuts the threshold here from the negatives the run already saved — no GPU — and pushes the new decision block to miLLM, which updates the row, the stored definition and the live runtime shape together. ⚠ It is the only way to change an imported probe's bar without losing its history: the alternative is disarm → delete → import, which assigns a new probe id and cascade-deletes every recorded verdict. It changes the BAR, never the detector — on_conflict=replace is still refused, and miLLM refuses any payload that could carry a detector or a cut it cannot match to the stored provenance.probe_id.
millm_runtime (5)millm_status, millm_list_profiles, millm_activate_profile, millm_deactivate_profile, millm_set_intensity
millm_sensing (5)millm_sensing_enable, millm_sensing_disable, millm_sensing_status, millm_sensing_events, millm_sensing_config

⚠️ = destructive or irreversible.

Long-running operations follow the platform's async pattern: the start tool returns a task id, and the agent polls get_task_status / get_steering_result. Task ids are durable database records, so they survive agent reconnects.

Probe monitors: running a detector on live traffic​

A probe monitor is a linear detector trained here — a vector, a threshold, and a record of how well it actually worked. millm_probes imports one into miLLM, arms it, and reads what it reported.

Two things about arming are worth knowing before an agent tries it:

  • Identity cannot be waived. A probe's weights are a direction in one specific model's residual space. Read in another model's space they produce numbers that are plausible, stable and about nothing, so millm_arm_probe refuses with PROBE_MODEL_MISMATCH and names every field that differs — which tells you whether the wrong model is loaded or the wrong probe was imported.
  • Below rung 2 needs explicit consent. The refusal is UNVALIDATED_PROBE, and it is resolved by passing acknowledge_below_rung2=true, not by retrying. That acknowledgement is recorded separately from the one inside the definition: whoever exported a weak probe and whoever arms it against live traffic are not necessarily the same, and only the second is choosing to monitor with it.

Read rung_language verbatim rather than composing a phrase from rung. Rung 3 reads "detects on unseen tasks, compared with a judge" — compared with, not better than. On the reference run here the judge won.

A quiet armed probe always says why, in paused_reason. Silence would read as "nothing detected", which is a claim no probe made.

Circuits: discovery, calibration, and recording​

The circuits category exposes the full mechanistic-interpretability pipeline. Two tools from the current arc are worth calling out:

  • Calibration (calibrate_circuit_strength → reproduce_calibration): a two-detector usable-band search. It finds the ONSET where steering starts to bite (output-drift, no judge needed) and the CORRECTNESS CLIFF where the model starts producing fluent-but-false output (an LLM judge scores generated neutral-topic falsifiable probes), bisects between them, and clamps the served dial to [onset, cliff]. It is a badge, not a gate — it annotates the circuit, it does not block serving. A judge too weak to grade its own probes reports judge_unreliable rather than a false no_band.
  • Steered Transcript Recorder (record_steering_samples → get_steering_samples): an instrument, not a judge. It records (dial, prompt, unsteered, steered) transcripts for a circuit, a cluster, or an ad-hoc feature set, so a stronger model can analyze the run afterwards. It never scores anything itself.

Both write manifests — self-contained, reproducible records — and both have a reproduce_* tool that re-runs from the manifest and reports a delta verdict, so a claim is reproducible rather than a one-off.

Guardrails & Provenance​

  • Label provenance: agent label edits carry label_source: mcp_agent, so you can always tell agent work from human work.
  • Protected labels: aqua-starred features (completed enhanced labels) return a 409 when an agent tries to edit their name/category/description; the agent must pass an explicit override_protected flag. Notes remain appendable — the documented convention is [MCP <date>] evidence: experiment <id> — <summary>.
  • Steering limits: concurrency cap and generation-length ceiling are enforced before the GPU is touched.
  • Operator-approval mode: with MCP_STEERING_APPROVAL=true, agent steering calls become pending requests instead of running. They appear as a banner in the Steering panel with Approve/Deny buttons; on approval the backend submits the stored request itself, and the agent's poll picks up the resulting task id.
  • Audit: every tool call is logged (tool name, argument digest, status, duration).

Worked Example: the Analyze → Group → Steer → Relabel Loop​

A Claude Code session pointed at the server, instructed with natural language only:

  1. "List extractions and build feature clusters for the newest one" → list_extractions, compute_feature_groups, get_task_status until complete
  2. "Show me the biggest groups" → get_feature_groups(sort_by="size") — say it finds a 12-member "love" group with cohesion 0.81
  3. "What do the members have in common?" → get_feature_group_members, then get_feature_examples + get_feature_logit_lens per member → agent hypothesizes "expressions of affection"
  4. "Validate that by steering" → enter_steering_mode, steer_sweep(feature_idx=4821, strength_values=[0, 10, 30]), get_steering_result → steered outputs turn affectionate; baseline doesn't
  5. "Save the evidence and update the labels" → save_experiment, then update_feature_label(name="expressions of affection", notes="[MCP 2026-07-12] evidence: experiment exp_… — sweep shows dose-dependent affection shift")
  6. Every record — the group index, the experiment, the relabeled features with mcp_agent provenance — is immediately visible in the miStudio UI.

Troubleshooting​

SymptomFix
Server exits at startup: "MCP_AUTH_TOKEN is required"Set MCP_AUTH_TOKEN in .env. MCP_ALLOW_ANONYMOUS will not help here — it is honoured on the stdio transport only, and the server refuses to start over HTTP with it set. That is deliberate: the HTTP port is LAN-reachable and exposes delete_circuit, GPU steering and label write-back.
401 from every callClient isn't sending Authorization: Bearer <token>, or tokens don't match
Tools error "backend unreachable"The mcp-server container can't reach backend:8000 — check both are on the same network and the backend is healthy
steer_* returns a guardrail messageConcurrency cap hit — poll or cancel existing tasks, or raise MCP_STEERING_MAX_CONCURRENT
Steering tools return pending_approvalApproval mode is on — approve the request in the Steering panel