udmi

UDMI Workbench: Architecture & Interface Contracts

Status: Describes the implementation on this branch. Items that are part of the vision but not built are listed only in §10.
Package Target: workbench/ (Workbench UI and gateway), mantis/ (agent and MCP tools)


1. Executive Summary & Vision

The Workbench vision converges three evolutions of UDMI device testing into one product:

  1. Mantis Workbench: LLM-assisted diagnostic reasoning that helps device developers understand why tests fail and how to fix them.
  2. Test Cadre: Reproducible local test infrastructure that lets lab operators run compliance suites without cloud dependencies.
  3. Unified Product: Standardized reporting that lets ecosystem curators certify devices.
                  ┌────────────────────────────────────────────────────────┐
                  │                   WORKBENCH VISION                     │
                  └──────────────────────────┬─────────────────────────────┘
                                             │
             ┌───────────────────────────────┼──────────────────────────────┐
             ▼                               ▼                              ▼
   ┌───────────────────┐           ┌───────────────────┐          ┌───────────────────┐
   │ Mantis Workbench  │           │    Test Cadre     │          │  Unified Product  │
   │ - Assisted Diag   │           │ - Local Stack     │          │ - Compliance View │
   │ - Root Cause RCA  │           │ - Pubber / DUT    │          │ - Device Reports  │
   │ - Dev Remediation │           │ - Swappable Stack │          │ - Support Bundles │
   └─────────┬─────────┘           └─────────┬─────────┘          └─────────┬─────────┘
             │                               │                              │
             ▼                               ▼                              ▼
      Device Developer               Test Lab Operator              Ecosystem Curator
      (Deep-dive 1 DUT)              (Broad multi-DUT)              (Certify & catalog)

Architectural Foundation: The MCP-First Design

The Workbench is one Python process, workbench/server/gateway.py, a stdlib ThreadingHTTPServer started by bin/workbench. It serves the single-page frontend and three backend planes:

The frontend is plain ES modules with no framework and no build step (workbench/static/).


2. The 3 Primary User Personas & Workflows

Persona Core Mission Key UI Needs Primary Contracts Used
Device Developer Develop and validate one IoT device before deployment. Failure triage, schema inspection, state-sync timelines, remediation. Contract 2 (Agent SSE), Contract 1 (MCP Tools), Contract 4 (Workspace State)
Test Lab Operator Manage physical and virtual test rigs, run suites across many units. Local stack setup (//mqtt/localhost:<port>), Pubber or a physical DUT, setup logs, live run logs. Contract 1 (Testbed REST), Contract 3 (Run Log Streaming), Contract 4 (Workspace State)
Ecosystem Curator Certify devices for installations. Per-device compliance by feature bucket and stage, device reports, support bundles. Contract 5 (Compliance & Reporting)

3. High-Level System Architecture

┌─────────────────────────────────────────────────────────────────────────────────────────┐
│                           Workbench UI (plain ES modules SPA)                           │
│  ┌───────────────────────────┐ ┌───────────────────────────┐ ┌────────────────────────┐ │
│  │ /sequencer                │ │ /devices                  │ │ Drawers & dialogs      │ │
│  │ - Run config & test list  │ │ - Compliance per device   │ │ - Mantis (Ctrl/Cmd+K)  │ │
│  │ - Live run log            │ │ - Bucket x stage matrix   │ │ - Logs (/logs)         │ │
│  │ - Local Test Setup drawer │ │ - Device reports          │ │ - Settings (gear)      │ │
│  └─────────────┬─────────────┘ └─────────────┬─────────────┘ └───────────┬────────────┘ │
│                └──────────────────────┐      │      ┌────────────────────┘              │
│                                       ▼      ▼      ▼                                   │
│                           ┌────────────────────────────────────────┐                    │
│                           │  WorkspaceStore (Contract 4, store.js) │                    │
│                           │  - siteModel, deviceId, activeTestId   │                    │
│                           │  - projectSpec, sessionId, running     │                    │
│                           └──────────────────┬─────────────────────┘                    │
└──────────────────────────────────────────────┼──────────────────────────────────────────┘
                                               │
     ┌─────────────────────────────────────────┼─────────────────────────────────────────┐
     │ 1. MCP Tool Calls (JSON-RPC)            │ 2. Agent SSE Stream    │ 3. Run Log SSE   │
     │ (POST /message or POST /rpc)            │ (POST /api/mantis/chat)│ (GET /api/       │
     │                                         │                        │  sequencer/stream│
     ▼                                         ▼                        ▼                  ▼
┌───────────────────────────────┐ ┌─────────────────────────┐ ┌──────────────────────────┐
│     MCP Server & Registry     │ │      Mantis Engine      │ │ Sequencer runner, testbed│
│   (mantis/mcp_server.py,      │ │    (mantis/agent.py)    │ │ & tmux sessions          │
│    mantis/tools/registry.py)  │ │ - Scoping               │ │ (workbench/server/       │
│ - 27 tools                    │ │ - Actor (ReAct)         │ │  runner.py, testbed.py;  │
│ - Site Model & Schema Read    │ │ - Critic                │ │  mcp/infra/              │
│ - DOT / Mermaid diagrams      │ │ - Arbitrator            │ │  session_manager.py)     │
└───────────────┬───────────────┘ └────────────┬────────────┘ └────────────┬─────────────┘
                └──────────────────────────────┴───────────────────────────┘
                                               │
                                               ▼
                              ┌──────────────────────────────────┐
                              │       Local Substrate / DUT      │
                              │  - Mosquitto, etcd, UDMIS        │
                              │  - DUT (Physical / Pubber)       │
                              └──────────────────────────────────┘

3.1. Routes, Drawers and Dialogs

The SPA has two routed views. Everything else is a drawer or dialog over the current view.

URL / Surface Primary Persona Purpose & Capabilities
/sequencer (default) Developer / Operator Pick a site model, device and target spec, select tests, run bin/sequencer, watch the live log, and inspect artifacts (RESULT.log, sequence.md, sequence.png). Hosts the Local Test Setup drawer (§3.2). Failed test rows offer Diagnose with Mantis; every row offers Explain this test.
/devices Curator / Operator Compliance view: per-device verdict, feature bucket × stage score matrix, run targets, and reports (Contract 5).
/logs Operator / Developer Not a view. Loading it opens the Logs drawer (also opened by the header terminal icon) over /sequencer. It shows the client diagnostic ring buffer and GET /api/diagnostics/logs?limit=.
Mantis drawer All Opened with Cmd/Ctrl+K, the header button, or a test row action. It inherits the active site model, device and test.
Settings dialog All Header gear icon. Email notification consent and recent deliveries (§5.4.1).

Deep links are query parameters: ?site_model=<name>&device=<id>. Any other path falls back to /sequencer.

Mantis Entry Points

  1. Diagnose with Mantis: on failed test rows in /sequencer and in the Artifact Viewer. It opens the Mantis drawer and sends a triage request with {site_model, device_id, test_id}.
  2. Explain this test: on every test row. It opens the drawer and asks how the test works.
  3. Free-form questions in the drawer, with the failed-test selector in the drawer toolbar.

3.2. Local Test Setup Drawer (Sequencer Screen Integration)

The Local Test Setup drawer docks to the right of /sequencer, so operators can manage the local stack without a terminal.

┌───────────────────────────────────────────────────────────────────────────────────────────────────┐
│ Sequencer Workspace (/sequencer)                                       │ [Local Setup Drawer]     │
│ ┌────────────────────────────────────────────────────────────────────┐ │ ┌───────────────────────┐│
│ │ Run Configuration (Site Model, Device, Target Spec)                │ │ │ [Start] [Stop] [Restart││
│ ├────────────────────────────────────────────────────────────────────┤ │ ├───────────────────────┤│
│ │ [Min stage: PREVIEW ▼] [Bucket ▼] [Select all] [Clear]             │ │ │ Topology Graph        ││
│ │ [✓ Passed] [– Skipped] [✗ Failed] [Search]                         │ │ │  [DUT / Pubber]       ││
│ ├────────────────────────────────────────────────────────────────────┤ │ │        │ (MQTT:port)  ││
│ │ Test list                                                          │ │ │        ▼              ││
│ │  pointset_publish     PASSED   [View Artifacts] [Explain]          │ │ │  [Mosquitto Broker]   ││
│ │  system_last_start    FAILED   [Diagnose with Mantis] [Explain]    │ │ │        ▼              ││
│ ├────────────────────────────────────────────────────────────────────┤ │ │  [Local UDMIS Pod]    ││
│ │ Live run log (SSE, Contract 3)                                     │ │ │        ▼              ││
│ └────────────────────────────────────────────────────────────────────┘ │ │  [etcd State Store]   ││
│                                                                        │ └───────────────────────┘│
└───────────────────────────────────────────────────────────────────────────────────────────────────┘

3.2.1. Drawer Affordance & Behavior

3.2.2. Topology Graph

3.2.3. Status Model

Probes:

3.2.4. Lifecycle Control Actions


4. Contract 1: MCP Tools & Workbench REST (Deterministic Operations)

4.1. Supported JSON-RPC Methods

resources/list and resources/read are not implemented and return -32601.

4.2. Tool Registry

All 27 tools are registered in mantis/tools/registry.py (get_mcp_tools()). Argument models live in mantis/models.py.

A. Environment & Test Setup Lifecycle

B. Sequencer Execution & Diagnosis

C. Site Model, Schemas & Source

D. Visualizations

4.3. Custom Site Model Roots

Site models may live outside the UDMI root. A directory is used only after the operator approves it in the consent modal. Approved roots are stored server-side in ~/.config/udmi/workbench.json under site_roots, never in the browser.

4.4. Local Testbed REST Endpoints

The Local Test Setup drawer manages the local stack through these endpoints. Every lifecycle and status call names the drawer’s own project spec, which must be //mqtt/localhost:<port> with an explicit unprivileged port (1024-65535, not 8883). All probed ports are derived from that spec per call (MQTT = port, etcd = port+1). A missing or non-local spec is rejected with HTTP 400. Start runs in the background: a later bin/udmi failure shows only as overall: "ERROR" with last_error quoting the tail of out/testbed_setup.log. Stop, and the stop phase of restart, return HTTP 500 when the command fails or the ports stay open.

4.5. Other Workbench REST Endpoints

| Method & Path | Purpose | | :— | :— | | GET /api/health | Gateway liveness. | | GET /api/site-models | Discovered site models (UDMI root and approved roots). | | GET /api/devices?site_model= / GET /api/devices/summary?site_model= / GET /api/device?site_model=&device_id= | Devices, with a summary form for the pickers. | | GET /api/sequences | Sequencer test catalog. | | GET /api/results?site_model=&device_id= | Previous results for the test list. | | GET /api/browse, GET /api/file, GET /api/repo-doc | Artifact and doc viewers. | | GET /api/sequencer/options, GET /api/sequencer/sessions | Run options and active runs. | | POST /api/sequencer/run | Start a run (§6). | | POST /api/sequencer/stop {session_id} | Stop a run. | | GET /api/results/commit/preview, POST /api/results/commit {site_model, message, branch, create_branch, push, remote} | Commit results to the site model repo. | | POST /api/support-bundle {site_model}, GET /api/support-bundle/download?bundle_id= | Build and download a gzip tarball support bundle. | | GET /api/diagnostics/logs?limit= | Server diagnostic log for the Logs drawer. |


5. Contract 2: Streaming AI Agent Contract (Mantis Cognitive Loop)

5.1. Execution Modes

Mode selection happens in mantis_adapter.py.

Mode A: Targeted Failure Triage

Mode B: General Exploration & Specification

5.2. Request

The adapter reads only session_id (default "sess-default"), message, notify, and context.{site_model, device_id, test_id}. The Workbench UI sends session_id: "workbench-assistant". The model provider is chosen by the server environment, not per request (§9).

Example 1: Targeted Failure Triage Request

POST /api/mantis/chat HTTP/1.1
Content-Type: application/json
Accept: text/event-stream

{
  "session_id": "workbench-assistant",
  "message": "Why did test pointset_publish fail for AHU-1?",
  "context": {
    "site_model": "sites/udmi_site_model",
    "device_id": "AHU-1",
    "test_id": "pointset_publish"
  },
  "notify": false
}

Example 2: General Exploration Request

POST /api/mantis/chat HTTP/1.1
Content-Type: application/json
Accept: text/event-stream

{
  "session_id": "workbench-assistant",
  "message": "What is the state field configAcked and how does it relate to config synchronization?",
  "context": {
    "site_model": "sites/udmi_site_model"
  }
}

5.3. Canonical Event Wire Schema

Event Data Notes
phase {"phase": "SCOPING" \| "ACTOR" \| "CRITIC" \| "ARBITRATOR"} Phase transitions.
thought {"text"} Interim prose from the agent.
tool_call {"call_id", "tool", "args", "timestamp"}  
tool_result {"call_id", "tool", "summary", "output"} summary is the record’s status, or "completed". It is not prose.
token {"text"} Sent once with the whole final answer (also once for a slash command). Diagrams are ` ```mermaid ` blocks inside it, rendered by the client.
hypothesis_matrix {"hypotheses": [...], "final": bool, "audit_markdown"?} Interim matrices have rows with verdict UNRESOLVED, rationale "Under investigation.", evidence_tier "NONE". The final matrix is sent after the token and carries audit_markdown; its rows may have null rationale or evidence_tier.
done {"session_id", "metrics": {"steps", "tool_calls", "tripartite_status", "agent_duration_sec", "total_duration_sec"}} A slash command’s done has metrics: {}.
error {"message"} Validation failure, busy session, notify refusal, offline provider, missing triage context, operator stop, agent failure, empty answer, or the 600 s view timeout. Ends the stream.

Hypothesis rows are {"hypothesis", "verdict": "PRIMARY" | "CONTRIBUTING" | "REFUTED" | "UNRESOLVED", "rationale", "evidence_tier": "LOCAL_FILE" | "CLOUD" | "NONE"}.

event: phase
data: {"phase": "ACTOR"}

event: tool_call
data: {"call_id": "call-1", "tool": "get_test_timeline", "args": {"test_id": "pointset_publish", "device_id": "AHU-1"}, "timestamp": "2026-09-15T15:00:00Z"}

event: tool_result
data: {"call_id": "call-1", "tool": "get_test_timeline", "summary": "completed", "output": {...}}

event: phase
data: {"phase": "CRITIC"}

event: phase
data: {"phase": "ARBITRATOR"}

event: token
data: {"text": "The test failed because ...\n\n```mermaid\nsequenceDiagram\n...\n```"}

event: hypothesis_matrix
data: {"final": true, "audit_markdown": "...", "hypotheses": [{"hypothesis": "State update arrived after the cutoff", "verdict": "PRIMARY", "rationale": "...", "evidence_tier": "LOCAL_FILE"}]}

event: done
data: {"session_id": "workbench-assistant", "metrics": {"steps": 9, "tool_calls": 7, "tripartite_status": "SUCCESS", "agent_duration_sec": 41.2, "total_duration_sec": 41.9}}

5.4. Session Control Endpoints

5.4.1. Email Notifications (“Email me when done”)

5.4.2. Desktop Notifications

5.5. Slash Commands (In Chat Prompt)

Every message starting with / is handled as a slash command without running the agent.

Any other slash command returns “Unknown command”. The Mantis drawer has a Copy transcript button for exporting a conversation.

5.6. Layer 0 Canonical Pydantic Models (mantis/models.py)

class ClaimStatus(str, Enum):
    CONFIRMED = "CONFIRMED"
    REFUTED = "REFUTED"
    UNVERIFIED_ASSUMPTION = "UNVERIFIED ASSUMPTION"
    NOT_ASSESSED = "NOT ASSESSED"

class MessageRole(str, Enum):
    USER = "user"
    ASSISTANT = "assistant"
    SYSTEM = "system"
    TOOL = "tool"

class ChatMessage(BaseModel):
    role: MessageRole
    content: str
    timestamp: str  # ISO-8601
    tool_name: Optional[str] = None
    tool_call_id: Optional[str] = None

class ExecutionMetrics(BaseModel):
    total_steps: int = 0
    total_duration_sec: float = 0.0
    api_calls_count: int = 0
    prompt_tokens: int = 0
    candidates_tokens: int = 0
    total_tokens: int = 0
    tool_calls: Dict[str, int] = Field(default_factory=dict)
    retry_count: int = 0
    tripartite_degraded: bool = False
    tripartite_status: str = "SUCCESS"

class SessionContext(BaseModel):
    active_site_model: str = "sites/udmi_site_model"
    active_session_id: Optional[str] = None
    active_device_id: Optional[str] = None
    active_test_id: Optional[str] = None
    history: List[ChatMessage] = Field(default_factory=list)
    metrics: Optional[ExecutionMetrics] = None

class VerificationClaim(BaseModel):
    claim: str
    status: ClaimStatus
    evidence: str

class HypothesisEvaluation(BaseModel):
    status: ClaimStatus
    evidence: str

class DiagnosticResult(BaseModel):
    status: str = "SUCCESS"
    test_id: str
    device_id: str
    site_model: str
    root_cause: str
    evidence: List[str] = Field(default_factory=list)
    fix: List[str] = Field(default_factory=list)
    verification_matrix: List[VerificationClaim] = Field(default_factory=list)
    competing_hypotheses: Dict[str, HypothesisEvaluation] = Field(default_factory=dict)
    timeline: Optional[Dict[str, Any]] = None
    report: str = ""
    error: Optional[str] = None

5.7. Native Python API

The Workbench calls the agent in-process from a worker thread (mantis_adapter.py):

import threading
from mantis.agent import MantisAgent
from mantis.models import SessionContext

agent = MantisAgent()
context = SessionContext(active_site_model="sites/udmi_site_model")
cancel_event = threading.Event()  # set() to stop at the next step boundary

answer: str = agent.run(
    "How does the test scan_single_future work?",
    context,
    stream_callback=lambda text: print(text, end=""),  # interim prose
    event_callback=lambda record: None,                 # phases, tool calls, hypotheses, metrics
    cancel_event=cancel_event,
)

agent.run_tripartite(prompt, context, stream_callback) returns a dict {"status", "prompt", "final_answer", "metrics"}.


6. Contract 3: Sequencer Run & Log Streaming

6.1. Stream Wire Format

| Event | Data | | :— | :— | | log | {"offset", "text"} | | test_event | {"type": "started", "test"} or {"type": "result", "test", "variant", "result", "bucket", "stage", "score", "message"} | | heartbeat | {"offset"} | | complete | {"session_id", "exit_code", "stopped", "offset"} | | error | {"message", "offset"} (for example after the 1800 s idle timeout) |

GET /api/sequencer/stream?session_id=seq-1&offset=0 HTTP/1.1
Accept: text/event-stream

event: log
data: {"offset": 118, "text": "Starting sequence pointset_publish on AHU-1...\n"}

event: test_event
data: {"type": "result", "test": "scan_single_future", "variant": "scan_single_future+bacnet", "result": "pass", "bucket": "discovery.scan", "stage": "...", "score": "...", "message": "..."}

event: complete
data: {"session_id": "seq-1", "exit_code": 0, "stopped": false, "offset": 20480}

7. Contract 4: UI State & Workspace Contract (Client-Side State Store)

7.1. State Shape

export interface WorkspaceState {
  // Selections (persisted)
  siteModel: string;
  deviceId: string;
  projectSpec: string;               // Sequencer target, e.g. '//mqtt/localhost:18833' or '//gbos/...'
  logLevel: 'INFO' | 'DEBUG' | 'TRACE';
  minStage: 'PREVIEW' | 'ALPHA' | 'ALPHA_ONLY'; // default 'PREVIEW'
  serialNo: string;
  selectedTests: string[];
  bucketFilter: string;
  searchQuery: string;

  // Layout (persisted)
  logPanelHeight: number;
  localSetupDrawerWidth: number;
  mantisDrawerWidth: number;

  // Catalog & results
  siteModels: object[];
  devices: object[];
  sequences: object[];
  results: Record<string, object>;
  resultsDirExists: boolean;
  stagesAdmitted: string[];

  // Current run
  sessionId: string | null;
  running: boolean;
  startedAt: string | null;
  elapsedDuration: string;
  exitCode: number | null;
  statusLabel: string;
  testStatus: Record<string, object>;
  commandLine: string;
  activeTestId: string;
  consoleLogs: object[];

  // Local Test Setup drawer
  testbedDrawerOpen: boolean;        // declared, not updated
  testbedStatus: object;             // raw GET /api/testbed/status payload (snake_case), plus client spec_error / UNKNOWN
  pubberMode: boolean;
}

7.2. View Preservation

  1. /sequencer and /devices are each created on first visit and kept mounted.
  2. Switching routes toggles the hidden attribute on the panes; DOM nodes are not destroyed.
  3. The run log stream, the Mantis stream (in its drawer), filters and scroll position survive route changes and drawer toggles.

8. Contract 5: Compliance & Reporting

8.1. Compliance Data (GET /api/compliance?site_model=)

{
  "site_model": "sites/udmi_site_model",
  "devices": [
    {
      "device_id": "AHU-1",
      "has_results": true,
      "reason": null,
      "last_run": "...",
      "start_time": "...",
      "udmi_version": "...",
      "status_message": "...",
      "counts": {},
      "features": {},
      "stages": {},
      "verdict": "pass | fail | not_evaluated",
      "sequences": [],
      "unscored": [],
      "targets": [],
      "provenance": {},
      "reports": {}
    }
  ],
  "totals": {
    "devices": 1, "devices_with_results": 1,
    "pass": 0, "fail": 0, "skip": 0, "total": 0,
    "devices_passing": 1, "devices_failing": 0, "devices_not_evaluated": 0
  },
  "stages": [],
  "stages_for_pass": []
}

By design there is no aggregate score, compliance rate or flakiness index.

8.2. Reports & Bundles


9. Frontend UI Components

  1. Mantis drawer & conversation feed:
    • Renders the final answer’s Markdown (token), including Mermaid diagrams with a zoom/pan modal.
    • Collapsible Agent Reasoning accordion with phases, tool calls and thoughts.
    • Tool call badges: tool name, arguments (JSON viewer) and tool → summary.
    • Stop, Clear, Copy transcript, the failed-test selector, and Email me when done.
  2. Hypothesis Audit Matrix:
    • Status chips: PRIMARY, CONTRIBUTING, REFUTED, UNRESOLVED.
    • Evidence badges: LOCAL_FILE, CLOUD, NONE.
  3. Local Test Setup drawer (§3.2): Start Setup, Stop, Restart, the Pubber/Physical switch, the topology graph and log tabs.
  4. Sequencer filter toolbar: Min stage (PREVIEW, ALPHA, ALPHA_ONLY), feature bucket, Select all, Clear, the selection buttons ✓ Passed / – Skipped / ✗ Failed (select tests by previous result), search, and Email me when done.
  5. Site Roots consent modal: register, list and revoke external site model directories (§4.3).
  6. Artifact Viewer modal: RESULT.log, sequence.md, sequence.png and other artifacts, with Diagnose with Mantis and Export support bundle.
  7. Settings dialog: email notification consent, a test email, and recent deliveries only. The model provider is chosen when the gateway starts: the one-time bin/mantis setup --vertex=<project>[/<region>] (saved under mantis in ~/.config/udmi/workbench.json; the gateway refuses to start if the environment contradicts it), otherwise MANTIS_OFFLINE, then GEMINI_API_KEY / GOOGLE_API_KEY (AI Studio), then Vertex AI via ADC.
  8. Logs drawer: client diagnostic ring buffer and the server diagnostic log.

10. Not Yet Implemented

These are part of the Workbench vision and are not built: