Skip to content
s1ns3nz0 | Known Unknowns
Go back

Diagnosing OpenCTI with Kagent (2) - Tools, Code-First Diagnosis, and Read-Only Access

5 min read

Part 2 of the series. Part 1 covered the architecture: Agent, ModelConfig, system prompt, and MCP server.

1. The four tools

Order diagnosis: diagnose_paid_order

The input is a tenant_id and an order_id. The tool reads order-related state from the dedicated diagnostics API and classifies which stage the order is in.

/internal/v1/diagnostics/tenants/{tenant_id}/orders/{order_id}

agent/paid_scan_diagnostics.py defines each tool’s name, description, and input schema. The order tool’s contract is:

{
    "name": "diagnose_paid_order",
    "description": (
        "Read one tenant-scoped order projection to locate "
        "payment, dispatch or result stage. "
        "Database facts only; no automatic repair."
    ),
    "inputSchema": {
        "type": "object",
        "required": ["tenant_id", "order_id"],
        "additionalProperties": False,
        "properties": {
            key: {"type": "string", "format": "uuid"}
            for key in ("tenant_id", "order_id")
        },
    },
}

This tells the model: “Order diagnosis needs a tenant ID and an order ID. No other input, such as an arbitrary URL or SQL, is accepted.”

The call the model chooses looks like this:

{
  "name": "diagnose_paid_order",
  "arguments": {
    "tenant_id": "<tenant UUID>",
    "order_id": "<order UUID>"
  }
}

The MCP server maps that name to a Python function, runs it, and returns the result. The model does not write and run new Python code.

This tool checks the processing state recorded in the database. It does not connect to LND to re-verify settlement.

Workload status: get_opencti_workload_status

This tool reads the Deployments, scan Jobs, Pods, and Warning Events in the OpenCTI namespace.

Comparing the order record’s dispatch_job_name with the actual Job name connects the application’s record to the Kubernetes execution state.

The tool does not pass raw logs or full resource definitions to the model. It returns only the fields the diagnosis needs.

Payment path diagnosis: diagnose_l402_funnel

This tool sends fixed queries to Prometheus to check how Aperture is handling L402.

What it observesWhy
Invoice issuance successes and failures in the last 15 minutesWhether payment instructions can be issued
Requests without a token vs. invoices issuedRequests arriving while no invoice comes out
Authentication results and storage errorsProblems while verifying payment proofs
Authentication rejections compared with a past baselineA security signal that needs a closer look
Component health probesState of the pricing service, merchant LND, and Aperture

Missing metrics or a failed query are never treated as healthy. Conversely, the absence of L402 traffic alone is not treated as an outage.

The tool’s scope is L402. It does not prove MPP or x402 state, or that an individual order has been paid.

Playbooks: get_playbook

Given a playbook name, this tool returns the reviewed response procedure.

PlaybookCovers
opencti-paid-order-stuckPayment, scan, or result problems for a specific order
opencti-l402-funnelAperture, invoice issuance, and L402 authentication problems

It also returns the playbook’s revision and hash, so you can tell which version of the procedure the agent followed.

2. The code and the model split the diagnosis

Not every judgment is left to the model. The tool code validates the state and classifies the stage first.

For example, the order diagnosis function has these branches:

elif order != "paid":
    stage, check = (
        "payment",
        "inspect_payment_confirmation_and_receipt_commit",
    )

elif facts["receipt_commit_state"] == "pending" \
        or facts["challenge_state"] not in {None, "settled"}:
    stage, check = (
        "payment_records_need_review",
        "inspect_payment_confirmation_and_receipt_commit",
    )

elif scan is None:
    stage, check = (
        "scan_creation",
        "inspect_backend_reconciliation",
    )

Read in order:

The order is not paid yet
    → check the payment stage

Payment records are not fully applied
    → check payment record consistency

Passed both, but there is no scan record
    → check the scan creation stage

The model then connects this result to the playbook, calls more tools if needed, and explains it to the operator.

Raw state
    ↓
Tool code: validate and classify
    ↓
Structured result
    ↓
Model: further investigation, playbook comparison, explanation

When the API lookup fails or its response cannot be trusted, the tool states the uncertainty explicitly:

return dict(
    result,
    status="unknown",
    reason=reason,
    observed_facts={},
    next_check="restore_diagnostic_evidence",
)

The decision not to treat a failed lookup as healthy lives in the code, not in the model.

3. Kubernetes permissions match the tools’ scope

The Role given to the MCP server is:

rules:
  - apiGroups: [apps]
    resources: [deployments]
    verbs: [get, list]

  - apiGroups: [batch]
    resources: [jobs]
    verbs: [get, list]

  - apiGroups: [""]
    resources: [pods]
    verbs: [get, list]

  - apiGroups: [""]
    resources: [events]
    verbs: [get, list]

Only get and list are allowed. This Role cannot modify a Deployment, create a Job, or delete a Pod.

A RoleBinding in the OpenCTI namespace connects this Role to the MCP server’s ServiceAccount. The agent runtime is not given cluster administration rights.

The main Kubernetes resources are:

ResourceRole
AgentConnects the model, prompt, and tools
ModelConfigModel connection settings
RemoteMCPServerRegisters the tool server’s address
Deployment, ServiceRun and expose the MCP server
ConfigMapSupplies the Python code and playbooks
SecretHolds the diagnostics API token and CA
ServiceAccount, Role, RoleBindingIdentity and permissions for lookups
NetworkPolicyLimits the allowed network paths

4. How a real question is handled

Suppose an operator asks:

“This order was paid, but there are no scan results.”

The agent reads the matching playbook and looks up the order. If it needs to confirm execution, it also checks Kubernetes.

get_playbook
    ↓
diagnose_paid_order
    ↓
get_opencti_workload_status
    ↓
Compare the order record with the actual Job state
    ↓
Explain observed facts, possible causes, and what to check next

For example, if the order is paid and its Job’s Pod is still Pending, that is evidence of a problem at the execution start stage.

But that alone does not prove a memory shortage or an image error. Anything not backed by further evidence should be reported as unconfirmed.

5. Where to change what

What you want to changeWhere
Diagnosis target API and namespacepaid-scan-wsl-e2e.values.yaml
Agent instructions and answer styletemplates/_helpers.tpl
Tools the agent may usetoolNames in templates/paid-scan.yaml
MCP server connection, deployment, permissionsRemoteMCPServer, Deployment, and RBAC in the same file
Actual lookup and diagnosis logicagent/paid_scan_diagnostics.py
Model and Ollama connectionThe deploy configuration read by ops/deploy-agent
Operational response proceduresDocuments in docs/playbooks/

Adding a new diagnostic capability usually goes in this order: implement the Python tool, register it in the tool list, then add usage instructions to the agent. Connection settings and permissions are extended only when a new data source is needed.

Next: Diagnosing OpenCTI with Kagent (3) - Designing Access Control in a GitOps Setup


Share this post:

Previous Post
Diagnosing OpenCTI with Kagent (3) - Designing Access Control in a GitOps Setup
Next Post
Diagnosing OpenCTI with Kagent (1) - Architecture: Agent, Model, and MCP Server