A self-evolving external harness for RCA

OpsHarness: A Self-Evolving Harness for Root Cause Analysis

OpsHarness turns diagnosis experience into reusable expertise—then verifies every evolution before it reaches the live harness.

Haiyu HuangJiewei LyuZhihan JiangJinyang LiuXiao HeTieying ZhangWu XiangMichael R. Lyu

The Chinese University of Hong Kong · ByteDance

Paper arXiv Code · Coming Soon Demo · Coming Soon
BeforeGeneral
Agent
strong reasoning · limited system context
Ops
Harness
AfterRCA
Expert
system knowledge · verified evolution
0159.0%Top-1 accuracy
02+63.4%vs. a bare general agent
034.02×vs. baseline RCA agents

Two public benchmarks + one industrial deployment

01 — The RCA gap

General models are capable.
The RCA gap is a harness that learns.

Modern models already provide strong reasoning, planning, and tool use. What RCA still lacks is an external harness that adapts those capabilities to each system—and keeps improving as incidents reveal new patterns.

The foundation

General capability is no longer the bottleneck

Modern models can reason, plan, and use tools, but they arrive without the context and operating experience of the target system.

Capable model · missing system context
The missing layer

RCA adaptation belongs in the harness

Put RCA skills, system knowledge, observability, and diagnosis tools around the model instead of rebuilding its general machinery.

General model · system-specific harness
OpsHarness

Static adaptation must become verified self-evolution

Distill successful and failed diagnoses into atomic updates, then test every proposal before it reaches the live harness.

Learn · verify · promote

RCA superpowers

Specialization belongs outside the model.

OpsHarness preserves the model’s general capabilities and supplies the external skills, knowledge, observability, verification, and learning loop that RCA requires.

Figure 1OpsHarness specializes a general-purpose agent through an external RCA harness.
Figure 2Human-SRE motivation: turn one corrected incident into reusable guidance for the next.

Why self-evolution matters

One corrected incident should improve the next.

A misleading symptom first sends the investigation down the wrong path. Once the real propagation chain and its caveat are recorded, a later incident can be localized with fewer wrong turns.

OpsHarness turns this manual learning loop into a reviewable, verified system process.

02 — The architecture

A harness with
memory and control.

OpsHarness separates what the agent knows from how that knowledge is acquired, tested, and promoted.

Per-system state

Data Plane

System-specific expertise is organized for progressive disclosure, so the agent loads only what each diagnosis needs.

  • K0
    General RCA knowledgePrinciples, constraints, and core workflows
  • K1
    System profileSchema, components, telemetry, and SOPs
  • K2
    Mined workflowsReusable diagnosis paths learned from use
  • K3
    Operations & rulesAtomic skills, evidence patterns, and caveats
Tool idea cardz-score anomaly detectioninput → procedure → output
Lifecycle orchestration

Control Plane

Four workflows carry a system from cold start to an expert harness while keeping every self-modification reviewable.

  1. 01
    SetupProfile a new telemetry system
  2. 02
    DiagnoseInvestigate and rank root causes
  3. 03
    EvolveMine correct and failed trajectories
  4. 04
    VerifyGate changes before promotion
Observability records every run; feedback closes the loop.
01Setupadapt to the system
02Diagnoseproduce evidence
03Evolvepropose improvements
04Verifypromote only if safe
Figure 3The complete OpsHarness data plane, control plane, and end-to-end usage flow.

Verified self-evolution

Learn from both success and failure—without learning the wrong lesson.

OpsHarness compares successful and failed trajectories, turns their evidence into atomic proposals, and evaluates a staged harness before any update reaches production.

  1. 01
    Mine trajectoriesContrast useful and failed diagnostic paths.
  2. 02
    Propose atomic updatesCreate reviewable skills, rules, and caveats.
  3. 03
    Pass two gatesImprove source cases and avoid regression on held-out cases.
Figure 4The staged evolution pipeline mines evidence, proposes atomic changes, and applies only updates that pass both gates.

03 — The evidence

It improves with use.
The gate makes it reliable.

Across OpenRCA, RCAEval, four model backbones, two agent frameworks, and an industrial deployment, the gains are consistent and practical.

Overall effectiveness59.0%

Final A@1 for full OpsHarness, versus 41.4% without evolution and 36.1% for Direct.

Learning with use0.83 vs 0.43

Final-window A@1 for full OpsHarness versus the non-evolving harness.

Verification matters37%

of evolution proposals are rejected; unverified evolution finishes at only 0.33 A@1.

Industrial deployment0.74 vs 0.24

Average A@1 for OpsHarness versus Direct across six production configurations.

Complete benchmark results

24 configurations across two benchmarks and six sub-datasets.

OpsHarness is evaluated with four model backbones. Specialized RCA agents are marked †; within each backbone, the OpsHarness row is emphasized.

Scroll horizontally to view all metrics →

Overall effectiveness of 24 configurations across OpenRCA and RCAEval.
Framework OpenRCA RCAEval Final
A@1
Telecom Bank Market Online Boutique Sock Shop Train Ticket
A@1A@3Avg A@1A@3Avg A@1A@3Avg A@1A@3Avg A@1A@3Avg A@1A@3Avg
GPT-5.5
RCA-Agent36.445.555.035.764.365.023.253.655.00.011.146.05.65.629.00.00.024.016.8
mABC9.118.214.010.721.435.03.19.832.016.733.376.011.122.261.05.611.161.09.4
Codex (Direct)54.554.564.057.164.372.023.246.459.066.783.395.061.177.889.044.472.283.051.2
Codex (ICL)27.336.450.046.453.665.020.040.049.055.677.893.061.161.185.061.188.994.045.3
OpsHarness (no-evolve)54.554.564.046.457.162.026.843.358.066.772.291.072.272.291.050.077.889.052.8
OpsHarness72.772.777.064.271.478.037.166.572.072.288.996.077.888.993.072.294.496.066.0
Claude Sonnet 4.6
RCA-Agent9.118.232.053.667.979.016.533.555.011.122.259.05.65.629.00.05.644.016.0
mABC0.00.00.03.625.040.03.16.730.011.133.376.011.111.159.05.622.269.05.8
Claude Code (Direct)36.445.555.024.236.361.017.026.338.038.961.185.027.844.476.016.722.276.026.8
Claude Code (ICL)45.554.564.021.439.358.026.736.749.544.472.289.027.844.472.038.955.676.034.1
OpsHarness (no-evolve)36.463.673.039.346.461.028.635.750.038.977.889.033.350.079.050.072.289.037.8
OpsHarness63.663.673.057.163.781.935.764.272.061.183.395.055.677.889.061.177.893.055.7
GLM-5.2
RCA-Agent36.454.559.060.764.377.013.423.740.516.722.256.016.716.733.00.00.028.024.0
mABC0.09.114.00.017.934.06.76.720.511.133.374.011.111.152.00.011.159.04.8
Codex (Direct)54.563.668.057.157.167.026.850.065.038.977.893.044.472.289.055.672.287.046.2
Codex (ICL)36.436.445.046.446.455.016.723.336.461.183.395.050.077.891.050.077.887.043.4
OpsHarness (no-evolve)54.563.680.057.164.374.028.642.952.066.788.989.044.455.674.055.688.994.051.2
OpsHarness72.781.887.057.164.374.042.950.061.077.888.989.055.666.776.088.994.498.065.8
DeepSeek-V4
RCA-Agent45.545.550.028.635.747.09.89.826.05.65.635.00.05.626.00.00.022.014.9
mABC0.00.05.03.614.332.00.03.116.011.133.372.00.05.643.00.00.048.02.5
Codex (Direct)18.227.341.032.139.352.010.323.239.027.844.476.022.227.863.011.122.259.020.3
Codex (ICL)36.436.441.028.639.350.020.030.042.861.172.291.022.233.361.016.738.967.030.8
OpsHarness (no-evolve)27.245.546.035.746.461.021.428.650.044.453.376.027.850.080.033.366.772.031.6
OpsHarness45.563.668.046.450.064.028.642.960.053.373.380.050.072.290.066.777.889.048.4
Finding 1Verified evolution steadily raises accuracy across all four backbones.
Finding 2The gains transfer to a real production system and an open agent stack.
Finding 3Per-case cost stays comparable to Direct and ICL while accuracy rises.

04 — Usage preview

One harness.
Two agent interfaces.

OpsHarness keeps the same skills, tools, knowledge, and ground-truth guard across Codex and Claude Code. Only the invocation syntax changes.

Code release coming soon.
The commands below preview the public workflow described in the paper and implementation.
opsharness / workflow

Codex

# Adapt to a telemetry system
$setup

# Diagnose an incident
$diagnose "2021-03-04 18:00 to 18:30,
service latency spike"

# Learn from feedback
$evaluate <sessions> "/path/to/ground-truth"
$evolve "last 10 diagnoses"

Claude Code

# Adapt to a telemetry system
/ops-harness:setup

# Diagnose an incident
/ops-harness:diagnose "2021-03-04 18:00 to 18:30,
service latency spike"

# Learn from feedback
/ops-harness:evaluate <sessions> "/path/to/ground-truth"
/ops-harness:evolve "last 10 diagnoses"

05 — Paper

Abstract

Automated root cause analysis (RCA) with large language models (LLMs) has drawn growing attention. Today, SREs typically automate RCA with LLMs in one of two ways: directly using a general-purpose agent (e.g., Codex or Claude Code) for diagnosis, or building a specialized RCA agent from scratch. As mainstream general agents grow more capable and iterate quickly, our quantitative study finds that the former now often surpasses the latter.

Its accuracy, however, still falls short of production needs, and this gap stems mainly from the external adaptation layer outside the agent's general capabilities, namely the harness. We therefore argue that LLM-based RCA should focus on this external harness, reusing the strong general capabilities of a modern agent rather than rebuilding an agent from scratch.

A key capability of such a harness is to self-evolve, accumulating system-specific experience from past diagnoses so that it gets better the more it is used. We introduce OpsHarness, a self-evolving RCA harness that turns diagnosis experience into reusable expertise. Its data plane combines layered operational knowledge with an idea-card tool library, while its control plane coordinates setup, diagnosis, evolution, and verification.

During evolution, OpsHarness contrasts successful and failed trajectories, converts their evidence into atomic proposals, and admits updates only through a dual-gate verification process designed to prevent overfitting and regression. Across two public benchmarks and an industrial deployment, OpsHarness achieves 59.0% top-1 accuracy, improving over a bare general agent by 63.4% and over baseline RCA agents by 4.02×.

06 — Cite the work

BibTeX

If OpsHarness is useful in your research, please cite the arXiv preprint.

@misc{huang2026opsharness,
  title={From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis},
  author={Haiyu Huang and Jiewei Lyu and Zhihan Jiang and Jinyang Liu and Xiao He and Tieying Zhang and Wu Xiang and Michael R. Lyu},
  year={2026},
  eprint={2608.25661},
  archivePrefix={arXiv},
  primaryClass={cs.SE},
  url={https://arxiv.org/abs/2608.25661}
}