FrameworkAnthropic

Expert Workflows Need Domain-Native Harnesses

The durable product-builder skill is no longer asking whether a frontier model can “do expert work” in the abstract. It is building domain-native harnesses where the model gets the right artifacts, reasoning loop, and expert verification for one specialized job.

What Changed

The strongest July 28 signal came from Digg Tech and Simon Willison surfacing Anthropic’s report on using Claude Mythos Preview to discover new cryptographic weaknesses. The immediate story is impressive, but the builder lesson is not “AI replaces cryptographers.” Anthropic shared the prompts, task framing, and human-research context, showing that the model was useful inside a carefully shaped workflow, not a blank general chat. That is the meaningful shift. In advanced domains, product value comes less from exposing raw model access and more from translating specialist practice into a harness the model can operate within.

Why Product Builders Should Care

Teams often fail in expert markets by shipping a generic AI shell into work that depends on notation, artifacts, review conventions, and asymmetric risk. The result is impressive demos that collapse under real scrutiny. Builders who create domain-native harnesses will unlock narrower but far more durable products because they make the model legible to the expert and the expert legible to the model. That is where adoption starts.

How To Use This

Turn one specialist workflow into a domain-native harness. Trigger: a repeated expert task such as legal issue spotting, growth-diagnosis review, security triage, scientific literature synthesis, or financial anomaly analysis. Context: load the domain’s native artifacts, notation, prior work, and acceptance criteria instead of generic prose summaries. Tools: give the model the exact calculators, datasets, structured forms, or scratch surfaces specialists already use. Verifier: require expert review or a domain-specific acceptance test that can reject impressive-looking but invalid work. Budget: set research depth, model attempts, and expert-review time before the workflow starts so the product stays economically legible. Artifacts: save the prompt frame, intermediate work product, verifier result, and final specialist-facing output. Stop condition: the loop ends only when the artifact is valid in the domain’s own terms, not when the model merely sounds confident.

Practice Drill

Pick one expert task your product touches and write down its native inputs, tools, proof standard, and final artifact. If the workflow still looks like “paste docs into chat,” it is not a domain-native harness yet.

What could make this wrong

For lightweight consumer tasks with loose acceptance criteria, a fully specialized harness may be more overhead than value.

Confidence · medium

The direct evidence is a single high-signal primary case study, but Simon Willison’s emphasis on the shared prompts and workflow structure supports the broader product inference that harness design, not raw chat access, makes expert use viable.

Revisit · Aug 4, 2026

Did translating the workflow into domain-native inputs and checks improve expert trust and usable output quality?

Watch: expert acceptance rate · time to valid result · manual corrections after AI output · usage concentration in one specialist workflow

Apply it now

Knowledge only counts when it changes the build.

Pick one expert task your product touches and write down its native inputs, tools, proof standard, and final artifact. If the workflow still looks like “paste docs into chat,” it is not a domain-native harness yet.

Stage
shape
Produce
Domain-native harness spec for one expert workflow

Full context at Anthropic. Bring back one decision, test, or workflow change.

Read the original ↗

Keep Going