How to Choose the Right Webflow Development Agency in Seattle (2026 Guide)

What to compare, what to ask and what a fair price looks like before you sign with a Webflow partner.

Published September 10, 20266 minute readENGINEERING
Your tools work. Will the agent use them right?
Key takeaways⌄
  • A passing tool call only proves the tool returned valid JSON, not that an agent used it correctly.
  • The harness runs real agents against disposable sites and checks both the outcome and the path taken.
  • Every run produces findings, so passing stories still surface specific product feedback.
  • Repeated runs and non-visual signals like semantic_html_ratio catch problems a single score hides.
  • Testing only with Claude missed Codex behavior that led to a rejected ChatGPT app submission.

Inside the MCP eval harness

A green checkmark in your agent doesn't tell you an agent will do the right thing with your tools. It tells you the tools returned valid JSON. Those are different claims, and the gap between them is where MCP servers actually fail in front of a real user, with a real agent making its own decisions. And sometimes those decisions are pretty bad when customers rely on their brand experience to be powered by a product.

In this post, we'll talk about how we build more confidence in how our MCP server behaves in agents we don't control.

The problem with testing an MCP server

A Model Context Protocol (MCP) server hands AI agents, whether Claude, ChatGPT, Cursor or whatever a person connects with, a set of tools for building pages, managing CMS content, and publishing sites. Traditional tests can establish that each tool is correct in isolation: it accepts the right inputs, enforces permissions, mutates the expected state, and handles failures properly. But they cannot answer a different question: when an autonomous agent sees the full tool surface and receives a plain-English task, will it discover and combine those capabilities correctly?

You only get an answer by watching an agent behave. So that's what our eval harness does. It drives real agents headlessly against real, disposable sites and scores what the agent chooses to do, not just whether the tools it called returned valid responses.

What it actually does

The harness runs stories: multi-step agentic tasks written as plain prompts, each with assertions about what should happen. A story might be "create a CMS collection with three fields and one entry" or "build a full SaaS landing page headlessly and publish it."

For each story, the harness provisions a fresh site and runs a real agent against it turn by turn. Deterministic checks verify expected and forbidden tool use and, where relevant, verify that a claimed write actually took effect. An LLM judge reads the transcript and final answer for task completion, tool selection, and correctness. Page-building stories add a visual judge over the rendered result, while structural signals can inspect what the generated content contains.

Every trace, score, finding, and screenshot ships to a dashboard so results are queryable and comparable over time. That's the skeleton. The interesting part is what we learned once a single pass/fail number stopped being enough.

Eval harness diagram

We separate outcome correctness from execution-path diagnostics. First, did the intended state change actually happen? Second, did the agent violate any genuine constraint, for example using a destructive tool or a capability unavailable in that host? Finally, did it take an unexpectedly difficult path that suggests the tool surface could be clearer? A different path is not necessarily wrong, but it is often worth understanding.

Every run leaves a punch list, not just a score

The same LLM judge call that scores a story also produces findings: concrete, actionable problems it noticed in the transcript whether or not the story passed. A tool call failed. The agent worked around a confusing capability. A missing capability had to be faked. A story can pass every assertion and still surface findings, because "the agent got there in the end" and "the tools made that easy" are different questions.

A passing story can still expose a bad tool experience

A few findings pulled straight from real runs:

  • data_whtml_builder rejects multiple root elements per call, forcing the agent to split three sections into separate inserts and reverse-order them; the tool should support multi-root fragments.
  • data_style_tool returns responses over 71 KB that exceed token limits, forcing the agent to shell out to Bash to parse them.
  • Auto-renamed CSS classes on collision (hero becomes hero-1-2-3), silently breaking the agent's intended class name with no warning from the tool.

None of those are pass/fail bugs. Every story still completed. But each is specific product feedback surfaced by an agent actually hitting the papercut, not by someone guessing what might go wrong. Those findings live in the dashboard alongside the score, so a green checkmark no longer means the path the agent took was clean.

What does a single run hide?

A single successful run can hide meaningful variation, so the harness can execute the same story repeatedly with a fresh site and agent session each time. This is initially a diagnostic tool, not a claim of statistical reliability. In one early three-run diagnostic, deterministic assertions passed in all three runs, while design-quality scores were 40%, 80%, and 85%. Those observations are not enough to estimate the underlying consistency distribution, but they immediately revealed something the original summary concealed: tool execution was stable while output quality was not.

The score that looked stable was the wrong score

The original console summary reported mean 1.00 and stddev 0.00 because the aggregation only saw the deterministic score. That result was mathematically correct and operationally misleading. We now report those dimensions separately and preserve the raw run-level results: deterministic checks passed 3/3 on all three runs, design quality scored 40%, 80%, and 85%. As we increase sample sizes, the same infrastructure can support more defensible reliability and regression analysis.

The obvious signal isn't the only signal

An LLM judge scoring a screenshot only sees the screenshot. Two pages can render pixel-identical while one is built from meaningful structure (nav, section, button) and the other is div soup all the way down. That is a real accessibility and maintainability difference, and it is invisible to an image.

Pixel-identical does not mean structurally identical

We added semantic_html_ratio, a deterministic, non-LLM signal computed from the rendered HTML already captured for the screenshot: the fraction of semantic tags versus generic wrappers. Combined across pages by mode, we can see not just how a page looks but how it was built. It's not a pass/fail gate. It's a second, independent signal for comparing runs and watching trends over time.

Signal
What it checks
Output
Deterministic checks
Expected and forbidden state changes on the site
Pass / fail
LLM judge score
Whether the story goal was met, from the transcript
Score
LLM judge findings
Concrete problems noticed during the run
Punch list
Repeated runs
Variation across fresh sites and sessions
Mean, stddev
semantic_html_ratio
Semantic tags versus generic wrappers in rendered HTML
Trend signal
Signals the harness records for each story run.
StepStageRuns againstDeterministicUses LLMOutput
01ProvisionFresh disposable siteYesNoClean site state
02Run agentClaude Code or Codex CLINoYesTurn-by-turn transcript
03Check stateExpected and forbidden changesYesNoPass / fail
04JudgeTranscript and screenshotNoYesScore and findings
05Measure structureRendered HTMLYesNosemantic_html_ratio
06ReportDashboardYesNoQueryable results
What happens in a single story run, in order.

Agent-host coverage is part of eval coverage

This is the one that mattered most in practice. OpenAI rejected one of our ChatGPT app submissions after an agent tried to use a Designer-canvas tool that requires a live Designer tab, something a ChatGPT session can never have. We ran the exact rejected prompts through our harness repeatedly and never reproduced the failure. Not once.

The reason was uncomfortably simple: the harness had only ever been driven by Claude, and Claude essentially never took the path from that piece of stale tool guidance. The production rejection came from a GPT-family model. We had built a real eval harness and never ran it against the host that was actually failing in production.

So we added a second driver that shells out to OpenAI's Codex CLI the same way the original shells out to Claude Code. Both share one orchestration core: provisioning, scoring, judges, screenshots, retries, and upload are identical. Then we ran a controlled comparison: the original prompt versus one hardened with an explicit "don't use Designer-canvas tools" instruction, five independent repeats each, on both drivers.

Same prompts, different agent-host behavior

The same task and MCP server produced materially different behavior across Claude Code and Codex CLI. This experiment doesn't isolate the underlying model from the surrounding host orchestration, and operationally, it doesn't need to. Users experience the combined system: model, system instructions, tool presentation, context management, and execution policy. The lesson wasn't that one model was better. It was that validating against a single agent host gave us confidence that did not transfer to another environment. Every confidence-building feature above, including repeat testing, semantic signals, and transcript judging, now runs identically against either host because those capabilities live in the shared core rather than Claude-specific code.

Looking for a Webflow-certified developer to build a website?
Talk to a Webflow Expert

Why this matters going forward

Chasing a green checkmark was never the point. Before we ship, we can say something specific and true: we ran this exact task, multiple times, against multiple real agents, and here's what actually happened. That's a different kind of confidence than tests passing.

  • Every tool-schema change runs through the story suite before it ships. The same stories can point at a locally running MCP server, so regressions can be caught on a pull request rather than after deploy.
  • Every ChatGPT app resubmission is checked against both drivers first, so a host-specific failure shows up in our harness before it shows up in a review.
  • A flaky story shows up as flaky, with variance, instead of getting lucky in CI and surprising someone later.

The suite now runs nightly, turning drift into a number that moves every day rather than a production complaint someone has to trace backward weeks later.

Running it has gone from knowing the incantation to asking, and getting an answer.

Where it's headed

The near-term work is operational: automatically flag tools with no story coverage, and turn each nightly run into a report of regressions, new findings, and notable variance that tools owners can act on directly.

The harder question is what a score actually means. Today every story's threshold, including what counts as "good", comes from whoever wrote it. Useful, but invented. We want a gold-reference ceiling hand-built on ideal versions of a flagship story, score it through the same judge, and express future runs relative to that ceiling instead of treating a bare 0-to-1 score as self-explanatory.

And the story set itself should not stay frozen at launch. Right now it reflects what we imagined users might ask. Over time, we want new stories grounded in real usage patterns and analytics, so our user data, not our guesses, determines what the harness tests. That's the version of this harness worth building toward: not just catching what we thought to check for, but learning what we should be checking.

Last updatedSeptember 10, 2026
CategoryEngineering
Gil Levin
AuthorGil LevinSenior Software Engineer
Not ready to book a call yet?

The Webflow Enterprise Evaluation Checklist (PDF)

What to ask Webflow sales, what your security team will need, and what to settle about roles before anyone starts designing.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

The execution framework, nothing else.