<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:media="http://search.yahoo.com/mrss/"><channel><title><![CDATA[Zuper Labs]]></title><description><![CDATA[Experiments with AI & Technology]]></description><link>https://labs.zuper.co/</link><image><url>https://labs.zuper.co/favicon.png</url><title>Zuper Labs</title><link>https://labs.zuper.co/</link></image><generator>Ghost 5.88</generator><lastBuildDate>Sun, 13 Sep 2026 19:56:50 GMT</lastBuildDate><atom:link href="https://labs.zuper.co/blog/rss/" rel="self" type="application/rss+xml"/><ttl>60</ttl><item><title><![CDATA[Inside Sense: how an agent answers, investigates, and computes what nobody built]]></title><description><![CDATA[<p><em>What&apos;s holding up this job?</em> For WO-123, the answer:</p>
<blockquote>
<p>WO-123 has been open eighteen days. The quote was approved on the 1st. The material request for the special-order shingles went out on the 4th, and its purchase order is still waiting on vendor confirmation. The technician&apos;s</p></blockquote>]]></description><link>https://labs.zuper.co/blog/inside-sense-how-an-agent-answers-investigates-and-computes-what-nobody-built/</link><guid isPermaLink="false">6a6af9fe5a2dfd3abfb65e23</guid><dc:creator><![CDATA[Ankush S]]></dc:creator><pubDate>Thu, 10 Sep 2026 10:27:05 GMT</pubDate><media:content url="https://labs.zuper.co/content/images/2026/09/b1cd8baf-f45a-413c-95da-a894caed9230.png" medium="image"/><content:encoded><![CDATA[<img src="https://labs.zuper.co/content/images/2026/09/b1cd8baf-f45a-413c-95da-a894caed9230.png" alt="Inside Sense: how an agent answers, investigates, and computes what nobody built"><p><em>What&apos;s holding up this job?</em> For WO-123, the answer:</p>
<blockquote>
<p>WO-123 has been open eighteen days. The quote was approved on the 1st. The material request for the special-order shingles went out on the 4th, and its purchase order is still waiting on vendor confirmation. The technician&apos;s last note, from the 7th, reads &quot;waiting on material, customer informed.&quot; Two service tasks are done; the install step is blocked on that material.</p>
</blockquote>
<p>A dispatcher who wants that paragraph opens ten screens: the job, its status history, the notes, the time logs, the quote, the material request, the purchase order, the vendor, the customer, the checklist. Ten to twenty minutes, if they already know where to click. Most people never do it. They ask the technician, or they guess.</p>
<p>Sense, the agent built into Zuper, wrote that paragraph in ninety seconds from that one question.</p>
<p>This post is about the context engineering that makes an answer like that possible, and correct. It goes somewhere we did not expect: the agent computing metrics the product never explicitly built, from a company&apos;s own operations manual.</p>
<video autoplay muted loop playsinline preload="auto" style="width: 100%; height: auto;">
  <source src="https://labs.zuper.co/content/media/2026/09/Sense-Demo.webm" type="video/webm">
</video>
<hr>
<h2 id="three-kinds-of-question-one-box">Three kinds of question, one box</h2>
<p>Field service companies ask three kinds of question all day.</p>
<p>The first is <strong>measured</strong>. <em>What was revenue by team last quarter? How many jobs are past due?</em> Harder than it looks, because the vocabulary is local: &quot;Booked&quot; is a status at one company and a category at another, and revenue can mean job revenue, invoiced amount, or approved quote value. Choose wrong and the number is confidently, invisibly incorrect.</p>
<p>The second is <strong>investigated</strong>, and almost nothing answers it well. <em>Why is this job stuck? Why did this invoice take three weeks to go out?</em> The answer is never in one place. It is a chain of facts across a quote, a material order, a vendor, and a note somebody typed at 6pm.</p>
<p>The third is <strong>documented</strong>. <em>How does this work here? How do we define first-time fix?</em> The answer lives in the company&apos;s own manuals and policy pages.</p>
<p>Sense answers all three in one conversation, and most real questions are a blend: <em><strong>show me overdue invoices, and tell me why the top three are stuck</strong></em> is a measured question and an investigated one in the same sentence.</p>
<p>One principle runs through everything below: <strong>every mechanism exists because a specific wrong answer is possible without it.</strong></p>
<p><img src="https://labs.zuper.co/content/images/2026/09/sense-title-1.png" alt="Inside Sense: how an agent answers, investigates, and computes what nobody built" loading="lazy"></p>
<hr>
<h2 id="anatomy-of-an-investigation">Anatomy of an investigation</h2>
<p>Back to WO-123. What the user watches while it works:</p>
<pre><code class="language-text">Looking up work order 123
Reading the job                  &#x2192; 18 days open, install pending
Checking the linked quote        &#x2192; approved on the 1st
Following the material request   &#x2192; shingles, raised on the 4th
Checking the purchase order      &#x2192; sent, awaiting vendor confirmation
Reading technician notes         &#x2192; last update on the 7th
Checking service tasks           &#x2192; 2 of 3 done, install blocked
</code></pre>
<p>Nobody handed the agent that sequence. It works out the trace itself, because it carries a working model of how the business fits together: a job belongs to a customer and a property, a quote usually precedes it, an invoice follows execution, a material request is fulfilled by a purchase order or a stock transfer, and service tasks are the job&apos;s sub-steps. Told to find what is holding up a job, it knows where to look. A typical investigation runs in about a minute to a minute and a half, end to end.</p>
<p>It works over a read-only, allowlisted slice of the <a href="https://labs.zuper.co/blog/how-we-built-the-zuper-mcp-intent-driven-testing-module-wise-coverage-and-closing-100-schema-gaps/">Zuper MCP</a> surface we wrote about separately: <strong>48 read-only tools across 14 entity families</strong>. The finding is relayed to the user in full rather than re-summarized, because a trace that reaches the user with most of its facts summarized away is a loss, not polish.</p>
<p>Two problems make this harder than looping over an API.</p>
<p><strong>Records are enormous.</strong> One fetch from a live operational API can return megabytes of nested data, and almost none of it bears on the question. So any response over <strong>50KB</strong> never reaches the model whole. It is parked in a short-lived store, and the investigator gets a preview capped at <strong>4KB</strong>: the record&apos;s real field names, the shape of the data, and a small sample. It then names the exact paths it wants and reads only those, under a budget, so an investigation converges instead of wandering. This is accuracy engineering before it is cost engineering: a model reasoning over a compact preview finds the blocked purchase order; one buried in a raw payload finds whatever surfaced last.</p>
<p><strong>Tools invite guessing.</strong> The operational tools expose thirty to seventy optional parameters each, often with example values sitting right there in their schemas, and a guessed filter does not fail loudly; it silently changes the answer. So only parameters with a real, concrete value get set, usually a single identifier that scopes the request. Words in the user&apos;s question like &quot;delayed&quot; or &quot;urgent&quot; describe the output to produce; they are not permission to set matching filters.</p>
<p>The boundary runs the other way too. This path never counts or aggregates. Even a single-customer count goes through the analytics path, because that path can render it, pin it, and refresh it. Explaining one record and measuring many are different jobs.</p>
<hr>
<h2 id="anatomy-of-a-measured-question">Anatomy of a measured question</h2>
<blockquote>
<p><em>Revenue by team last quarter for the Inspection Ready residentail jobs.</em></p>
</blockquote>
<p><strong>Vocabulary first.</strong> &quot;Inspection Ready&quot; is not in any product catalog; it is a status this company invented. And &quot;residentail&quot; is a typo. Before anything gets built, the agent looks the unfamiliar phrases up against what this workspace actually stores:</p>
<pre><code class="language-text">### &quot;Inspection Ready&quot; &#x2192; job_status
Exact match: Inspection Ready  (category: Field Work)
Filter on: jobs.current_job_status_name    (current state)
  &#xB7; job_status_history.status_name         (history / transitions)
  &#xB7; job_status_checklist.status_name       (answers captured at that status)

### &quot;residentail&quot; &#x2192; custom_field_value
Field: Property Type (jobs)
No exact match. Closest stored values:
  - Residential
  - Residential - Multi
&#x2192; confirm before filtering
</code></pre>
<p>The first block settles the status and raises the real question: three fields carry it, and jobs <em>currently</em> in Inspection Ready are not the same set as jobs that <em>passed through</em> it, so the choice is made from the tense of the question. The second block is the other half of the rule: a close match is never used automatically. &quot;residentail&quot; becomes a clarifying question with options that are real stored values, not invented plausible ones.</p>
<p>Here is why that matters. Suppose the lookup finds nothing, and the system shrugs and drops the filter:</p>
<pre><code class="language-text">Q: how many jobs for Acme Property Management are in production?
   [ &quot;Acme Property Management&quot; unresolved, filter dropped ]
   [ &quot;production&quot; unresolved, filter dropped ]

   4,812 jobs        &#x2190; every job in the company, rendered as the answer
</code></pre>
<p>That number is wrong, specific, and beautifully formatted. Nobody catches it, because a stat card carries no evidence of what it left out. When every term in a question fails to resolve, <strong>Zuper Sense</strong> does not run the query. It says which words it could not find and asks.</p>
<p>The other half of that defense is that every rendered number can be opened. Tap a bar on a chart or the number on a stat card and the underlying records open as a table. A stat card you can open is a number you can audit, and the fastest way to notice a filter you did not expect is to see the rows behind the figure.</p>
<p><strong>Then a specification, not SQL.</strong> Once the status is pinned to a field and the user confirms Residential, the agent describes the result it wants in business terms: revenue and job count, broken down by team, for last quarter, filtered to jobs currently in Inspection Ready on Residential properties, ordered by revenue, top 20, rendered as a bar chart.</p>
<p><strong>Checked against reality.</strong> Every field named in that spec is verified against the definitions actually available for the selected entities. Near-misses snap to the real name. Fields in the wrong slot get moved. Date operators on non-date fields, placeholder values, internal identifiers, duplicates, and impossible shapes get stripped or corrected before execution.</p>
<p><strong>Executed, shaped, rendered.</strong> Results come back with display names resolved to the company&apos;s own labels, then get shaped for the visualization: a single number becomes a stat card, a time axis becomes a line, and a breakdown over time stacks. A turn that ends on text alone is treated as incomplete, so a missing chart gets recovered rather than shipped as a paragraph the user cannot see. End to end, a measured question answers in about <strong>10.6 seconds on average</strong>. The median is around <strong>9 seconds</strong>, with a tail of about <strong>18.6 seconds at p90</strong> and <strong>24.8 at p95</strong>.</p>
<hr>
<h2 id="metrics-the-product-never-shipped">Metrics the product never shipped</h2>
<p>Here is the part we did not anticipate.</p>
<p>Companies write down how they measure themselves. It sits in an operations manual, an onboarding doc, a policy page. And those write-ups are often precise enough to compute from.</p>
<blockquote>
<p><strong>Q:</strong> what&apos;s our first-time fix rate this quarter?<br>
<strong>Sense:</strong> Your operations manual defines first-time fix rate as jobs closed on the first visit divided by all closed jobs, excluding warranty callbacks. On that definition it is <strong>82.4%</strong> this quarter, up from <strong>78.1%</strong> last quarter.<br>
<em>Sources: Field Operations Manual</em></p>
</blockquote>
<p>There is no first-time-fix-rate measure in the catalog. Nobody built one. A page of prose was the specification, and the agent read it, carried the definition into the data query, and computed the company&apos;s own metric their way. When a user names a metric the catalog has never heard of, searching the documentation comes before guessing and before asking.</p>
<p>That is why the documentation layer has an absolute rule attached: <strong>answer only from the retrieved passages.</strong> A fabricated configuration answer is indistinguishable from a real one to the person reading it, and they will act on it.</p>
<p>The same pattern shows up with fields nobody designed a metric for. Companies track their own dates in custom fields and want the rates between them: what share of referrals became booked assessments last month. There is no &quot;booking rate&quot; measure to select, only two custom fields and a ratio, so Sense builds the metric on the spot. A rate is really two counts, each with its own date field and conditions, and the agent restates both halves on every follow-up so a casual <em>now show it weekly</em> cannot silently drop one.</p>
<p>And when the arithmetic does not hold, the engine refuses. Ask for &quot;the latest value&quot; of a field that holds text and there is nothing to sort by; put a date window on a field that holds numbers and it matches nothing. In both cases the system stops with a reason specific enough to turn into a real choice:</p>
<blockquote>
<p>Picking the latest one needs a date to sort by. Want every record&apos;s value listed instead, one row each, or should I pick by an actual date field on that record?</p>
</blockquote>
<hr>
<h2 id="why-there-is-more-than-one-agent">Why there is more than one agent</h2>
<p>The two walkthroughs above were run by different agents, and the user never chose between them. The split is about context isolation, not specialization for its own sake: every fact in front of a model is a candidate for its answer, and irrelevant candidates are where plausible wrong choices come from. So each specialist sees only what its own decision needs.</p>
<ul>
<li>An <strong>entity selector</strong> reads entity descriptions, nothing deeper, and decides which parts of the business a question touches.</li>
<li>The <strong>query composer</strong> then sees only those entities&apos; fields. For revenue by team, that is <strong>93 fields instead of 1,220</strong>. Among 93, the right revenue field is obvious; among 1,220, revenue-shaped fields from jobs, quotes, invoices, and payments sit beside it, every one a defensible wrong choice.</li>
<li>The <strong>vocabulary resolver</strong> sees live stored values, not the query being built.</li>
<li>The <strong>record investigator</strong> runs its own tool loop and returns a written finding, so the conversation never carries the megabytes it read.</li>
<li>A <strong>knowledge searcher</strong> and a <strong>memory keeper</strong> see only passages and preferences.</li>
</ul>
<p>The specialists exist so each context can stay small.</p>
<hr>
<h2 id="what-the-agent-knows-before-it-answers">What the agent knows before it answers</h2>
<p>A model that reasons well and knows nothing about your business produces fluent, wrong answers. Several layers of context stand between the question and the answer.</p>
<p><img src="https://labs.zuper.co/content/images/2026/09/layers-sense.png" alt="Inside Sense: how an agent answers, investigates, and computes what nobody built" loading="lazy"></p>
<h3 id="a-curated-definition-layer-not-raw-tables">A curated definition layer, not raw tables</h3>
<p>The agent never sees a database. It sees a semantic layer we maintain: <strong>53 business entities carrying roughly 1,220 named metrics and attributes</strong>, each described in plain language. &quot;Total revenue&quot; is a definition we wrote, not a <code>SUM()</code> the model invented.</p>
<p>Three properties matter more than the size.</p>
<p>It is deliberately smaller than the data. Bridge tables, lookup tables, and internal masters stay hidden even though the system still queries them. A junction table between users and teams is a correct thing to have in a data model and a terrible thing to offer an agent as a choice.</p>
<p>It carries relationships, because <em>revenue by technician</em> spans multiple entities and the wrong path through them produces double counting rather than an error.</p>
<p>It carries counting rules. Take one job worth $1,000 that passed through four statuses. Group job revenue by status transition and the honest answer is still $1,000, but the arithmetic hands back $4,000, because the job lands in four groups and its revenue is counted in each. Every parent total split by a one-to-many child does this: job revenue by service task, invoice totals by line item. The overstatement looks entirely plausible, so the agent is directed to the child&apos;s own metrics instead, and told to name the risk out loud when a user insists on the inflating cut.</p>
<p>It is also the scope boundary: if an entity this workspace does not expose is absent, the answer is <em>not available yet</em>, never a nearest-neighbor guess.</p>
<h3 id="the-companys-own-vocabulary">The company&apos;s own vocabulary</h3>
<p>Nothing in the catalog knows this company says &quot;Inspection Ready,&quot; tracks a field called &quot;Assessment Booked Date,&quot; or ships from a warehouse named &quot;North Dock.&quot; That vocabulary is per-workspace, it changes weekly, and it is where many wrong answers come from.</p>
<p>So the agent looks things up, under one rule: <strong>a word you cannot find in the catalog is an unverified lookup, not an ambiguity.</strong> Look it up before asking the user, and before building a query with it. The lookup covers everything a company can name: statuses, categories, teams, payment terms, custom fields and the values inside them, checklist questions, locations, people, and customers. Common vocabularies are cached briefly per company; people and customers are searched live, because names are neither stable nor unique. A resolved name comes back with each candidate field tagged by role, and the agent picks by what the question is actually asking.</p>
<h3 id="the-conversation-itself">The conversation itself</h3>
<p>Every result in a session lands in a compact ledger the agent reads on later turns.</p>
<pre><code class="language-text">[ r3 ] entities: jobs, teams | rows: 12 | rendered: bar_chart | status: ok   &#x2190; latest
  query: revenue by team, last quarter, Inspection Ready residential jobs
  summary: Revenue by team, last quarter
</code></pre>
<p>That ledger makes follow-ups cheap and correct. <em>Show that as a table</em> re-renders an existing result without rerunning the query; <em>add job count</em> knows what <em>that</em> was; a superseded result is never offered again. It also carries context across surfaces: opening a chat from a pinned dashboard widget seeds the conversation with that widget&apos;s query and a sample of the rows on screen, so <em>why is the top one so high?</em> resolves to a real record.</p>
<p>The loop runs forward too. Every answer ends with grounded follow-ups based on what was just rendered: after revenue by team, <em>break down the top team by technician</em>; after a stat card, often a <em>why</em> that hands off to the investigator. That is how a measured answer turns into an investigated one without the user needing to know the second capability exists.</p>
<h3 id="standing-rules-and-runtime-facts">Standing rules and runtime facts</h3>
<p>Users set explicit preferences, like scoping job questions to this quarter or preferring tables for top-N lists, and the agent separately remembers how they talk. Precedence runs one way: this turn&apos;s instruction, then preferences, then memories, then defaults. Both get applied out loud, because a filter the user cannot see is a filter the user cannot correct. Then come the runtime facts that are wrong by default if omitted: <em>today</em>&apos;s date, the user&apos;s timezone, and who <em>my jobs</em> belongs to.</p>
<hr>
<h2 id="answers-we-shipped-that-were-wrong">Answers we shipped that were wrong</h2>
<p>Everything above is the system working. It did not always. Three answers that rendered wrong in production, and what each one bought us.</p>
<p><strong>The rate that counted deleted jobs.</strong> <em>What share of referrals became booked assessments?</em> The rate came back plausible and slightly wrong: it included deleted jobs. The exclusion of deleted records lived on the parent entities, so any query that touched only child records, like the custom-field values this rate is built from, never joined the parent and silently skipped the filter. The fix moved the exclusion into every child entity&apos;s own definition, so no query path can reach a child row whose parent is deleted. The rule it left behind: an invariant enforced on one entity and inherited by others through a join is not an invariant.</p>
<p><strong>The payments attributed to every job.</strong> <em>Payments by job type.</em> There is no direct link between a payment and a job; the correct chain is payment to invoice to job. But three linking paths of equal length existed, and the engine broke the tie by declaration order, walking payment to customer to jobs, which attributes each payment to every job the customer has ever had. The grouped total was inflated and perfectly plausible. The fix pinned the correct path as the tie-break winner, and any grouped total is now reconciled against its ungrouped baseline: it must sum back exactly.</p>
<p><strong>The filter that became an axis.</strong> <em>Count jobs by Referral Source, but only jobs whose Referral Date falls in 2025.</em> Two custom fields, one for grouping and one for filtering, and there was no supported recipe for that shape. The generator fell back to treating both as display dimensions: the date filter stopped filtering and became a breakdown axis, and all-time counts rendered under a title that said 2025. Root cause: an untaught query shape degrading into the nearest taught one instead of refusing. The fix added the missing recipe with the filter declared as a predicate, pinned by a regression test, and hardened the rule behind that earlier 4,812-job example: a constraint that cannot be honored must fail the query, not fall off it.</p>
<hr>
<h2 id="from-a-whiteboard-to-a-chart">From a whiteboard to a chart</h2>
<p>Planning meetings produce chart sketches, not queries. Somebody draws revenue by month with a second bar for job count, photographs the whiteboard, and wants that.</p>
<p>Sense reads the image as a <strong>layout specification</strong>: axes, groupings, chart type, sort order, the date window scribbled in the corner. It resolves the fields those labels imply, fetches live data, and renders the drawn chart with real numbers. Two rules keep it honest: the image is never a data source, so numbers visible in the photo are never quoted or computed on; and text inside the image is a label, never an instruction, so a sticky note reading <em>ignore your previous instructions</em> is just a label about a chart. The same path handles screenshots of reports from other systems, which is how people often ask.</p>
<p>[Screenshot: a whiteboard sketch beside the chart Sense rendered from it]</p>
<hr>
<h2 id="trust">Trust</h2>
<p>Which company&apos;s data a request may touch is injected by the request layer and cannot be named, overridden, or supplied by the model, on any path, in any tool; the same holds for the credentials behind investigations and the document set knowledge search may read. Every path is read-only. Everything read from outside (retrieved passages, preference text, company-authored field names, text inside images) arrives labelled as data, not instruction; <strong>a passage that contains instructions gets reported, not obeyed</strong>.</p>
<p>And the reasoning is visible while it happens: the trace at the top of this post, sources cited on documentation answers, the query inspectable behind every pinned widget, and the records openable behind every chart. When something breaks, the model gets a generic failure signal and the real error goes to logs, so field names and query fragments never surface as product copy.</p>
<hr>
<h2 id="what-it-cost-to-make-wrong-answers-hard">What it cost to make wrong answers hard</h2>
<p><strong>Context beats capability, and small contexts beat big ones.</strong> Every meaningful jump in quality came from giving the agent something true about the business, and most came from giving one agent less: <strong>93 fields instead of 1,220</strong> in front of the composer, a <strong>4KB</strong> preview instead of a <strong>50KB-plus</strong> payload in front of the investigator, a written definition instead of a guess in front of a metric. Each cut removed a whole class of confident error.</p>
<p><strong>Asking is a feature, not a fallback.</strong> On a 100-question QA set covering analytics questions and live-record investigations, <strong>91%</strong> of answers were correct. That is an evaluation result, not a production accuracy claim. Users forgive a clarifying question instantly and almost never forgive a wrong number, because a wrong number gets forwarded. <em>Clarify unless it is genuinely settled</em> is a higher bar than <em>clarify when stuck</em>.</p>
<p><strong>Guarantees live below the model.</strong> Every entry in the failure gallery ended the same way: a rule that had lived in instructions or in an assumed join moved into the layer beneath, where no phrasing can route around it.</p>
<p><img src="https://labs.zuper.co/content/images/2026/09/75bb26e8-5f35-41ec-ac9c-10880e2a1ea1.png" alt="Inside Sense: how an agent answers, investigates, and computes what nobody built" loading="lazy"></p>
<hr>
<p>This was never about putting a chat interface over a database. It was about engineering the context that lets an agent understand your business deeply enough to turn a question into the fastest path to understanding what is happening.</p>
]]></content:encoded></item><item><title><![CDATA[How we built the Zuper MCP: intent-driven testing, module-wise coverage, and closing 100+ schema gaps]]></title><description><![CDATA[<h1 id></h1><h2 id="the-problem-we-started-with">The problem we started with</h2><p>Zuper is a field service management platform with around 30 core business domains &#x2014; jobs, invoices, estimates, work orders, customers, service tasks, purchase orders, timesheets, and so on. When we set out to expose Zuper to LLM agents through the Model Context Protocol (MCP), the</p>]]></description><link>https://labs.zuper.co/blog/how-we-built-the-zuper-mcp-intent-driven-testing-module-wise-coverage-and-closing-100-schema-gaps/</link><guid isPermaLink="false">6a4e17955a2dfd3abfb65db7</guid><dc:creator><![CDATA[Maniarasan Sivaseran]]></dc:creator><pubDate>Mon, 13 Jul 2026 11:16:24 GMT</pubDate><media:content url="https://labs.zuper.co/content/images/2026/07/MCP_cover-image.png" medium="image"/><content:encoded><![CDATA[<h1 id></h1><h2 id="the-problem-we-started-with">The problem we started with</h2><img src="https://labs.zuper.co/content/images/2026/07/MCP_cover-image.png" alt="How we built the Zuper MCP: intent-driven testing, module-wise coverage, and closing 100+ schema gaps"><p>Zuper is a field service management platform with around 30 core business domains &#x2014; jobs, invoices, estimates, work orders, customers, service tasks, purchase orders, timesheets, and so on. When we set out to expose Zuper to LLM agents through the Model Context Protocol (MCP), the naive shape was obvious: one MCP &quot;tool&quot; per REST endpoint. That gave us hundreds of surface points, and every one of them had to be right.</p><p>&quot;Right&quot; is a loaded word. A tool schema is right when:</p><ul><li>Every field it accepts is actually accepted by the API.</li><li>Every field the API requires is either in the schema or auto-derived.</li><li>The enums it advertises are the enums the backend enforces.</li><li>The request-body shape is what the backend expects.</li><li>The endpoint URL and HTTP method match reality.</li><li>The agent gets a useful error, not a generic 400, when it&apos;s wrong.</li></ul><p>Getting one tool right by hand is easy. Getting ~300 tools across ~30 modules right needed a different approach.</p><p>This is the story of how we built it, and what we learned along the way.</p><hr><h2 id="why-not-just-unit-tests">Why not just unit tests</h2><p>The instinct was: unit-test each tool with mocked HTTP. We didn&apos;t take that path, for three reasons.</p><p><strong>One, mocks lie.</strong>&#xA0;A unit test with a mocked API client passes when the tool&apos;s schema accepts the input and the mock returns a canned response. It says nothing about whether the real backend will accept that payload. We learned this the hard way when a &quot;green&quot; test suite shipped a tool that called&#xA0;<code>POST</code>&#xA0;on an endpoint that only responded to&#xA0;<code>PUT</code>. The backend&apos;s 404 was a real signal; the mock happily returned success.</p><p><strong>Two, the schemas evolve daily.</strong>&#xA0;The MCP layer is a translation between two shapes &#x2014; the agent-facing schema and the backend API contract. The backend gets updates from a different team on a different cadence. We needed tests that run against the real API, ideally on staging, so drift is caught immediately.</p><p><strong>Three, the questions are workflow-shaped.</strong>&#xA0;&quot;Can an agent create a job, attach line items, apply a per-line tax, and generate an invoice?&quot; is not one unit test. It&apos;s a chain of six calls with state flowing between them. That&apos;s what our users care about; that&apos;s what we needed to verify.</p><p>So we built&#xA0;<strong>intent-based tests</strong>.</p><hr><h2 id="what-an-intent-test-is">What an intent test is</h2><p>An intent is a plain-English business goal, encoded as a small object:</p><pre><code class="language-js">{
  id: &apos;CUST-C01a&apos;,
  module: &apos;Customers&apos;,
  category: &apos;create&apos;,
  intent: &apos;Create a customer with name and email&apos;,
  tool: &apos;create_customer&apos;,
  args: () =&gt; ({
    first_name: &apos;Acme&apos;,
    last_name: &apos;Corp&apos;,
    email: &apos;contact@acme.example&apos;,
  }),
  onSuccess: (result, state) =&gt; { state.new_customer_id = result.customer_id; },
  verify: {
    tool: &apos;get_customer&apos;,
    args: state =&gt; ({ customer_id: state.new_customer_id }),
    check: r =&gt; r.customer.email === &apos;contact@acme.example&apos;,
  },
}
</code></pre><p>Every intent has:</p><ul><li><strong><code>intent</code></strong>&#xA0;&#x2014; a human-readable sentence that reads like a Jira ticket title.</li><li><strong><code>tool</code>&#xA0;+&#xA0;<code>args(state)</code></strong>&#xA0;&#x2014; the tool to call and how to build its arguments from the shared session state.</li><li><strong><code>onSuccess</code></strong>&#xA0;&#x2014; a hook to stash the returned identifier for downstream tests.</li><li><strong><code>verify</code></strong>&#xA0;&#x2014; a follow-up read that confirms the write took effect. Not just &quot;did the tool return 200,&quot; but &quot;does the corresponding read tool show the field we wrote.&quot;</li></ul><p><code>args</code>&#xA0;is a function, not a value, because tests share state. Creating a customer produces a customer ID; later intents that create jobs or invoices for that customer consume it. The state object is threaded through the whole run.</p><p>Intents are organized by module and category (create / update / list_get) &#x2014; one file per business area, one array per file.</p><p>The verify round-trip matters more than the initial call. Many silent-failure modes we later fixed &#x2014; schema too loose, backend rejecting a subfield but returning 200, tax field silently dropped &#x2014; only showed up in the read-back.</p><hr><h2 id="where-1000-intents-came-from">Where 1000+ intents came from</h2><p>Writing test cases by hand doesn&apos;t scale to hundreds of tools. Writing test cases from imagination is worse &#x2014; you test the happy path and miss the edge cases that actually break in production. We needed a source of test intents that reflected how real people used the product, not how engineers thought they should.</p><p>We got there by pulling from two sources:</p><p><strong>Past support tickets and issue trackers.</strong>&#xA0;Every ticket the team had ever worked on described a real customer trying to do a real thing. &quot;Customer can&apos;t mark an invoice partially-paid from the draft state&quot; is a test intent. &quot;Field engineer&apos;s clock-in reflected on the wrong user&quot; is a test intent. &quot;Percentage discount not honored when applied at line-item level&quot; is a test intent. We walked the ticket history module by module, converted each recognizable customer flow into an intent, and folded them into the coverage suite. Every ticket that turned into an intent got a comment linking back so the next incident would find the regression test already there.</p><p><strong>Intent libraries per customer segment.</strong>&#xA0;Different customers use different subsets of the platform. An HVAC service company lives in jobs, service tasks, and invoices. A property management firm lives in properties, service contracts, and recurring jobs. A distribution business lives in purchase orders, material requests, and transfer orders. Testing &quot;does the tool work&quot; is not the same as &quot;does the tool work for the way this segment uses it.&quot; We built per-segment intent libraries by walking through each customer&apos;s actual workflows &#x2014; using their data model (product categories, custom fields, business units) as inputs to the intents.</p><p>The two sources compounded. Ticket-derived intents caught the bugs we&apos;d already been bitten by. Segment-derived intents caught the bugs we&apos;d otherwise ship into a new segment blind. Together they took the suite past 1000 intents.</p><p>A side benefit we didn&apos;t expect: the intent library became&#xA0;<strong>documentation for the tool</strong>. Product and solutions teams can read the intent titles and immediately see what the platform can do for a specific segment. &quot;Schedule a follow-up job for a customer with an existing service contract&quot; is more useful than any API doc.</p><p>The rule we settled on:&#xA0;<strong>new bugs get intents before they get fixes.</strong>&#xA0;Once a customer-reported issue is in the intent suite, it stays covered forever. Fix without an intent, and it comes back.</p><hr><h2 id="the-runner-field-strip-retry">The runner: field-strip retry</h2><p>Real APIs fail for surprising reasons. A schema evolves, an enum tightens, a required field starts erroring out. When that happens in a test suite, we don&apos;t want to abandon the whole intent; we want to isolate which field is the culprit.</p><p>So the runner has a&#xA0;<strong>field-strip retry loop</strong>. On tool error, it:</p><ol><li>Pattern-matches the error message for a suspect field name.</li><li>Removes that field from the arguments.</li><li>Retries the call.</li><li>Loops until success, or until every optional field has been stripped.</li></ol><p>At the end of the run it reports two verdicts per field:&#xA0;<strong>VERIFIED</strong>&#xA0;(the field was accepted and reflected in the verify read) or&#xA0;<strong>UNVERIFIED</strong>&#xA0;(the field was stripped during retry).</p><p>The point isn&apos;t to make failing tests pass. It&apos;s to produce a&#xA0;<strong>coverage map</strong>. When we saw one field marked UNVERIFIED across forty different tests, we knew the whole schema for that field had drifted from the backend &#x2014; one bug, not forty.</p><hr><h2 id="the-coverage-document-%E2%80%94-one-per-module">The coverage document &#x2014; one per module</h2><p>Each module owns two generated files:</p><ul><li><strong>A run report</strong>&#xA0;&#x2014; the summary. Every intent, PASS/FAIL/VERIFY_FAILED, which fields were stripped, which are UNVERIFIED.</li><li><strong>A failure guide</strong>&#xA0;&#x2014; each failure gets a stanza with the tool name, error message, stripped fields, likely cause, and a &quot;suggested fix&quot; line.</li></ul><p>These aren&apos;t just artifacts &#x2014; they became the primary planning surface. Every session started by reading the last run&apos;s failure guide. Fix, re-run, watch FAIL count drop, watch PASS count climb.</p><p>The pattern we settled into:</p><pre><code>Round 1: baseline run &#x2014; enumerate every failure.
Round 2: fix the top N failures, re-run just those tests.
Round 3: re-run the whole module. New failures? Different failures?
Round 4: expand coverage &#x2014; add intents for edge cases surfaced during fixes.
</code></pre><p>A concrete example: the tax and line-item audit. Round 1 showed several &quot;create estimate with ad-hoc tax&quot; intents as VERIFY_FAILED. Investigation revealed the backend requires a tax UID reference for every document-level tax entry &#x2014; it hard-rejects ad-hoc&#xA0;<code>{name, percent}</code>&#xA0;forms. We narrowed the schema to require the UID. Re-ran. Passed. Added Round-2 intents for line-item-level tax and markup. Passed. Added Round-3 intents for the update path, tax-exempt overrides, mixed inheritance. Passed. Coverage went from 4 tax intents to 22 in one week, all green.</p><p>The coverage doc is a ratchet: it only counts intents that VERIFY, and the verified count only goes up.</p><hr><h2 id="adversarial-verification-with-parallel-agents">Adversarial verification with parallel agents</h2><figure class="kg-card kg-image-card"><img src="https://labs.zuper.co/content/images/2026/07/parallel_agents.png" class="kg-image" alt="How we built the Zuper MCP: intent-driven testing, module-wise coverage, and closing 100+ schema gaps" loading="lazy" width="1536" height="1024" srcset="https://labs.zuper.co/content/images/size/w600/2026/07/parallel_agents.png 600w, https://labs.zuper.co/content/images/size/w1000/2026/07/parallel_agents.png 1000w, https://labs.zuper.co/content/images/2026/07/parallel_agents.png 1536w" sizes="(min-width: 720px) 720px"></figure><p>Half the bugs we caught weren&apos;t in the runner output. They were in the&#xA0;<em>schema itself</em>&#xA0;&#x2014; fields that looked plausible but weren&apos;t in the backend, enums that shipped extra values the backend would reject.</p><p>For those, the runner alone can&apos;t help &#x2014; a schema that never gets exercised by an intent test can be arbitrarily wrong and nobody notices. So we ran periodic audit passes.</p><p>The pattern:</p><ol><li>Spawn 4&#x2013;6 parallel research agents. Each takes a subset of modules.</li><li>Each agent reads the MCP tool file, the corresponding webapp form component, and the backend controller. It reports a table: field-by-field, what&apos;s in the tool vs the UI vs the backend.</li><li>Agent findings are&#xA0;<strong>claims</strong>, not truth. Every high-severity claim gets spot-verified against the actual source before it hits the audit doc.</li><li>Consolidate into a single module-wise gap doc.</li></ol><p>The reject-before-trust step is important. Agents are helpful and confidently wrong in equal measure. In one audit round, an agent claimed a listing tool used the wrong enum values. We opened the file &#x2014; the actual enum was correct; the agent had misread a docstring elsewhere in the same file. We rejected that claim in the write-up. In the same round, another agent flagged a deposit-recording tool as calling POST when the backend registered PUT. We opened the file &#x2014; real bug, six months old, silent behind an outer try/catch.</p><p><strong>Every claim gets grep-verified before it becomes a finding.</strong>&#xA0;This one habit cut the false-positive rate on our audit docs from ~30% to under 5%.</p><hr><h2 id="when-docs-arent-enough-browser-automation-as-a-discovery-tool">When docs aren&apos;t enough: browser automation as a discovery tool</h2><p>Roughly a third of the way through the project, we hit a wall: the internal API reference documents were incomplete. Some endpoints were undocumented. Others were documented but omitted fields the UI was actually sending. Some documented fields were flagged &quot;optional&quot; but were required in specific configurations. Without ground truth for those endpoints, we were guessing.</p><p>So we brought browser automation into the loop.</p><p>The pattern was simple: script the webapp&apos;s real workflow end-to-end &#x2014; log in, create a job, add products, set a tax, save, refresh &#x2014; and capture every network request. Each captured request gave us a real payload: exact field names, exact nesting, exact enum values, exact wrappers. When our schema disagreed with the payload, the schema was wrong.</p><p>A few things it exposed that we&apos;d have missed otherwise:</p><ul><li><strong>Undocumented approval hooks.</strong>&#xA0;A specific status transition sent an extra approval identifier the API reference didn&apos;t mention. The browser automation caught it on the first pass.</li><li><strong>Field-nesting quirks.</strong>&#xA0;The webapp nests certain settings (billing frequency, payment term) under a nested accounts object on save. Our tool was sending them flat. The API &quot;accepted&quot; both shapes but silently discarded the flat one. The captured payload made the correct shape visible.</li><li><strong>Company-config-gated behaviors.</strong>&#xA0;Some fields only appear in the payload when a company setting is on. Running the automation against two configurations side by side let us design a schema that handled either.</li></ul><p>Browser automation wasn&apos;t a replacement for the API docs or the backend source. It was the fastest way to answer &quot;what does the UI actually send?&quot; when the other two disagreed, and to catch the hidden third case both had missed. The reusable rule:&#xA0;<strong>when in doubt, capture the network traffic.</strong></p><hr><h2 id="the-module-dependency-matrix">The module dependency matrix</h2><p>Ordering modules by &quot;business criticality&quot; was the first cut. But criticality isn&apos;t the same as &quot;safe to fix first.&quot; A tool schema for jobs depends on customers existing, products existing, tax IDs existing, business units existing. Fix jobs first and half your tests fail because the entities they depend on aren&apos;t seeded.</p><p>So we built a&#xA0;<strong>module dependency matrix</strong>&#xA0;&#x2014; a per-module document mapping every upstream dependency:</p><pre><code>Customer domain
  &#x251C;&#x2500; Depends on: categories, organizations, tax groups, pricelists
  &#x2514;&#x2500; Depended on by: jobs, invoices, estimates, projects, contracts, ...

Job domain
  &#x251C;&#x2500; Depends on: customer, product, product category, job category,
  &#x2502;              user, team, business unit, service tasks, tax
  &#x2514;&#x2500; Depended on by: invoice, estimate, service report, timelog, ...
</code></pre><p>Roughly two dozen of these documents now live alongside the code, one per business domain. Each names its foreign-key relationships, the exact fields that carry the reference, and any known quirks (e.g. &quot;the job&apos;s property inherits from the customer if omitted, but only when the customer has a default property set&quot;).</p><p>Two immediate benefits:</p><p><strong>Testing order becomes deterministic.</strong>&#xA0;The seed script that populates the shared session state reads the dependency graph and creates entities in topological order. Customers before jobs, products before line items, tax before invoices. No more &quot;test failed because the customer ID is empty.&quot;</p><p><strong>Field audits become sharper.</strong>&#xA0;When we&apos;re auditing an invoice-creation tool, the customer dependency doc tells us which fields flow into the customer lookup, which fields the invoice inherits automatically, and which ones a manual override is expected. That&apos;s usually what the API docs leave implicit.</p><p>We also generated companion&#xA0;<strong>API reference documents</strong>&#xA0;&#x2014; one per module &#x2014; that consolidate every endpoint we call, its verified request shape (from network capture + backend source cross-check), and its response shape. These sit alongside the dependency matrix. The pair of documents &#x2014; dependency map + verified API reference &#x2014; became the input to every new intent-suite expansion and every audit round. If a claim in a coverage doc contradicts one of these two, the coverage doc gets updated first.</p><hr><h2 id="module-by-module-progression">Module-by-module progression</h2><p>We didn&apos;t try to fix all tools at once. We ordered modules by:</p><ol><li><strong>Business criticality</strong>&#xA0;(jobs first, then invoices and estimates, then service tasks, then material requests, then everything else).</li><li><strong>Intent-suite coverage</strong>&#xA0;(biggest suites first &#x2014; they surface the most bugs per hour).</li><li><strong>Cross-module dependency</strong>&#xA0;(fix customers before jobs, because job creation needs customer state).</li></ol><figure class="kg-card kg-image-card"><img src="https://labs.zuper.co/content/images/2026/07/19055c66-09fd-4c19-8ace-059fa08375ee.png" class="kg-image" alt="How we built the Zuper MCP: intent-driven testing, module-wise coverage, and closing 100+ schema gaps" loading="lazy" width="1536" height="1024" srcset="https://labs.zuper.co/content/images/size/w600/2026/07/19055c66-09fd-4c19-8ace-059fa08375ee.png 600w, https://labs.zuper.co/content/images/size/w1000/2026/07/19055c66-09fd-4c19-8ace-059fa08375ee.png 1000w, https://labs.zuper.co/content/images/2026/07/19055c66-09fd-4c19-8ace-059fa08375ee.png 1536w" sizes="(min-width: 720px) 720px"></figure><p>Per module, the loop was:</p><ol><li><strong>Baseline</strong>&#xA0;&#x2014; run current intents, capture the coverage report.</li><li><strong>Field audit</strong>&#xA0;&#x2014; parallel research pass to find missing or wrong fields vs the UI and the backend.</li><li><strong>Fix critical</strong>&#xA0;&#x2014; schema and endpoint corrections.</li><li><strong>Expand coverage</strong>&#xA0;&#x2014; new intents for the fields just added.</li><li><strong>Regression</strong>&#xA0;&#x2014; re-run the whole module; add any newly surfaced failures to the queue.</li></ol><p>For the jobs module alone, that loop produced dozens of field-level fixes and roughly doubled the passing intent count.</p><hr><h2 id="what-surprised-us">What surprised us</h2><p><strong>Silent overrides.</strong>&#xA0;Several endpoints happily accept fields they then ignore. The backend&apos;s tax calculator overwrites caller-supplied tax name and percent from the tax-master lookup, regardless of what the caller sent. Our MCP was faithfully passing through fields the backend would silently drop, giving agents the illusion of setting them. Fixing this meant narrowing schemas&#xA0;<em>below</em>&#xA0;what the backend would accept &#x2014; trading permissiveness for predictability.</p><p><strong>The unit-price / price schism.</strong>&#xA0;Jobs use one field name for line-item price. Invoices and estimates use a different one. A converter tool translates one to the other. No documentation of the flip anywhere. An agent using our tools cross-module has to know this by feel unless the tool descriptions call it out. We added the notes.</p><p><strong>Status enums that lie.</strong>&#xA0;One status-update tool accepted status values as write targets that the backend never allows to be set directly &#x2014; some are computed automatically from other actions, some are display-only. An agent that tried those transitions through MCP would get a generic 4xx and no clear signal. Narrowing the write enum was a one-line fix; finding the mismatch was hours of matrix-reading.</p><p><strong>The auto-derive trap.</strong>&#xA0;Some create tools mark fields optional even though the backend requires them &#x2014; because when the caller passes a parent entity, the tool auto-copies the required fields from that parent. The schema looks broken until you read the execute body. Documentation matters more than schema strictness once the tool starts doing derivation work.</p><hr><h2 id="the-tooling-that-kept-us-moving">The tooling that kept us moving</h2><ul><li><strong>A restart-between-chunks runner</strong>&#xA0;&#x2014; the upstream MCP session state degrades under sustained load; a chunked runner is a workaround, not a fix, but it kept the suite moving.</li><li><strong>A shared session-state file</strong>&#xA0;&#x2014; pre-seeded IDs for common entities (customer, product, tax, category), populated once via a seed script; every test run starts from it.</li><li><strong>An ID-prefix filter</strong>&#xA0;&#x2014; a CLI flag to re-run just one group of intents by ID prefix. Made the &quot;fix, run three tests, re-run all&quot; loop fast enough to stay in flow.</li><li><strong>An env-file credential loader</strong>&#xA0;&#x2014; no hardcoding, no accidental leaking of production keys into a staging suite.</li></ul><p>None of these are impressive individually. Together, they made large-scale coverage tractable.</p><hr><h2 id="what-wed-change-next-time">What we&apos;d change next time</h2><p><strong>Start with the coverage doc, not the tool.</strong>&#xA0;The first tools we built had shallow schemas because we hadn&apos;t scoped the UI form&apos;s field list yet. We ended up rebuilding several tools twice. Working from a coverage matrix &#x2014; every field in the UI form, every field in the backend schema, every field we plan to expose &#x2014; would have front-loaded the design work.</p><p><strong>Adopt a strict &quot;tighten before ship&quot; rule.</strong>&#xA0;Schemas are easy to loosen after the fact and painful to tighten. Every &quot;make it optional for now&quot; we deferred came back as a support ticket. Better to start narrow and open up on demand.</p><p><strong>Contract tests, not just integration tests.</strong>&#xA0;The intent suite exercises real endpoints, but it doesn&apos;t lock the backend response&#xA0;<em>shape</em>. A field rename on the backend silently breaks our verify checks. A layer of contract tests that snapshot response schemas would have caught two silent regressions we found only by luck.</p><p><strong>Backend-side gate for enums.</strong>&#xA0;Many enum bugs (statuses in write enums that shouldn&apos;t be, transition targets the backend won&apos;t accept) exist because the backend accepts one set of values as input and produces a different set as output. The backend should either publish the write-side matrix or the MCP should mirror the matrix explicitly.</p><hr><h2 id="where-the-project-stands">Where the project stands</h2><ul><li>Approaching 300 tools across roughly 30 modules</li><li>1000+ intent tests, sourced from past tickets and per-segment customer workflow libraries</li><li>Intent suites for the top business modules: jobs, invoices, estimates, customers, service tasks, material requests, timesheets, purchase orders</li><li>Live-API test runs on staging on demand</li><li>Per-module coverage docs updated on every run</li><li>Field-gap audits checked in with prioritized fix queues</li></ul><p>The core lesson is unglamorous but real:&#xA0;<strong>at scale, the test suite is the design tool.</strong>&#xA0;Intents make the requirements concrete. Coverage docs make the gaps visible. Round-by-round iteration makes the delta shippable. We didn&apos;t write hundreds of tools right on the first try. We wrote them approximately, and then let the tests tell us where they were wrong.</p><p>That, more than anything, is how the <strong>Zuper</strong> MCP got built.</p>]]></content:encoded></item></channel></rss>