Building a read-only web audit MCP server that never opens a browser
Clientinternal tooling
ServicesMCP
StackTypeScript 5.7 on Node.js 22, official @modelcontextprotocol/sdk, undici, cheerio, express 5, zod, pino
The problem
There is no shortage of tools that will score a web page. The problem is where they sit. Most of them live behind a dashboard you have to open in a browser, paste a URL into, wait for a spinner, and then read a report you cannot feed back into the work you were already doing. The audit happens in one tab, the fix happens in another, and nothing connects them.
The better version of this is to let the AI client you are already working in run the audit itself, as a tool call. Ask it to check a page, get the findings back as structured data in the same conversation, and act on them without leaving the editor. That is what an MCP server is for.
So the requirement was narrow. Build an MCP server that audits a public URL, returns findings an agent can actually use, and does it without becoming a security liability on the machine or the network it runs on.
What the server does
mcp-server-web-audit is a read-only MCP server that takes a URL and returns an audit report. Every tool performs exactly one GET of the page the user asked for, then analyses the static HTML and the response headers it got back.
It ships five focused tools and one combined tool:
Tool
What it covers
audit_seo
Title and meta description length, indexability, canonical, H1, Open Graph, Twitter card, hreflang, lang
GA4, GTM, Meta Pixel, TikTok, LinkedIn, Hotjar, Clarity, plus consent coverage
audit_accessibility
Image alt text, form labels, landmarks, heading order, lang
audit_performance
TTFB, HTML size, compression, caching, plus optional CrUX field data
audit_full
All five engines from one fetch, with a weighted overall score
Each tool takes a url and a format (markdown or json) and returns either. Every tool is annotated readOnlyHint: true, openWorldHint: true, idempotentHint: true, destructiveHint: false, so a client can reason about what it is about to call before it calls it.
The decision that shaped everything: no headless browser
The obvious way to audit a page is to launch a headless browser, render the page, and inspect the live DOM. That is what Lighthouse and most commercial scanners do. We did not.
A headless Chromium means a large dependency footprint, a few seconds of cold start, and a much larger attack surface if anything about the input is hostile. Running one per audit call, on the same box that also accepts requests, is a lot of weight to carry for checks that are mostly about what came back in the response.
So the performance engine reports synthetic metrics measured from the audit server's own fetch: TTFB, decompressed HTML size, whether the response was compressed, and what Cache-Control says. For real user experience numbers, it optionally pulls p75 phone data from the Chrome UX Report API with a CRUX_API_KEY, falling back from URL-level to origin-level when the URL has no CrUX data of its own.
That choice buys a small container, a fast cold start, and a much simpler security story. It costs a set of checks we cannot make:
Tags injected at runtime are invisible. If GTM injects a pixel after load, the tracking engine will not see it.
There is no rendered DOM, so no colour contrast check and no computed-style checks in the accessibility engine.
No layout means no real Core Web Vitals measurement, which is exactly why CrUX is an add-on rather than a substitute.
Those limits are stated in the tool documentation rather than hidden. An audit tool that quietly reports "no issues" because it could not see an issue is worse than useless.
Architecture
The layering is deliberate. tools/common.ts owns one pipeline that every tool runs through, so deadline handling, validation, caching, fetching and rendering behave identically no matter which engine you asked for. The engines themselves are pure functions over a fetch result or parsed HTML. That means they are unit-testable without a network, and it means a bug in the CSP parser cannot affect the SEO report.
Technical decisions worth explaining
DNS pinning, not a URL blacklist
The standard way to block server-side request forgery is to check whether a hostname looks private. That check has a hole: the hostname is resolved once for the check and again for the connection, and between those two moments DNS can answer differently. A host that resolves to a public address during validation and to 127.0.0.1 at connect time walks straight past a URL-level guard.
So validation in this server happens at connect time, not before it. The fetcher uses undici with a per-call Agent whose connect.lookup resolves the hostname, validates all returned addresses, and returns only the addresses that passed. The socket cannot be opened against anything that was not just checked.
On top of that:
Ports are restricted to a default set of 80, 443, 8080, 8443, configurable, or * if you really mean any port.
Redirects are followed manually, with every hop re-validated before it is fetched, capped at 5 by default.
3xx response bodies are cancelled rather than read.
Set-Cookie headers are collected across every hop in the chain, which is how the security engine catches a cookie set on a redirect that would otherwise be invisible.
A Content-Length pre-check runs alongside a streamed byte cap, and the reader is cancelled the moment the cap is passed.
Error messages are mapped to codes without leaking resolved IPs.
Errors that name the reason
A server that returns "audit failed" is not useful to an agent trying to decide what to do next. Every failure maps to a code an agent can branch on: SSRF_BLOCKED, POLICY_BLOCKED, DNS_FAILED, HTTP_ERROR, NOT_HTML, ROBOTS_DISALLOWED, TOO_MANY_REDIRECTS, RESPONSE_TOO_LARGE, TIMEOUT, CANCELLED, FETCH_FAILED.
Two of those encode real decisions. Error pages are not audited, so a 404 or a 500 returns HTTP_ERROR rather than a report on the error page. And the security audit works on any response type, because headers exist whether or not the body is HTML, so a non-HTML response returns NOT_HTML for the HTML-based tools while still allowing audit_security to run.
In-flight de-duplication with per-caller cancellation
Two reasons drive caching here: the same URL gets audited repeatedly during a fix cycle, and repeated audits of the same URL are wasted work against someone else's server.
The cache is an in-memory LRU with a TTL, default 5 minutes and 100 entries. The interesting part is what happens when two callers ask for the same URL at the same time. Rather than issuing two fetches, the second call waits on the first. But the shared fetch runs on the cache's own abort signal, not either caller's. A caller that gives up only stops waiting. The fetch is aborted only once every waiting caller has cancelled.
That distinction matters in an agent loop, where a client will frequently abandon a call it no longer needs. Without it, one impatient caller would kill a fetch that three others were still waiting on. Partial results are never cached, so a failed engine cannot poison the cache for the next caller.
A deadline that covers the whole call
There are two timeouts, not one. AUDIT_TIMEOUT_MS (default 15 seconds, capped at 30) applies per network operation. AUDIT_TOTAL_TIMEOUT_MS (default 45 seconds) covers the entire audit, and it is combined with the client's own abort signal.
The total deadline is the one that got the most attention, because it has to hold across every stage: DNS resolution during validation, robots.txt fetch if enabled, the page fetch, and each engine. A per-request timeout alone lets a pathological target drip-feed bytes slowly enough to keep an audit running far past any reasonable bound.
Scoring that survives a partial failure
audit_full runs all five engines from a single fetch, with the CrUX call in parallel. Each engine runs in its own try/catch. If one fails, its category is dropped from the result, listed under errors in JSON or under "Partial Result" in Markdown, and the overall score is re-weighted over the categories that did complete:
The weights are SEO 25 percent, Security 25 percent, Tracking 15 percent, Accessibility 15 percent, Performance 20 percent. Scores map to ratings at 80 and above for good, 50 and above for warning, below 50 for poor.
Re-weighting rather than zero-filling is the difference between "we could not measure performance, so here is what the rest says" and "this site scored 75 because our own fetch timed out". The first is a usable result. The second is a bug report about the auditor.
Writing the score as penalties, not points
Every engine computes a penalty total and converts it once: score = clamp(100 - penalty, 0, 100). Adding a new check means adding a penalty for the specific thing it finds, and the score for every other site stays stable. The alternative, adding points for each passing check, makes every score change whenever a check is added, which makes historical comparisons meaningless.
Cookie penalties are capped on purpose. A site with many cookies set across redirect hops should be flagged, not driven to zero on that alone.
Challenges
Crawler-scoped header directives
X-Robots-Tag is not a single value. It can carry a crawler name, as in otherbot: noindex. Reading that as a blanket noindex would report a blocking directive on a page that is perfectly indexable to Google.
The SEO engine parses each directive against its scope and ignores rules aimed at a crawler other than Googlebot. It is a small parsing decision that removes a whole class of false positives, and false positives are expensive in an audit tool because they teach the person reading the report to distrust it.
Decorative images are not missing alt text
The naive accessibility check flags every <img> without a non-empty alt. That buries the real findings in noise on any modern page, because decorative images are supposed to have alt="".
The engine treats three cases as intentionally decorative and excludes them: alt="", role="presentation" or role="none", and aria-hidden="true". What is left is images that genuinely lack a text alternative, which is a much shorter and more actionable list.
A tracking score that is not just a tag count
Counting trackers is easy and mostly meaningless. Four well-configured tags on a site with a consent manager is not the same problem as four tags with no consent handling at all.
So the tracking score drops for specific, defensible reasons: the same tracker loaded under multiple IDs, ad or heatmap trackers with no recognisable consent manager and no Google Consent Mode default, more than four distinct trackers, and tracker scripts loaded without async or defer where they block rendering. The consent list covers the managers that actually show up in the wild: Cookiebot, OneTrust, CookieYes, Usercentrics, Didomi, iubenda, Osano, Termly, Quantcast Choice, Complianz and Klaro.
Keeping the HTTP transport from being the weak point
The server defaults to stdio, which has no network surface at all. The HTTP transport exists for self-hosting, and it only turns on when TRANSPORT=http, at which point the configuration is validated and refuses to start if it is unsafe: a missing MCP_AUTH_TOKEN, a token shorter than 32 characters, or a non-loopback HOST with no MCP_ALLOWED_HOSTS list.
In front of that sit a Host header check, an Origin allowlist, an IP rate limit of 30 requests per minute by default, and a separate failure limit of 10 for authentication attempts, so a token-guessing loop gets shut down faster than ordinary traffic. The JSON body is capped at 256 KB.
Implementation walkthrough
An audit call runs through this sequence.
runAudit creates one deadline for the whole call and combines it with the client's abort signal.
validateTarget checks the scheme, rejects embedded credentials, applies the domain allow and block lists, and runs the SSRF check with DNS resolution, all still under the deadline.
The cache is consulted. A hit returns immediately with a notice prepended to the Markdown output. A miss either computes, or joins an existing in-flight computation for the same URL.
fetchTarget checks robots.txt if RESPECT_ROBOTS_TXT=true, then calls safeFetch. A status of 400 or above returns HTTP_ERROR without auditing the error page.
HTML-based engines require an HTML content type and return NOT_HTML otherwise. The security engine runs on any response.
renderResult produces Markdown or pretty-printed JSON. Errors come back as isError: true results carrying their code.
If the client supplied a progressToken, audit_full sends three stage notifications: fetching the page, running the audit engines, and rendering the report.
Each tool's input schema is a fresh zod shape rather than a shared referenced type, so the published JSON Schema contains no $ref. That is a small thing that makes the tool definitions easier for every client to consume, because not every MCP client resolves references the same way.
Testing
The test suite runs with no real network access at all. DNS is faked globally, so the SSRF and validation paths can be exercised deterministically, and HTTP-level tests run against a local server started inside the test process. Integration tests drive the server through an in-memory MCP client and over the HTTP transport.
Coverage is tracked with vitest, the build runs through tsconfig, and the repo carries a CI workflow, an eslint config, a Dockerfile, and a security note in docs/SECURITY.md.
Current state and results
The server is buildable and tested, and runs both as a local stdio server for an editor client and as a self-hosted HTTP endpoint behind Docker.
On a real audit of our own site (flowagenz.com, October 2026) the combined tool returned 88 out of 100, with SEO 90, security 80, accessibility 100, tracking 100 and performance 75. The findings it surfaced were specific enough to act on directly: the title tag ran 63 characters against a 60 character target, the CSP allows unsafe-inline and unsafe-eval, X-Content-Type-Options: nosniff is missing, and TTFB measured around 1003 ms.
That is the shape of output the tool is built to produce. Not a grade, a list of edits.
What is planned next
A rendered-DOM mode, opt-in. The static approach is right for the default path, but colour contrast and runtime-injected tags are real gaps. The plan is an optional headless path, off by default, so the fast and lightweight behaviour stays the default and the heavier mode is a deliberate choice.
Structured output separate from rendered output. Returning Markdown is convenient for a chat client and awkward for a program. The JSON format exists, and the next step is making the schema stable and documented so a scheduled job can consume it directly.
Historical diffing. An audit is a snapshot. The more useful question is what changed since the last run, which needs stored results and a comparison call rather than a new engine.
Takeaways for anyone building something similar
Resolve at connect time, not at check time. Every SSRF guard that validates a hostname and then hands that hostname to an HTTP client has a DNS rebinding window between the two. Pinning the validated addresses at the socket is the version that actually holds.
Not running a browser is a feature, as long as you document what it costs. The static approach is faster, smaller and safer, and it cannot see three categories of problem. Writing those limits into the tool description is what keeps the tool trustworthy.
And build the error codes before the features. An agent deciding what to do next reads the error, not the stack trace. NOT_HTML and HTTP_ERROR are the difference between an agent retrying a transient failure and an agent realising the URL was wrong.