The problem
Competitor content research is slow in a way that is hard to notice until you count the hours. You open a competitor's page, skim it, open your own, skim it too, and try to hold both in your head while you work out what they cover that you do not. Do that across ten competitors and the conclusion you reach is a feeling, not a finding.
The mechanical part of this is entirely automatable. Word counts, heading structures, keyword overlap, readability, which pages link out and how much, whether the page declares schema. None of that requires judgement, and all of it is better done consistently across ten pages than carefully across two.
The judgement part is what you want the AI to help with, which means the numbers have to reach it first. That is the gap this server fills.
There was a second, less obvious requirement. Fetching a competitor's page is making a request to someone else's server, sometimes through a headless browser, sometimes against a site that is actively hostile to being scraped. A tool that does that carelessly is not a research tool. It is a liability.
What the server does
mcp-server-competitor-content exposes eight tools across three roughly distinct jobs.
Multi-URL tools fetch up to three pages at a time and still space requests to the same host by RATE_LIMIT_DELAY_MS, which defaults to 1000 ms. Concurrency across hosts and politeness per host are two different settings, and conflating them is how a scraper ends up hammering one site while claiming to be throttled.
content_gap_analysis, compare_headings and cluster_competitors send notifications/progress, one per fetched page, when the client supplies a progressToken. On a ten page comparison that turns a silent wait into a visible one.
The SPA problem, and the fallback that solves it
The server reads static HTML first. That is cheap, fast and copies nothing beyond text. It also returns almost nothing useful on a single page application, where the served HTML is a shell and the content arrives later over fetch calls.
Rather than choosing between speed and coverage, the server does both in order. If a static fetch yields less than HEADLESS_MIN_CONTENT_CHARS, which defaults to 200 characters, and ENABLE_HEADLESS_FALLBACK is on, which it is by default, the page is rendered with Playwright instead. Playwright is an optional dependency, so the fallback needs a one-time npx playwright install chromium before it can work. The render waits for domcontentloaded and then up to 3 seconds for network idle, within an overall HEADLESS_TIMEOUT_MS of 15 seconds.
This is the right default for research work. Most pages are static and most of the time is spent on pages that are, so paying the browser cost only when the cheap path fails keeps the average fast without giving up on the modern web.
The hard part: pinning DNS inside a browser
Once a headless browser is in the picture, the security work stops being a simple URL check.
If your HTTP client validates a hostname before connecting, you can pin the validated address at the socket and close the DNS rebinding window. A browser does not offer you that control. It resolves hostnames itself, follows redirects itself, and makes requests you did not hand it: preconnect hints, WebRTC, and anything a script decides to fetch.
The approach here is to take the network away from Chromium entirely.
Chromium is launched with its own DNS disabled:
Anything that slips past request interception, whether a preconnect or a missed request, fails to resolve rather than reaching a host on its own.
Every request is intercepted and fetched by the pinned Node client. The route handler checks the URL against the same public-address policy the rest of the server uses, and either fulfils the request from the validated fetch or aborts it. Skipping images, media and fonts is a deliberate bandwidth decision: they are not needed to extract text, and not fetching them removes most of the request volume.
Redirects are followed in Node, not in the browser. When a fetch returns a 3xx, the handler resolves the location, re-validates it, and fetches the next hop itself, up to MAX_REDIRECTS = 5. That happens because Chromium, on receiving a fulfilled redirect, would try to resolve the target itself. Re-checking every hop is the only way the policy applies to the whole chain rather than just the first request.
The render is bounded. MAX_REQUESTS_PER_RENDER is 150. Service workers are blocked at the context level, so a page cannot install a background fetch path around the intercept. WebSockets are closed on open with code 1008 when the Playwright build supports intercepting them, and skipped when it does not.
The browser instance is reused, the context is not. One Chromium process is launched lazily and shared, while each render gets its own context that is closed in a finally block. Launch cost is paid once, and page state is still thrown away between renders.
The per-host verdicts are cached for the duration of a single render, so a page pulling twenty subresources from one host is resolved and validated once rather than twenty times. Over a render, that is the difference between a fast check and a slow one that happens to produce the same answer.
Technical decisions worth explaining
A robots.txt parser with no regular expressions
Most robots.txt implementations build a regular expression per rule and match against the path. That is convenient and it is a denial-of-service hole. A rule containing pathological alternation or nesting can be made to backtrack catastrophically, and the pattern comes from the site you are fetching, not from you.
So patterns are matched without regular expressions, in linear time. Disallow patterns longer than 512 characters are cut to a prefix, which makes the match stricter, Allow patterns that long are dropped, and files larger than 512 KiB are truncated rather than parsed in full. A hostile robots.txt costs you a bounded amount of work and nothing more.
Correctness follows RFC 9309 where it matters in practice: product-token matching, groups merged when shared between user agents, case-sensitive * and $ pattern handling, and status handling that is easy to get backwards. A 4xx on robots.txt means allow, because an absent file is permission. A 5xx means deny, because the site is telling you something is wrong. An unreachable host denies for 60 seconds and then retries, rather than caching the failure forever or retrying on every call. A redirect to a different host is checked against that host's robots.txt, not the original's.
robots.txt verdicts are cached separately from page content, with a 3600 second TTL against 300 seconds for content. A site's crawl policy does not change as often as its pages do.
Decoding pages the way browsers do
Text extraction sounds like a solved problem and is not, because character encoding on the web is a chain of fallbacks and every step exists because the previous one failed somewhere real.
The decoder tries four sources in order: the byte order mark, then the charset parameter on the Content-Type header, then a <meta charset> or <meta http-equiv> tag found in the first 2 KB, and finally UTF-8. The meta sniff only runs for markup content types, since a JSON or text response with <meta in it is not declaring an encoding.
There is a further fix for a Node quirk. Some Node releases decode windows-1252, and its latin1 and ascii aliases, as ISO-8859-1. Those differ in the 0x80 to 0x9F range, where ISO-8859-1 has control characters and windows-1252 has the curly quotes and dashes that pages actually use. When the decoder reports windows-1252, those bytes are remapped to their real windows-1252 characters. Without it, a competitor's smart quotes arrive as invisible control codes in the extracted text.
Unknown or malformed charset labels fall back rather than throwing. A TextDecoder that rejects a label is caught and the next source in the chain is tried, so a page declaring an encoding nobody has heard of still gets read as UTF-8 instead of failing the tool call.
Clustering that does not slow down as pages get longer
cluster_competitors builds a TF-IDF vector per page and clusters them by cosine similarity. The naive version scales with total corpus size in a way that gets slow as pages get long, because a 20,000 word page produces far more distinct terms than a 500 word one, and k-means compares every vector against every centroid on every iteration.
So each document's vector is reduced to its heaviest MAX_CLUSTER_VECTOR_TERMS = 2000 terms before clustering. The cost per document becomes bounded instead of proportional to page length, and the terms dropped are the ones that carry least signal anyway, since TF-IDF has already pushed them to near zero.
Initialisation uses farthest-first seeds rather than random ones. The first seed is the first document, and each subsequent seed is the document least similar to any existing seed, which spreads the starting centroids out and makes the result deterministic. Seeding stops early when the remaining documents duplicate an existing seed, which is the common case for a group of competitors writing about the same subject, and it means asking for more clusters than the corpus supports returns fewer rather than empty ones.
Empty clusters are dropped and the remaining clusters are renumbered, so a caller never receives a cluster with no members. Each cluster reports avgIntraSimilarity, the mean pairwise cosine similarity across its members, which is the number that tells you whether a grouping is meaningful. A cluster of four pages at 0.31 average similarity is not a cluster. It is a coincidence.
An honest gap analysis
content_gap_analysis is the tool most likely to produce a confident and wrong answer, so it is worth being precise about what it does.
It builds a TF-IDF model over your page plus every competitor page, so a term is weighted by how rare it is across the comparison set rather than how often it appears. Then it records every term that appears in a competitor and not in yours, keeping the number of competitor pages that used it and the best score it achieved.
Results rank by competitor hits first, then by score. A term three competitors use matters more than a term one competitor used with a higher weight, because the question is what the topic is expected to cover, not what one page happens to emphasise. Each page contributes its top 30 keywords to the comparison, and the merged list is capped at 25 results.
Each competitor entry also carries a similarity score: the cosine similarity between your page's vector and theirs. That single number is often more useful than the gap list, because it tells you whether you are writing about the same thing at all before you go chasing missing terms.
Failing in parts instead of failing entirely
Every multi-URL tool returns a result even when some fetches fail. Failures are collected and returned alongside the successful work in an errors array, each entry carrying the URL and the error code and message.
One competitor behind a bot wall does not cost you the analysis of the other nine. The alternative, returning a single error for the whole call, means the caller retries the entire comparison because one URL in a batch of ten was slow. For a tool that is usually called in a loop, partial results are the difference between a workflow that degrades and one that stops.
Twenty thousand characters, chosen because of a token cap
scrape_page returns at most maxChars of body text, defaulting to 20,000 and allowing up to 100,000, and it sets truncated: true when it cuts. The default is not arbitrary. Some MCP clients limit tool output to tens of thousands of tokens, and a raw page dump can blow through that on a long article, which produces a failed tool call rather than a large result.
The tool also never returns raw HTML. Body text comes back wrapped as untrusted content, and the clean extractors return structured fields rather than markup. That keeps the token cost predictable and keeps page-controlled markup out of the model's context.
Only one SERP provider, on purpose
serp_features works with SerpApi and nothing else. DataForSEO and Google Custom Search Engine were removed rather than left in a half-working state, and the config loader rejects those provider names with an explicit message telling you to use SerpApi, rather than silently falling back.
An integration that is known to be broken and is still advertised is worse than no integration, because somebody will build on it.
Security, stated plainly
Every hostname is resolved and every A and AAAA record must be public. The blocked set covers private ranges, loopback, CGNAT, link-local, documentation and benchmark ranges, the 6to4 relay range at 192.88.99.0/24, and the NAT64, 6to4 and IPv4-mapped forms of all of those.
IPv6 is an allow-list rather than a block-list, because the block-list approach does not scale to a space with that many ways to encode an address. Only global unicast 2000::/3 passes, minus 2001::/23 which includes Teredo, minus the documentation range 2001:db8::/32, minus 3fff::/20.
The HTTP client connects only to addresses it validated itself, so a hostname cannot rebind between the check and the connection. Redirects use redirect: "manual" and are re-validated per hop.
Response bodies are capped at 2 MB by default and extracted text at 100,000 characters. API keys are stripped from error messages and stderr logs, with query strings redacted, so a failed SerpApi call cannot write your key into a log file.
The server is stdio only, and the README says plainly not to expose it as a remote HTTP endpoint. That is the correct boundary for a tool that fetches arbitrary URLs on request: the only way to reach it should be a process you started yourself.
Testing and coverage
An 80 percent threshold is enforced across statements, branches, functions and lines. The suite covers the engines directly, the encoding path with deliberately mislabelled payloads, the robots.txt matcher including the cases that are easy to get backwards, and the headless renderer with a browser-based integration test (it needs Chromium installed to run) alongside unit tests that exercise the request guard and redirect handling without launching Chromium.
The security paths have their own test file, which is where the IPv6 allow-list, the blocked ranges and the redirect re-validation are asserted.
What is planned next
A stable JSON schema for every tool. The results are structured, and the schemas are not yet documented as a public contract. Anything scheduled or scripted against this tool needs that before it can rely on a field name.
Sitemap-aware competitor discovery. Right now the caller supplies the competitor URLs. Given a domain, the server could read the sitemap and pick the pages most likely to be worth comparing, which is the step that takes the longest by hand.
Caching keyed on content, not just URL. Content TTL is 300 seconds. Pages that have not changed should be reusable for longer, and a content hash would let the cache tell the difference between a page that was refetched and a page that was actually updated.
Takeaways for anyone building something similar
If you put a headless browser behind a scraper, the browser is now part of your attack surface, and it does not respect the network policy your HTTP client enforces. Taking the network away from it, and doing all fetching yourself, is more work than launching it and hoping. It is also the only version that holds up.
Do not build regexes from data you did not write. A robots.txt parser that compiles patterns supplied by the site being scraped hands that site a way to stall your process. Linear-time matching is not paranoia here, it is the basic requirement.
And decide what a partial result means before you need it. In a multi-URL tool, one failure out of ten should never cost the caller the other nine.