Marketing · Weekend build

Build your own Screaming Frog SEO Spider

A polite spider that only bites your own site. Screaming Frog SEO Spider charges $18–$22 per month — that’s $216–$264 a year — for something you can replace with focused software of your own. Here is the honest scope, the honest timeline, and the exact prompt to hand your coding agent.

from 10 hoursCrawl your own backyard

Who this replacement is for

This build targets a site owner auditing the handful of domains they actually control. The goal: crawl your own site on a schedule and get a readable, prioritized list of what to fix. If you need more than that, keep paying — the point of building it yourself is owning a tool shaped exactly like your workflow, not re-implementing a venture-funded roadmap.

How long it actually takes

One number would be a lie, so here are three. Each tier is a real, usable product — pick the one that matches how much of Screaming Frog SEO Spider you actually use.

EstimateWhat you get
10 hourspolite crawler, broken link report
4 daysredirect chains, meta audits, sitemap diff, scheduling
1 month+javascript rendering, million-url crawls, rank tracking — the part you should probably skip

What a minimal Screaming Frog SEO Spider alternative needs

Data model

Site, Crawl, Page, LinkEdge, Finding

Integrations

transactional email, Cloudflare Cron Triggers

Capability context

Astro, Workers, D1

The guardrail

Crawl only domains the owner has verified, honor robots directives and rate limits, and back off when a site starts answering with errors.

Deliberate non-goals

Do not crawl third-party sites, scrape competitors, track rankings, or grow into a general SEO suite.

The complete build prompt

Copy this into your coding agent of choice. It is scoped for a useful v1 — journeys, screens, business rules, data model, security, tests, and acceptance scenarios included. Pick your stack:

You are building a production-ready software product named “Crawlbug”, a deliberately focused alternative to Screaming Frog SEO Spider. Build a complete, usable vertical slice—not a landing page, static mockup, or disconnected collection of components.

WORKING AGREEMENT
Before writing implementation code, produce a short technical plan that names the routes or pages, server actions or endpoints, data tables, important state transitions, authorization boundaries, background jobs, and external adapters. Resolve contradictions in favor of the narrow audience and non-goals below. Prefer a small, legible architecture over speculative abstraction, but do not omit persistence, validation, error handling, or tests.

PRODUCT BRIEF
Primary user: a site owner auditing the handful of domains they actually control.
Primary outcome: crawl your own site on a schedule and get a readable, prioritized list of what to fix.
Product principle: optimize the exact workflow below instead of copying the full breadth of Screaming Frog SEO Spider. A first-time user should understand what to do from the interface itself, without a tour or documentation.

END-TO-END USER JOURNEYS
Implement all of these flows through the real interface and persistent data layer:
1. The owner verifies domain ownership, sets crawl scope and politeness limits, runs a first crawl, and reviews a prioritized findings report.
2. The owner fixes a batch of broken links and redirect chains, re-crawls, and confirms the findings show as resolved in the comparison with the previous crawl.
3. A scheduled weekly crawl finishes, diffs the sitemap against discovered pages, and emails a summary only when something actually changed.

SCREENS AND INFORMATION ARCHITECTURE
Build these as coherent responsive views. Each screen must specify its primary action, secondary actions, visible status, validation feedback, empty state, loading or pending state, success confirmation, and recoverable failure state.
1. Sites: verified domains, last crawl time, page count, open findings by severity, and run-crawl plus schedule actions.
2. Crawl detail: live progress, pages fetched, response-code breakdown, duration, and politeness stats such as request rate and robots exclusions.
3. Findings: broken links, redirect chains, title and meta issues, and image alt gaps, filterable by type and severity, each with its referencing pages.
4. Sitemap diff: pages in the sitemap but not crawlable, crawled pages missing from the sitemap, and changes since the previous crawl.

CORE CAPABILITIES
1. polite crawling with robots.txt respect, rate limits, and per-domain scope
2. broken internal and external links with the pages that reference them
3. redirect chains and loops with hop-by-hop detail
4. missing, duplicate, and overlong titles and meta descriptions plus image alt gaps
5. a sitemap diff showing pages missing from the sitemap or from the crawl

DETAILED BEHAVIOR AND BUSINESS RULES
Treat these as server-enforced product requirements, not interface suggestions:
1. Crawl only domains that have passed ownership verification and stop any crawl that strays outside its configured scope.
2. Honor robots.txt, enforce a per-domain request rate and concurrency cap, and back off or abort when the server returns repeated errors.
3. Follow redirects to a bounded depth, record every hop, and file chains and loops as findings instead of following them forever.
4. Attach findings to a specific crawl and compute new and resolved findings by comparing crawls, never by mutating historical results.

DATA MODEL AND LIFECYCLE
Design a small relational schema centered on Site, Crawl, Page, LinkEdge, Finding. Before implementing it, document:
1. Each table’s purpose, primary key, ownership or tenant boundary, timestamps, status fields, and important attributes.
2. Foreign keys, uniqueness constraints, check constraints, indexes needed by the named screens, and transaction boundaries for multi-record changes.
3. The allowed lifecycle or state transitions, who may trigger each transition, which transitions are terminal or reversible, and what audit history must remain immutable.
4. Archive, retention, and deletion behavior, including what happens to dependent records and external files.
5. Idempotency strategy for submissions, jobs, imports, notifications, webhooks, or retries where applicable.
Use migrations rather than ad-hoc schema creation. Store time instants consistently and retain named timezone context whenever local schedules or dates matter. Never rely on a counter, disabled button, or client-side check to preserve a business invariant.

USERS, AUTHENTICATION, AND PERMISSIONS
Implement only the roles required by the stated audience. Make the ownership and visibility model explicit before coding. Enforce authorization in every server-side query and mutation, including search, exports, attachments, live updates, and guessed URLs—not merely by hiding controls. Use secure session defaults, protect state-changing requests, and provide an understandable signed-out, expired-session, and forbidden state. Seed distinct users when multiple roles are required so permissions can be demonstrated and tested.

INTERACTION AND VISUAL DIRECTION
The product should feel fast, calm, focused, and credible rather than like a generic admin template. Use a clear visual hierarchy, restrained color, readable typography, generous hit targets, and consistent placement for primary actions. Start with server-rendered HTML and progressively enhance only the interactions that benefit from it. The core workflow must remain understandable if enhancement fails.

Start with server-rendered Astro pages and ordinary HTML forms. Use HTMX for form submissions, partial navigation, and server-driven updates, then Alpine.js only for small local browser state. The core workflow must remain understandable if either enhancement layer fails.

Design mobile layouts intentionally instead of simply stacking desktop panels. Support keyboard navigation, visible focus, semantic landmarks, explicit labels, useful page titles, reduced-motion preferences, and screen-reader announcements for asynchronous results. Never use color alone to communicate state. Destructive actions require clear scope and confirmation; safe repeated actions should be idempotent.

TECHNICAL DIRECTION
Build this version with the AHA stack: Astro for routing, layouts, and server-rendered pages; HTMX for interactions that benefit from HTML fragment responses; and Alpine.js for small, local interface state. Prefer Cloudflare D1 for relational persistence, R2 for object storage, Workers for server endpoints and scheduled work, Durable Objects only for coordinated real-time state, and Workflows or Queues for durable background jobs—but only when the product requirements call for them.

Keep domain rules in testable server-side modules instead of route handlers or UI components. Separate persistence, external providers, and background work behind small interfaces without building a framework. Prefer ordinary HTML forms and URLs for durable navigation; use optimistic interaction only when failure can be reconciled clearly.

The product brief currently identifies Astro, Workers, D1 as capability context. Preserve any required native, browser-only, edge, storage, real-time, or background-processing capability through a narrow adapter appropriate to the selected framework. If the core workflow genuinely requires native or browser APIs, keep that runtime as the primary execution surface rather than simulating inaccessible capabilities or inventing an unnecessary web surface.

Integrate with transactional email and Cloudflare Cron Triggers. For every integration:
- List required environment variables in an .env.example without real secrets.
- Add a small adapter with timeouts, normalized errors, and a deterministic local fake or development path.
- Verify inbound signatures and deduplicate provider events where supported.
- Keep credentials server-side, encrypt long-lived provider tokens at rest, and redact secrets and sensitive payloads from logs.
- Define retry, backoff, and idempotency behavior for any side effect that can be repeated.

SECURITY AND PRIVACY
Crawl only domains the owner has verified, honor robots directives and rate limits, and back off when a site starts answering with errors.
Validate, normalize, and length-limit all untrusted input on the server. Escape rendered content by default, sanitize any intentionally accepted markup, rate-limit public or abuse-prone actions, and use private object storage plus short-lived authorized URLs for sensitive files. Collect the minimum personal data necessary for the named workflow. Document retention and deletion behavior. Add specific protections for the riskier surfaces in this app, such as uploads, redirects, outbound requests, email delivery, OAuth, webhooks, CSV import or export, and real-time connections.

ACCEPTANCE SCENARIOS
Automate these app-specific scenarios at the most appropriate level:
1. Given a page returning 404 that three other pages link to, the report shows one broken-link finding listing all three referencing pages.
2. Given a redirect that eventually points back to itself, the crawl records the loop, stops following it, and files a single loop finding.
3. Given a crawl request for an unverified domain submitted through a forged form, the server rejects it and logs the attempt.

TESTING
Add focused unit tests for state transitions, authorization predicates, normalization, date or money calculations, and other risky domain rules. Add integration tests for persistence constraints and each external adapter’s success, timeout, retry, and rejection paths. Add at least one browser-level test for every end-to-end journey above, including one small-screen viewport. Tests must use isolated data and run through a documented single command.

OPERATIONS AND FAILURE RECOVERY
Add structured server logs with request, job, or event correlation IDs but no secrets or unnecessarily sensitive data. Make failures actionable in both the interface and logs. Background work must expose pending, succeeded, failed, and retrying states where relevant; do not silently swallow errors. Include safe database migration and rollback guidance, seed data, backup and restore notes, external-data cleanup behavior, and a basic health or diagnostic path appropriate to the stack.

DELIVERABLES
Ship the working application, migrations, representative seed data, tests, .env.example, and a concise README. The README must cover prerequisites, local setup, environment variables, migrations, seed and test commands, deployment, integration setup, backup and restore, security decisions, and known limitations. Seed data should exercise the happy path plus at least one empty, failed, overdue, expired, archived, or permission-restricted state relevant to the product.

DEFINITION OF DONE
The app is complete when a fresh developer can follow the README, create and migrate the database, run the app, sign in as each relevant role, complete every named journey using real persisted data, refresh without losing state, recover from common failures, and use the core interface on phone and desktop. All acceptance scenarios pass, permission boundaries are covered by tests, and no core screen is left as a placeholder.

NON-GOALS
Do not crawl third-party sites, scrape competitors, track rankings, or grow into a general SEO suite.

Frequently asked questions

How long does it take to build your own Screaming Frog SEO Spider?

A basic version — polite crawler, broken link report — takes about 10 hours. Roughly 4 days gets you a solid v1 with redirect chains, meta audits, sitemap diff, scheduling. Matching everything Screaming Frog SEO Spider really does (javascript rendering, million-url crawls, rank tracking) is closer to 1 month+, which is exactly why you should scope down instead.

How much does Screaming Frog SEO Spider cost if I keep subscribing?

Screaming Frog SEO Spider runs $18–$22 per month on public paid plans, which is $216–$264 per year. A focused self-built replacement costs your build time plus close-to-zero hosting.

What stack should I use to build a Screaming Frog SEO Spider alternative?

The build prompt on this page ships in four flavors: the AHA stack (Astro, HTMX, Alpine.js), Next.js, Laravel, and Ruby on Rails. The capability context for this product is Astro, Workers, D1. Pick the stack you already know — the scope matters more than the framework.

What features does a minimal Screaming Frog SEO Spider replacement need?

A useful v1 needs: polite crawling with robots.txt respect, rate limits, and per-domain scope; broken internal and external links with the pages that reference them; redirect chains and loops with hop-by-hop detail; missing, duplicate, and overlong titles and meta descriptions plus image alt gaps; a sitemap diff showing pages missing from the sitemap or from the crawl. Everything else is scope creep until you personally miss it.

What should I deliberately not build?

Do not crawl third-party sites, scrape competitors, track rankings, or grow into a general SEO suite.

Prompt copied. Go ship it.