Productivity · Side quest build

Build your own Wispr Flow

Talk naturally. Get clean text where the cursor is. Wispr Flow charges $15–$39 per month — that’s $180–$468 a year — for something you can replace with focused software of your own. Here is the honest scope, the honest timeline, and the exact prompt to hand your coding agent.

from 2 daysA focused native side quest

Who this replacement is for

This build targets one Mac user who writes frequently by voice. The goal: turn push-to-talk dictation into corrected text without retaining a library of recordings. If you need more than that, keep paying — the point of building it yourself is owning a tool shaped exactly like your workflow, not re-implementing a venture-funded roadmap.

How long it actually takes

One number would be a lie, so here are three. Each tier is a real, usable product — pick the one that matches how much of Wispr Flow you actually use.

EstimateWhat you get
2 dayspush-to-talk capture, transcription, paste at cursor
1 weekstreaming transcription, personal dictionary, cleanup modes, undo
1 month+low-latency on-device models, per-app formatting, windows port — the part you should probably skip

What a minimal Wispr Flow alternative needs

Data model

DictationSession, TranscriptDraft, DictionaryEntry, RewritePreset

Integrations

macOS microphone, Accessibility API, speech-to-text API, LLM API

Capability context

SwiftUI, Swift, macOS

The guardrail

Request microphone and Accessibility permissions separately, never capture outside an active session, and delete audio immediately after transcription.

Deliberate non-goals

Do not build meeting recording, voice cloning, team analytics, mobile apps, or background listening.

The complete build prompt

Copy this into your coding agent of choice. It is scoped for a useful v1 — journeys, screens, business rules, data model, security, tests, and acceptance scenarios included. Pick your stack:

You are building a production-ready software product named “Sayflow”, a deliberately focused alternative to Wispr Flow. Build a complete, usable vertical slice—not a landing page, static mockup, or disconnected collection of components.

WORKING AGREEMENT
Before writing implementation code, produce a short technical plan that names the routes or pages, server actions or endpoints, data tables, important state transitions, authorization boundaries, background jobs, and external adapters. Resolve contradictions in favor of the narrow audience and non-goals below. Prefer a small, legible architecture over speculative abstraction, but do not omit persistence, validation, error handling, or tests.

PRODUCT BRIEF
Primary user: one Mac user who writes frequently by voice.
Primary outcome: turn push-to-talk dictation into corrected text without retaining a library of recordings.
Product principle: optimize the exact workflow below instead of copying the full breadth of Wispr Flow. A first-time user should understand what to do from the interface itself, without a tour or documentation.

END-TO-END USER JOURNEYS
Implement all of these flows through the real interface and persistent data layer:
1. The user grants microphone permission, configures a global shortcut, holds it in a text editor, speaks a paragraph, previews the transcript, and inserts it at the cursor.
2. The user adds a frequently misheard name to the dictionary, dictates it again, chooses a polished cleanup mode, accepts the preview, and can undo the insertion.
3. The user encounters a microphone or insertion-permission failure, sees exactly which permission is missing, copies the transcript manually, and retries after changing settings.

SCREENS AND INFORMATION ARCHITECTURE
Build these as coherent responsive views. Each screen must specify its primary action, secondary actions, visible status, validation feedback, empty state, loading or pending state, success confirmation, and recoverable failure state.
1. Setup: microphone and Accessibility permission states, shortcut recorder, provider configuration, language, retention choice, and a test-dictation action.
2. Floating capture panel: listening state, timer, input level, cancel, stop, processing, and an obvious boundary between captured and idle states.
3. Preview: literal transcript, cleanup mode, highlighted rewrite differences, edit, copy, insert at cursor, retry, and undo guidance.
4. Dictionary and history: personal terms, import and export, recent text-only sessions when enabled, deletion controls, and provider health.

CORE CAPABILITIES
1. global push-to-talk shortcut and clear listening indicator
2. streaming transcription with punctuation and paragraph breaks
3. small personal dictionary for names and technical terms
4. optional concise, polished, and literal cleanup modes
5. preview, copy, and type-at-cursor delivery with undo

DETAILED BEHAVIOR AND BUSINESS RULES
Treat these as server-enforced product requirements, not interface suggestions:
1. Open the microphone only while an explicit session is active, display system and in-app capture indicators, and discard captured audio immediately after transcription or cancellation.
2. Keep transcription and rewriting separate so a failed rewrite never destroys the literal transcript; show a diff and require acceptance before inserting rewritten text.
3. Inject text only into the currently focused editable control after fresh Accessibility authorization; otherwise fall back to copying without simulating hidden keystrokes.
4. Normalize and length-limit dictionary entries, keep them local by default, and never send unrelated clipboard, window, or surrounding document content to a provider.

DATA MODEL AND LIFECYCLE
Design a small relational schema centered on DictationSession, TranscriptDraft, DictionaryEntry, RewritePreset. Before implementing it, document:
1. Each table’s purpose, primary key, ownership or tenant boundary, timestamps, status fields, and important attributes.
2. Foreign keys, uniqueness constraints, check constraints, indexes needed by the named screens, and transaction boundaries for multi-record changes.
3. The allowed lifecycle or state transitions, who may trigger each transition, which transitions are terminal or reversible, and what audit history must remain immutable.
4. Archive, retention, and deletion behavior, including what happens to dependent records and external files.
5. Idempotency strategy for submissions, jobs, imports, notifications, webhooks, or retries where applicable.
Use migrations rather than ad-hoc schema creation. Store time instants consistently and retain named timezone context whenever local schedules or dates matter. Never rely on a counter, disabled button, or client-side check to preserve a business invariant.

USERS, AUTHENTICATION, AND PERMISSIONS
Implement only the roles required by the stated audience. Make the ownership and visibility model explicit before coding. Enforce authorization in every server-side query and mutation, including search, exports, attachments, live updates, and guessed URLs—not merely by hiding controls. Use secure session defaults, protect state-changing requests, and provide an understandable signed-out, expired-session, and forbidden state. Seed distinct users when multiple roles are required so permissions can be demonstrated and tested.

INTERACTION AND VISUAL DIRECTION
The product should feel fast, calm, focused, and credible rather than like a generic admin template. Use a clear visual hierarchy, restrained color, readable typography, generous hit targets, and consistent placement for primary actions. Start with server-rendered HTML and progressively enhance only the interactions that benefit from it. The core workflow must remain understandable if enhancement fails.

Start with server-rendered Astro pages and ordinary HTML forms. Use HTMX for form submissions, partial navigation, and server-driven updates, then Alpine.js only for small local browser state. The core workflow must remain understandable if either enhancement layer fails.

Design mobile layouts intentionally instead of simply stacking desktop panels. Support keyboard navigation, visible focus, semantic landmarks, explicit labels, useful page titles, reduced-motion preferences, and screen-reader announcements for asynchronous results. Never use color alone to communicate state. Destructive actions require clear scope and confirmation; safe repeated actions should be idempotent.

TECHNICAL DIRECTION
Build this version with the AHA stack: Astro for routing, layouts, and server-rendered pages; HTMX for interactions that benefit from HTML fragment responses; and Alpine.js for small, local interface state. Prefer Cloudflare D1 for relational persistence, R2 for object storage, Workers for server endpoints and scheduled work, Durable Objects only for coordinated real-time state, and Workflows or Queues for durable background jobs—but only when the product requirements call for them.

Keep domain rules in testable server-side modules instead of route handlers or UI components. Separate persistence, external providers, and background work behind small interfaces without building a framework. Prefer ordinary HTML forms and URLs for durable navigation; use optimistic interaction only when failure can be reconciled clearly.

The product brief currently identifies SwiftUI, Swift, macOS as capability context. Preserve any required native, browser-only, edge, storage, real-time, or background-processing capability through a narrow adapter appropriate to the selected framework. If the core workflow genuinely requires native or browser APIs, keep that runtime as the primary execution surface rather than simulating inaccessible capabilities or inventing an unnecessary web surface.

Integrate with macOS microphone and Accessibility API and speech-to-text API and LLM API. For every integration:
- List required environment variables in an .env.example without real secrets.
- Add a small adapter with timeouts, normalized errors, and a deterministic local fake or development path.
- Verify inbound signatures and deduplicate provider events where supported.
- Keep credentials server-side, encrypt long-lived provider tokens at rest, and redact secrets and sensitive payloads from logs.
- Define retry, backoff, and idempotency behavior for any side effect that can be repeated.

SECURITY AND PRIVACY
Request microphone and Accessibility permissions separately, never capture outside an active session, and delete audio immediately after transcription.
Validate, normalize, and length-limit all untrusted input on the server. Escape rendered content by default, sanitize any intentionally accepted markup, rate-limit public or abuse-prone actions, and use private object storage plus short-lived authorized URLs for sensitive files. Collect the minimum personal data necessary for the named workflow. Document retention and deletion behavior. Add specific protections for the riskier surfaces in this app, such as uploads, redirects, outbound requests, email delivery, OAuth, webhooks, CSV import or export, and real-time connections.

ACCEPTANCE SCENARIOS
Automate these app-specific scenarios at the most appropriate level:
1. Given the shortcut is pressed while idle, one capture session opens; releasing it stops capture and produces one transcript without retaining the audio.
2. Given Accessibility permission is denied, transcription still completes, insertion is not attempted, and copy remains available with actionable recovery text.
3. Given a polished rewrite changes wording, the preview shows differences and rejecting it leaves the literal transcript available and unchanged.

TESTING
Add focused unit tests for state transitions, authorization predicates, normalization, date or money calculations, and other risky domain rules. Add integration tests for persistence constraints and each external adapter’s success, timeout, retry, and rejection paths. Add at least one browser-level test for every end-to-end journey above, including one small-screen viewport. Tests must use isolated data and run through a documented single command.

OPERATIONS AND FAILURE RECOVERY
Add structured server logs with request, job, or event correlation IDs but no secrets or unnecessarily sensitive data. Make failures actionable in both the interface and logs. Background work must expose pending, succeeded, failed, and retrying states where relevant; do not silently swallow errors. Include safe database migration and rollback guidance, seed data, backup and restore notes, external-data cleanup behavior, and a basic health or diagnostic path appropriate to the stack.

DELIVERABLES
Ship the working application, migrations, representative seed data, tests, .env.example, and a concise README. The README must cover prerequisites, local setup, environment variables, migrations, seed and test commands, deployment, integration setup, backup and restore, security decisions, and known limitations. Seed data should exercise the happy path plus at least one empty, failed, overdue, expired, archived, or permission-restricted state relevant to the product.

DEFINITION OF DONE
The app is complete when a fresh developer can follow the README, create and migrate the database, run the app, sign in as each relevant role, complete every named journey using real persisted data, refresh without losing state, recover from common failures, and use the core interface on phone and desktop. All acceptance scenarios pass, permission boundaries are covered by tests, and no core screen is left as a placeholder.

NON-GOALS
Do not build meeting recording, voice cloning, team analytics, mobile apps, or background listening.

Frequently asked questions

How long does it take to build your own Wispr Flow?

A basic version — push-to-talk capture, transcription, paste at cursor — takes about 2 days. Roughly 1 week gets you a solid v1 with streaming transcription, personal dictionary, cleanup modes, undo. Matching everything Wispr Flow really does (low-latency on-device models, per-app formatting, windows port) is closer to 1 month+, which is exactly why you should scope down instead.

How much does Wispr Flow cost if I keep subscribing?

Wispr Flow runs $15–$39 per month on public paid plans, which is $180–$468 per year. A focused self-built replacement costs your build time plus close-to-zero hosting.

What stack should I use to build a Wispr Flow alternative?

The build prompt on this page ships in four flavors: the AHA stack (Astro, HTMX, Alpine.js), Next.js, Laravel, and Ruby on Rails. The capability context for this product is SwiftUI, Swift, macOS. Pick the stack you already know — the scope matters more than the framework.

What features does a minimal Wispr Flow replacement need?

A useful v1 needs: global push-to-talk shortcut and clear listening indicator; streaming transcription with punctuation and paragraph breaks; small personal dictionary for names and technical terms; optional concise, polished, and literal cleanup modes; preview, copy, and type-at-cursor delivery with undo. Everything else is scope creep until you personally miss it.

What should I deliberately not build?

Do not build meeting recording, voice cloning, team analytics, mobile apps, or background listening.

Prompt copied. Go ship it.