Open to roles · San Francisco & the Midwest

Product Engineer · Production AI · San Francisco and the Midwest

I build production
AI systems end-to-end.

7 years full-stack, backend-leaning. I'm a software engineer at Pawservation, building a booking, payments, and Claude agent platform on Cloudflare Workers, with Stripe Connect, WhatsApp booking, and an OAuth 2.1 + PKCE MCP server. Looking for senior Product Engineer or Software Engineer roles in San Francisco, or in the Midwest if the role is worth relocating for.

Previously Salesforce, Total Brain, and Front. Since 2022, building and running Pawservation.

More than half of bookings now arrive self-serve across 30+ active clients and 60+ bookings a month, replacing 5 to 15 hours a month of manual work. 4+ outside pet-sitting businesses run on the platform.

Summary

Brad Burch is a product engineer in San Francisco with 7 years of experience across Salesforce, Total Brain, Front, and Pawservation. At Pawservation he builds a booking, payments, and AI platform on TypeScript and Cloudflare Workers: a Claude agent and an OAuth 2.1 + PKCE Model Context Protocol (MCP) server with human approval on every AI-made request, Stripe Connect payments, and WhatsApp booking, used by 30+ active clients and 4+ outside pet-sitting businesses.

He is looking for senior Product Engineer and Software Engineer roles, and selectively Forward Deployed Engineer roles, at Series B or later companies, in San Francisco or relocating to Chicago, Minneapolis-St. Paul, Detroit, Columbus, Indianapolis, or Ann Arbor.

The client portal, the multi-tenant booking platform, the AI surfaces on top of both, and open-source tools.

01

Brad Paws Client Portal

A React 19 + Hono portal on Cloudflare Workers that replaced 5 to 15 hours a month of manual invoicing, reconciliation, and scheduling for a live SF pet-care business.

Live

A scheduling, booking, and billing system that runs unattended for Brad Paws, a San Francisco pet-care business: 30+ active clients, 60+ bookings and 30+ invoices a month. It replaced 5 to 15 hours a month of manual invoicing, payment reconciliation, and scheduling. It is a TypeScript monorepo on Cloudflare Workers, D1, KV, R2, Queues, and Workflows with no origin server, fed by four sources it does not control: Notion, Gmail, Venmo exports, and Google Calendar. Most of the work is deciding which source wins when two of them disagree about what a client owes, and keeping four booking channels from double-booking a calendar that has no transactions.

React Router 8React 19HonoTypeScriptCloudflare WorkersD1KVR2QueuesWorkflowsRechartsVitestPlaywright
30+active clients
5 to 15 hrsof manual work replaced / month
154 ms to 18 msCPU time per invoice PDF
Architecture & Infrastructure
  • Built a TypeScript monorepo (client portal, Hono API Worker, scheduled sync Worker, shared domain package) on Cloudflare Workers, D1, KV, R2, and Workflows with no origin server, serving 30+ active clients and 60+ bookings a month
  • Chose a serverless stack so the system has nothing to scale and nothing to patch, with a documented path to Durable Objects if write concurrency demands it
  • Deduplicated records across four sources with ExternalId matching, UUID primary keys, and E.164 phone normalization, holding one client identity across Notion, Gmail, Venmo, and Calendar
Google Calendar Integration
  • Integrated Google Calendar v3 via service account, with KV-backed token caching on a 55-minute TTL so a token round trip happens once an hour instead of once a request
  • Guarded against double-booking across four booking channels plus reschedules and recurring series with a D1 lease that serializes every booking write, because Google Calendar has no transactions, backed by a capacity re-check at approval; a 10-way race test confirms one holder
  • Implemented multipart/mixed HTTP batching for Calendar writes with raw boundary construction, 50 PATCH requests per HTTP call, and 429/500/503 retry with exponential backoff, staying inside Cloudflare subrequest limits during migrations
Reliability: Workflows Backfill
  • Rebuilt a calendar backfill blocked by a 50-subrequest platform ceiling into a durable Cloudflare Workflows fan-out with one instance per date window, crash-safe cursor checkpointing, and an idempotent fan-in
  • Recovered backfills per date window after a crash, retrying the failed window without skipping or double-writing data
  • Isolated each sync source behind its own error boundary, with the daily summary reporting per-source failures rather than a false success on zero rows
  • Built observability for unattended automation: structured JSON logging, workflow run tables, heartbeat rows, cron failure emails, and a watchdog that catches silent never-ran failures
  • Consolidated four cron triggers into one hourly trigger, with heavy jobs as Cloudflare Workflows keyed by the hour, so a double-fired cron runs once and invoices cannot email twice
  • Added auth-failure and parser-gap alerts after an expired Google token froze calendar and payment sync for about two months and a regex bug of mine rejected 3 weeks of payment emails
Booking, Invoicing & Payment Pipeline
  • Built conflict-aware booking across boarding, house-sit, and walk/check-in with a server-enforced cancellation policy and automatic Google Calendar rollback on DB failure
  • Corrected booking capacity to count animals rather than calendar events, closing two over-limit paths
  • Anchored cancellation-fee tiers to the business's local business day, correcting a UTC-midnight bracket that mispriced late-evening cancellations
  • Single-sourced combined-set multi-pet pricing behind the quote, the calendar cost stamp, and the invoice, ending multi-pet stays quoted at double ($170 instead of $85)
  • Replaced date-ordered payment matching with closest-first allocation, ending a defect class that cleared a $600 stay with a $40 payment and flagged a $2,000 prepayment as delinquent; owners who share pets are grouped with union-find
  • Built Venmo payment ingestion with no Venmo API: CSV uploads from R2 and Gmail receipts accepted only when SPF and DKIM pass for venmo.com, normalized with Notion into idempotent events with per-source cursor commits, replacing manual cross-referencing of three payment systems
  • Automated monthly invoicing (pdf-lib synthesis, R2 storage, emailed delivery) with authenticated batch ZIP download behind ownership-verified endpoints, hardened by RFC 5987 filename encoding against header injection
  • Recovered four months of invoice runs killed partway by the Workers CPU limit: the logo was re-decoded for every PDF, and resizing it cut per-PDF CPU time from 154 ms to 18 ms; explicit completion markers with a 3-month catch-up replaced "month done once any invoice exists"
  • Added self-service invoice downloads and client-initiated cancellations against a fee preview running the server's own enforcement code, so the quoted fee and the charged fee cannot diverge
Testing & CI/CD
  • Audited the test suite assertion by assertion rather than trusting the coverage number, closing four security controls whose tests ran the code path without ever asserting the control held
  • Covered auth, booking, cancellation, calendar parsing, AI tools, and every dashboard component with Playwright end-to-end suites behind that audit
  • Added a web test suite inside the real Workers runtime alongside the Node suite (the sync Worker already ran there), after a fetch option Node accepts and Workers rejects took down the Claude connector
  • Moved an integration suite that ran only after merge into PR CI, after finding 5 of its 9 tests failing behind a green PR build
  • Gated every push on type generation, typecheck, ESLint, Prettier, and tests, with Cloudflare auto-deploy on merge to main and production source maps
02

Pawservation Booking Platform

A multi-tenant booking platform for independent pet sitters, with paid plans, Stripe Connect payments, an AI assistant, and WhatsApp booking, on Cloudflare Workers, D1, and Durable Objects. 4+ outside pet-sitting businesses now on it.

Live

Built from customer conversations and product feedback: booking, payments, and client messaging for pet sitters who do not have a developer, with 4+ outside pet-sitting businesses now on it. Sitters sign up self-serve for Solo ($15/mo) or Pro ($29/mo) on Stripe Checkout and install with one script tag or a hosted booking link. The paid tier adds Stripe Connect card payments with a post-stay auto-charge that runs at most once per stay, an AI assistant that checks every dollar figure against tool output, and WhatsApp booking on the sitter's own number. Every SQL query is scoped by business id, and the test data gives two businesses' clients the same email and phone, so any unscoped lookup fails visibly.

TypeScriptReactHonoCloudflare WorkersD1Durable ObjectsKVStripe ConnectStripe CheckoutWhatsApp Cloud APIClaudeVercel AI SDKMCPTurnstileViteVitestGitHub Actions
4+outside pet-sitting businesses on the platform
At most onceauto-charge per stay, raced with 20 concurrent writers
To the centevery dollar in an AI answer checked against tool output
Plans, Data & Import
  • Launched Solo ($15/mo) and Pro ($29/mo or $290/yr) plans on Stripe Checkout and the Customer Portal with a 30-day trial, keyed on paid invoices so a renewal in flight grants no free month
  • Built a Google Calendar import that turns past events into bookings; on a month of real calendar data it resolved 53 of 55 events and sent only the two it could not place for review
  • Fixed a D1 write-cap problem: the 15-minute calendar reconcile rewrote every external event each pass, the source of 97% of daily database writes (56,424 of 58,121) against a 100,000/day platform cap; a null-safe change guard on the upsert fixed it a week before the cap took effect
  • Caught a 16x payment over-allocation ($42,430 proposed against $2,640 owed) by testing the matcher on real data before release, so the defect never reached a customer
Payments: Stripe Connect
  • Guaranteed each stay is auto-charged at most once by claiming a D1 uniqueness constraint before Stripe is called, with a Stripe idempotency key so a retry is the same charge; 20 concurrent writers on the real D1 runtime land one row
  • Built card payments on each sitter's own Stripe Connect account: deposits, a client pay page, revocable saved-card consent, and an hourly auto-charge after each stay ends, with the sitter as merchant of record, card entry on Stripe-hosted pages, and no platform fee
AI & Channels
  • Made the assistant unable to misquote money: every dollar figure in an answer is checked to the cent against tool output, and any answer with an unverified figure is replaced
  • Required a server-rendered preview and a single-use, expiring token for every AI-initiated write, re-validated against live data on confirm; a typed "yes" never confirms anything
  • Built one shared tool layer behind web chat, WhatsApp, and the MCP server, so each new capability ships on all three channels at once and they cannot drift apart
  • Ran each sitter and each client conversation as its own Durable Object with built-in timers, so alerts, reminders, and AI turns fire on schedule and one conversation cannot slow another
Isolation & Security
  • Isolated tenant data across config, pricing, capacity, and bookings by scoping every SQL query through a single business-id module, tested with data that gives two businesses' clients the same email and phone, so any lookup not scoped by tenant fails visibly
  • Made paid features fail closed: entitlement is checked live (3 s timeout, 60 s cache) over whole route groups, and CI reads the router's route table so any new ungated route fails the build; a config outage denies paid features, never booking
Why it matters: the install path and the money path are the product. A sitter with no developer pastes one tag or shares a link, and 4+ outside businesses are now on the platform. Underneath, the guarantees a sitter cares about are enforced by the database and CI rather than by convention: a stay auto-charges at most once, an AI answer cannot state a dollar figure the tools did not return, and a paid route without an entitlement gate cannot merge.
03

Pawservation AI Chat Agent

A production Claude Haiku 4.5 booking agent on web chat and WhatsApp, with a human approving every request and nearly all AI-made requests approved as submitted.

Live

Clients book in plain language on web chat or WhatsApp. The agent runs on Claude Haiku 4.5 with fifteen tools wired to the live system. The agent proposes, the client confirms, and the business owner approves each request in one tap; nearly all AI-made requests are approved as submitted. The safety envelope is most of the work: every write needs a single-use token claimed atomically in D1, per-user and global USD caps, and a circuit breaker. I designed on the assumption that an injection eventually lands, so a compromised turn can propose a mutation but cannot execute one. A 47-case live-model eval suite gates every prompt change.

Claude Haiku 4.5Vercel AI SDK v6Prompt cachingLLM evalsWhatsApp Cloud APICloudflare QueuesTypeScriptHonoCloudflare WorkersD1
15 toolswired to live availability, quotes, and bookings
1 tapowner approval, nearly all approved as submitted
47 caseslive-model eval suite gates every prompt change
Agent & Safety Envelope
  • Shipped a customer-facing Claude Haiku 4.5 agent (Vercel AI SDK, 15 tools) on web chat and WhatsApp plus a remote MCP server that shares the same tool definitions
  • Designed human-in-the-loop booking: the client confirms, the business owner approves or declines in one tap from an admin queue or WhatsApp through one channel-agnostic approval service, and nearly all AI-made requests are approved as submitted; a pending request holds capacity
  • Gated every write behind an explicit confirmation card and a single-use token claimed atomically in D1, shared by every channel; it replaced a KV get-then-delete token, deleted about 170 lines, and 20 concurrent redemptions produce one winner
  • Kept cancel tokens minimal: a chat cancellation token carries only the event id, and the server re-derives the fee and everything else from the live event on confirm
  • Implemented layered cost controls: per-user rate limiting (20 msg/hour), daily token budgets (50k tokens/day), hard USD spend caps at the user and global level, and a circuit breaker (3 failures, then a 60s cooldown)
  • Made AI turn billing settle once per turn through one guarded path across stream finish, error, abort, and transport failure, so a timed-out turn cannot refund twice or reset the circuit breaker
  • Cut agent cost 4.9x per eval run by restructuring the system prompt into a byte-stable cached prefix plus a per-request context block, pinned by a test; about 95% of eval-suite input tokens now come from cache
  • Built a 47-case live-model eval suite (23 chat, 24 WhatsApp; 18 at launch) as the acceptance bar for prompt and model changes; within a day it caught a prompt bug that exposed three date defects
  • Replaced a prompt rule that was not working with a code fix: the WhatsApp agent invented dates in 2 to 3 of 5 runs, so relative-date messages now force a date-resolution tool call first
  • Benchmarked Claude Haiku against 8 open-weight models on Workers AI for cost and accuracy: Haiku, cached and uncached, scored 100% on the 10-config benchmark at 2.1 s mean latency, validating the model the agent was built on, while three open models broke a booking-safety rule
  • Closed a prompt-injection path by rebuilding conversation history on the server instead of trusting the client, where it could forge tool results such as a $0 quote
  • Built the WhatsApp pipeline on the Meta Cloud API with a webhook and Cloudflare Queues: HMAC signature verification, deduplication by message ID, per-client ordering, and capped retries
  • Fixed a client's WhatsApp cutoff within days of the complaint by removing per-client limits, keeping a global daily dollar cap against auto-reply loops, and adding per-message logging so the next refusal shows up in production
  • Designed chat persistence with a 20-message sliding window, automatic session rotation at 100 messages, and a 90-day retention cleanup cron
  • Threat-modeled the agent's attack surface (prompt injection, tool authority, data exfiltration) before wiring it to live billing and scheduling data, and bounded a compromised turn to proposals only: scope gating limits which tools it reaches, and the USD cap limits what it can cost before a human approves anything
Why it matters: a chat agent in front of a live billing and scheduling system is a liability without guardrails, and an agent nobody uses is not worth the guardrails. This one does both jobs: clients book through it, and the owner approves nearly everything it proposes as submitted. I put the circuit breaker at the agent layer rather than the model layer, so it still protects the downstream tools when the model itself is the thing misbehaving.
04

Pawservation MCP Server

A standards-compliant Model Context Protocol server with OAuth 2.1 + PKCE, stateless on Workers V8.

Spec-strict

Clients can book from their own Claude, Claude Desktop, Cursor, or any spec-compliant client. The MCP SDK supplies the protocol layer; the OAuth 2.1 + PKCE authorization server is hand-built and stateless across the Workers V8 fleet. Every write goes through the same preview, confirmation, and owner approval as web chat. The connector once went quiet for weeks because a fetch redirect option Node accepts is one Workers rejects, and every Node test stayed green. The fix was one line, followed by a CI suite that runs in the real Workers runtime and an auth-health alert. MCP Auth Kit (below) is the open-source version of the auth and safety layer.

MCP specOAuth 2.1PKCE (RFC 7636)HonoCloudflare Workers
RFC 7636hand-written PKCE
Statelessacross the V8 fleet
Real-runtime CINode-only bugs now fail the build
MCP Booking Server
  • Built the MCP booking server (@modelcontextprotocol/sdk v1.x, Streamable HTTP) so clients can book from Claude through 16 tools: the chat agent's 15 shared definitions plus a confirm tool that Claude clients use
  • Aligned MCP transport with Workers' V8 isolate model: stateless per-request server instances with closure-cached tool state, so no server-side sessions and no Durable Objects required
  • Centralized tool definitions across the chat agent and MCP server in a single source of truth (names, Zod schemas, MCP annotations), eliminating behavioral drift between access channels
  • Kept every MCP write behind the same preview, client confirmation, and owner approval as web chat, so a client's own model gets no more authority than the portal's agent
Reliability: the connector outage
  • Traced a Claude connector outage (9 weeks with no new connections, fully down for the last 5) from a client's screenshot to a fetch redirect option Node accepts and the Workers runtime rejects, using live log tailing and a local repro
  • Switched to a manual redirect that still refuses any 3xx, keeping the SSRF guard intact, and added a real-Workers-runtime CI suite and an auth-health alert so that bug class fails the build and a silent auth failure pages
Authentication & Security
  • Hand-wrote the OAuth 2.1 authorization server with PKCE (RFC 7636, S256 challenge verification, single-use codes, rotating hashed refresh tokens, dynamic client registration, SSRF-guarded client metadata fetches, revocation that fails closed), letting any spec-compliant client connect self-service
  • Rate-limited authorization attempts by IP against brute-force enumeration, with hashed-token audit logging on every OAuth event
  • Implemented cookie-based JWT auth (HS256 via jose) with HttpOnly/Secure/SameSite=Lax cookies and revocable server-side sessions, protecting the client dashboard from XSS and CSRF without a third-party auth provider
  • Added non-blocking D1 audit logging for all OAuth events, indexed by owner, action, and timestamp for queries that stay off the request path
05

MCP Auth Kit

A production-minded MCP server kit: OAuth 2.1 + PKCE, rate limiting, scope-gated tools, and two-phase confirmation.

Open source

The hard parts of a safe MCP server, packaged: OAuth 2.1 with PKCE, rate limiting, scope-gated tools, and two-phase confirmation on sensitive actions, unopinionated about tools, identity provider, and storage. The open-source version of the auth layer behind the MCP server above.

06

Debrief

A native macOS app that records job-interview calls locally, transcribes them on-device, and turns them into LLM-scored coaching feedback.

Open source

Records interview calls as two separate tracks (mic and ScreenCaptureKit), transcribes on-device with WhisperKit, and uses the Claude API to score them against a rubric, including a pasted company leveling guide. Audio flushes to disk in chunks, so a crash never loses a session. Swift, SwiftUI, and Swift Package Manager.

07

Roadrunner

A multi-user Django app that shares the nature you saw on your Strava activities, matching eBird and iNaturalist observations by time and location.

Open source

Writes the species you logged on eBird or iNaturalist into the matching Strava activity, using Strava OAuth and webhooks and a deferred 2/4/8-hour re-check queue for late checklists. Python, Django, and PostgreSQL on Vercel and Neon, with tests covering matching, sync, OAuth, and the queue.

7 years of customer-facing engineering, from Salesforce to Pawservation.

2022 to PresentSan Francisco
PawservationSoftware Engineer

Booking, payments, and AI platform for pet-care businesses, shaped by customer conversations and product feedback.

  • Shipped a client portal and three AI booking channels (web chat, WhatsApp, Claude via MCP), moving more than half of bookings to self-serve across 30+ active clients and 60+ bookings/month.
  • Replaced 5 to 15 hours/month of manual invoicing, payment reconciliation, and scheduling by syncing Google Calendar, Gmail, Notion, and Venmo into one booking and payment record.
  • Launched Solo ($15/mo) and Pro ($29/mo) plans on Stripe Checkout with a 30-day trial and Turnstile-protected self-serve signup, with 4+ outside pet-sitting businesses now on the platform.
  • Designed human-in-the-loop AI booking with client confirmation and one-tap owner approval from an admin queue or WhatsApp, with nearly all AI-made requests approved as submitted.
  • Guaranteed each stay is auto-charged at most once by claiming a D1 uniqueness constraint before calling Stripe, verified with 20 concurrent writers on the real D1 runtime.
  • Hand-built an OAuth 2.1 + PKCE authorization server for a remote MCP server (S256, rotating hashed refresh tokens, SSRF-guarded client metadata), letting clients book directly from Claude.
2022San Francisco
FrontSoftware EngineerImpacted by RIF

Shipped a Node.js REST API enabling account-level email rules for multi-domain enterprise customers, previously configurable only per-inbox. Performed a security upgrade on the core JavaScript email library responsible for 200K+ outbound messages per month. Role ended in a company-wide layoff.

2019 to 2021San Francisco
Total BrainSoftware Engineer

Engineering point of contact on enterprise customer-success and sales calls. Owned the HubSpot, Box, and SAML+JWT SSO integrations; migrated the backing store from MSSQL to DynamoDB (50% latency drop) and shipped a Scala backend integration for Apple, Google, and Facebook auth, contributing to a 60% lift in new sign-ups. Enterprise onboarding went from 2+ weeks to 3 days. Mentored junior engineers on backend patterns.

2017 to 2018San Francisco
SalesforceSoftware Engineer

Built an OAuth2 security filter for microservice API access control with remote identity validation. Separately, built a React management UI that internal engineers used to configure, launch, and inspect microservice test runs, cutting internal test setup time by 40%. Led delivery of a config builder and validator used by internal teams, scoping milestones and coordinating partner teams.

FocusFull-stack, backend-leaning · Serverless · distributed systems · REST APIs · webhooks
LanguagesTypeScript · JavaScript · Python · Swift · Java · Scala · SQL
FrameworksReact 19 · React Router 8 · Node.js · Hono · Vercel AI SDK · Django · SwiftUI
AI & LLMClaude (Haiku, Sonnet) · Model Context Protocol (MCP) servers · tool calling · LLM evals · prompt caching · human-in-the-loop approvals · agent safety envelopes (scope gating, spend caps) · Claude Code · Cursor
Cloud & DataCloudflare Workers, Durable Objects, Queues, Workflows · D1 · KV · R2 · AWS (DynamoDB) · PostgreSQL · MSSQL · SQLite
Payments & APIsStripe Connect, Checkout, and Billing · idempotency · webhooks · WhatsApp Cloud API · Google Calendar · Notion · Gmail · HubSpot · Box
Auth & SecurityOAuth 2.1 · PKCE · SAML 2.0 · JWT · SSO · HMAC webhook verification · SSRF prevention · multi-tenant isolation · rate limiting
Testing & DevOpsVitest (Node and Workers runtime) · Playwright · contract tests · GitHub Actions · CI/CD · monitoring and alerting

I'm a backend-leaning full-stack engineer in San Francisco, 7 years in. I started at Salesforce in 2017, worked through Total Brain and Front, and have been at Pawservation since 2022. The thread across all of it: I like the seam where engineering meets the customer. Pawservation came out of customer conversations and product feedback, and I own the loop from that feedback to the shipped, tested change.

At Total Brain I was the engineering point of contact on enterprise customer-success and sales calls while owning the integrations and a MSSQL-to-DynamoDB migration. That is where I learned what the "Forward Deployed Engineer" job description actually describes: sitting with the customer, scoping the integration, and then building it.

The work I am proudest of is the unglamorous money work. I replaced date-ordered payment matching with closest-first allocation after date ordering let a $40 payment clear a $600 boarding stay, and I made each stay auto-charge at most once by claiming a database uniqueness constraint before Stripe is ever called.

Next I want a senior Product Engineer or Software Engineer role at a Series B+ company in San Francisco or the Midwest, or a Forward Deployed Engineer role where the stack fits. San Francisco is home, and I would relocate to Chicago, Minneapolis-St. Paul, Detroit, Columbus, Indianapolis, or Ann Arbor for the right team.

How I verify AI-assisted code

I build with Claude Code and Cursor and treat the output as a draft I have to prove, through automation:

  • Tests in the real runtime. A CI suite runs inside the actual Workers runtime, after a fetch option that Node accepts took a Claude connector down.
  • Live-model evals. Prompt and model changes ship only after a 47-case suite passes against the live model.
  • Real data before release. Payment matching runs against real history first, which is how a 16x over-allocation got caught.

Brad Burch is looking for senior Product Engineer and Software Engineer roles (and selectively Forward Deployed Engineer) at Series B or later companies. Based in San Francisco, open to relocating to Chicago, Minneapolis-St. Paul, Detroit, Columbus, Indianapolis, or Ann Arbor. US citizen, no sponsorship required. Fastest reply is email.

→ Email Brad
Passing this along?

Copy this: "Brad Burch, product engineer in SF, 7 years, ex-Salesforce and Front. Software engineer at Pawservation: booking and payments platform with a Claude agent and an OAuth 2.1 + PKCE MCP server, used by 4+ outside pet-sitting businesses. Looking for senior Product or Software Engineer roles at Series B or later, SF or Midwest. bradburch.github.io, bradburch.jobs@gmail.com"