Skip to content

Case Study

Stream of Worship

Role
Product Owner & Engineer
Team
Solo
Stack
Next.js 16 FastAPI AWS Lambda/SQS Neon Postgres + pgvector Cloudflare R2 PyTorch / Demucs Jetpack Compose (Android)

Problem

Chinese worship teams have no tooling for the space between songs. English-language worship has mature set-planning software; Chinese worship music (中文敬拜詩歌) lives in scattered catalogs with fragmented audio sources, no transition planning, and no offline-ready playback. Every Sunday, someone fumbles with a laptop between songs and the room loses the moment.

Goal: go from “here is my song list” to a seamless, offline-capable worship set — one continuous audio file plus a lyrics video with chapter markers — without manual audio engineering.

Constraints

Three hard realities shaped everything:

  1. Venue Wi-Fi is the least reliable thing in the room. A worship set that stalls mid-service is worse than no system at all — offline playback was a launch requirement, not a nice-to-have.
  2. Catalog data is messy. Song titles mix Traditional characters and pinyin, tempo detection on worship arrangements is ambiguous (octave errors), and synced lyrics don’t exist for most of the catalog.
  3. Budget. This is a solo side project — every GPU job and every render had to fit free-tier quotas and serverless cost envelopes.

Decisions

  • Monorepo, hard boundaries (9 components). Only the Dockerized Analysis Service may hold ML dependencies; the admin CLI and Android client are architecturally forbidden from touching Postgres, R2, or SQS directly (enforced by convention, documented in AGENTS.md). A solo multi-year project survives on boring, explicit interfaces.
  • Transition quality is a boundary problem, not a song-level problem. Compatibility is judged on the key a song leaves through versus the key the next one arrives through, plus boundary BPM — whole-song averages produce jarring seams.
  • Worship practice encoded as constraints. Sets follow a fixed five-phase arc (讚美 → 感恩 → 敬拜 → 奉獻 → 差遣); a beam-search songset constructor does constrained search over the catalog with a scored objective (theme, tempo flow, harmonic mix, diversity) and tunable hard-constraint relaxation.
  • Domain priors beat generic models. Tempo estimation adds worship-music BPM priors plus a lognormal prior derived from lyrics characters-per-second to resolve octave ambiguity that fools off-the-shelf estimators.
  • A real ML pipeline is mostly plumbing. The lyrics pipeline chains YouTube transcripts (with a circuit breaker) → cloud ASR → local Whisper → LLM alignment → a forced aligner; per-API quota waiters park jobs until free tiers reset; per-job-class semaphores keep batch backfill from starving live traffic.
  • Serverless render pipeline. Crossfaded multi-song MP3s and ~29 000-frame lyric MP4s render in an AWS Lambda container triggered via SQS, with SSE progress streamed to the browser — no long-running Vercel functions.
  • Offline-first playback. The web app caches rendered artifacts in a service worker (byte-range serving for large videos), and share links carry token-scoped frozen copies so a shared set never changes under the recipient (ADR 0009).

Showcase

This is real output from the songset constructor’s eval harness (eval/run1_standard, 2026-07-26): a beam search across a 438-song pool — 191,406 scored pairwise transitions — producing five ranked three-song sets. Pick a proposal to inspect its songs, transition notes, and objective scores.

Audio files are not committed to the repo — waveforms are illustrative. Song, key, BPM, crossfade, and score data are verbatim constructor output.

Showcase — songset constructor output

run songset-20260726T070852Z-3s-top5

Waveform bars are synthetic placeholders — no audio is shipped. Song, key, BPM, crossfade, and score values are verbatim constructor output.

The ranking itself is the interesting surface: proposal 1 dominates on tempo flow (0.642 vs 0.526) and harmonic mix (0.780 vs 0.620), yet proposals 3 and 4 land within 0.15 composite using entirely different openings — and the constructor accepted a 4-second crossfade with a +1 semitone key shift into the closing song in proposals 2 and 5 rather than rejecting them. Ranked options with visible tradeoffs, not a single “best” answer, is the product.

Artifacts

Outcome

  • Built to handle a 400+ song catalog with thousands of scored pairwise transitions (191,406 in the eval harness)
  • Process: ~500 spec docs with v1–v5 iteration chains, 11 ADRs, CI running real Postgres with pgvector, 150+ Vitest files on the web app alone.