Files
Lyra/docs/superpowers/specs/2026-07-11-lyra-library-manager-design.md
Jonathan dcddcac86c docs: add library-manager design (slice 2)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 11:46:41 +02:00

17 KiB

Lyra — Library Manager (Design)

Date: 2026-07-11 Status: Approved design, ready for implementation planning Slice: 2 of 3 (Library manager)

Overview

Slice 2 turns Lyra from a one-shot acquisition tool into a standing library manager: you follow artists, monitor the specific releases you want, and Lyra keeps working in the background to obtain and upgrade them. It is Lidarr-style monitoring built directly on the slice 1 acquisition engine — the monitor's only output is the same Request + Job a manual request produces, so the entire existing pipeline is reused unchanged.

Two loops are equally central and share one monitoring engine:

  1. Watched artists → auto-grab. Follow an artist, browse their discography, and monitor the releases you want. Optionally flip a per-artist toggle so future releases auto-monitor as they appear.
  2. Wanted list → keep retrying. Every monitored release without an acceptable copy is "wanted" and is re-attempted on a schedule until a good copy lands — and, within a window, upgraded to better quality.

Goals

  • Follow artists and monitor releases at release-group granularity; obtain and keep them current automatically.
  • Reuse the slice 1 pipeline verbatim — monitoring enqueues normal jobs and reads their outcomes; the pipeline itself is untouched.
  • Never silently drop a want: a monitored release without a copy stays wanted and keeps being retried.
  • Keep sources and secrets in the worker. The web app performs only read-only, credential-free MusicBrainz metadata lookups for interactive search/browse.
  • Fully developable and testable with zero real credentials via a fake MusicBrainz browser and seeded Postgres.

Non-goals (this slice)

  • Notifications / webhooks on grab or needs-attention (a later slice).
  • Discovery / recommendations (slice 3).
  • A release-calendar / upcoming view.
  • Multi-user; this remains a single-user home-server tool.
  • Any change to how downloading, tagging, or the six-stage pipeline work.

Scope

In scope: watched artists (live MusicBrainz search + follow), a per-artist discography detail view with monitor toggles, a per-artist "auto-monitor future releases" toggle, a background monitor sweep (discovery + retry/upgrade), a cross-artist wanted list, and the outcome-reconcile that feeds job results back into monitoring state.

Out of scope: notifications, discovery, calendar. Deferred hardening from slice 1 is unaffected and remains deferred.

Key decisions

  • Default is passive. Following an artist marks their entire discography unmonitored. You opt in per release. A per-artist autoMonitorFuture toggle (default off) makes genuinely-new releases auto-monitor from the moment it is enabled.
  • No type filtering when auto-monitoring. When autoMonitorFuture is on, all official release-group types (albums, EPs, singles, live, compilations, …) auto-monitor; you unmonitor the noise by hand. ("Everything, filter later.")
  • Quality upgrades within a window. A quality cutoff and an upgrade window (N days) govern re-searching: below cutoff and within the window → keep chasing better quality; at/above cutoff, or window expired → frozen (done).
  • One table for discography and wanted list. A "wanted" item is simply a monitored release-group without a copy at cutoff. No separate wanted table, no sync problem.
  • Monitor lives in the worker. A scheduled sweep runs inside the existing worker loop. The web app owns CRUD + display, plus read-only MusicBrainz search for the interactive picker.

Architecture

No new containers. Two additions to the existing three-container stack:

  • Worker gains a monitor. Beside the job-claim loop, a cheap ~60s tick runs monitor.sweep(); after each pipeline run the loop calls monitor.reconcile(). All scheduled MusicBrainz work and all enqueuing live here, next to the pipeline it feeds.
  • Web gains monitoring UI + read-only MB reads. New pages/routes for watched artists, discography detail, and the wanted list, plus a thin MusicBrainz read module used only for the live search/browse picker.
┌───────────────────────────┐   Postgres    ┌──────────────────────────────┐
│  Next.js app (UI + API)   │──────────────▶│  Python worker               │
│                           │  requests/    │                              │
│  /artists  /artists/[id]  │  jobs +       │  claim → run_pipeline (S1)   │
│  /wanted   live MB picker │  monitor rows │  monitor.sweep()  (~60s tick)│
│                           │◀──────────────│  monitor.reconcile(job)      │
│  read-only MB search ─────┼──▶ MusicBrainz│  scheduled MB discovery ─────┼─▶ MB
└───────────────────────────┘  (metadata)   └──────────────────────────────┘

Separation of concerns. The web reasons about watched artists, monitored releases, requests, and jobs; the worker reasons about candidates, files, and scheduled acquisition. The web's only external reach is read-only, cred-free MusicBrainz metadata for the picker — every download source and every credential remains worker-only, as does the scheduled discovery poll and all pipeline-time MB resolution.

Data model

Two new tables, one new nullable column, and config keys in the existing key/value Config table (no migration for config).

WatchedArtist
  id
  mbid                (unique)                 # MusicBrainz artist MBID
  name
  autoMonitorFuture   Bool  default false      # auto-monitor new releases going forward
  monitorFrom         DateTime default now()   # boundary for "future" when the toggle is on
  lastPolledAt        DateTime?                 # throttles the discovery poll per artist
  createdAt
  monitoredReleases   MonitoredRelease[]

MonitoredRelease                                # discography row AND wanted-list row
  id
  watchedArtistId     FK?                       # null = standalone manual wanted add
  artistMbid, artistName                        # denormalized: standalone rows need no join
  rgMbid              (unique)                   # release-group MBID
  album
  primaryType         String?                    # "Album" | "EP" | "Single" | ...
  secondaryTypes      String[]                   # ["Live", "Compilation", ...]
  firstReleaseDate    String?                    # "YYYY-MM-DD" as MB reports it
  monitored           Bool  default false        # does the sweep act on it?
  state               enum(wanted|grabbed|fulfilled)  default wanted
  currentQualityClass Int?                        # quality of the copy we have (null = none)
  firstGrabbedAt      DateTime?                   # drives the upgrade window
  lastSearchedAt      DateTime?                   # throttles retry
  createdAt

Request  (existing) gains:
  monitoredReleaseId  String?                     # set when the monitor enqueues; null for manual

Config keys (worker-read, via existing Config):

key default meaning
monitor.enabled false master switch; off → worker behaves exactly as slice 1
monitor.pollIntervalHours 24 how often discovery re-browses each artist
monitor.retryIntervalHours 6 how often a wanted item is re-searched
monitor.qualityCutoff 2 the qualityClass at/above which a release is "done"
monitor.upgradeWindowDays 14 days after first grab to keep chasing upgrades

qualityClass (from slice 1's quality.py): 3 = hi-res lossless, 2 = CD lossless, 1 = lossy. The default cutoff 2 means "any lossless satisfies; only a lossy copy stays wanted and keeps chasing an upgrade."

How rows come to exist

  • Follow (live picker): the web already fetched the artist + discography from MusicBrainz to render the picker. On Follow it writes the WatchedArtist and persists the discography as MonitoredRelease rows, all unmonitored. The detail view is therefore instant.
  • Background discovery: the worker's poll re-browses each watched artist and inserts any newly appearing release-groups. If autoMonitorFuture is on and the release is new relative to monitorFrom, it is inserted monitored + wanted; otherwise unmonitored.
  • Standalone wanted add: adding an arbitrary album to the wanted list (not via a followed artist) resolves it to a release-group and writes a single MonitoredRelease with watchedArtistId = null, monitored = true.

The monitor (worker)

New module worker/lyra_worker/monitor.py. It uses an MbBrowser seam (below) and plain SQL over the shared connection; scheduling is driven entirely by timestamps in the tables, so the loop tick is just a cheap gate.

Scheduling

run_forever() records the last tick time; each loop iteration, if monitor.enabled and ≥ ~60s elapsed, it calls monitor.sweep(conn, browser, cfg) before continuing to claim jobs. After every run_pipeline it calls monitor.reconcile(conn, job_id). One process, one loop, no new infra.

sweep() — two phases

Phase 1 — Discovery (per artist where now - lastPolledAt > pollIntervalHours):

  • browser.browse_release_groups(artist.mbid) → release-group list.
  • For each rg not already a MonitoredRelease, insert it. monitored = artist.autoMonitorFuture and firstReleaseDate > artist.monitorFrom; state wanted. (All types eligible — no filtering.)
  • Set lastPolledAt = now.

Phase 2 — Retry / enqueue (rows with monitored = true, state != fulfilled):

  • Skip rows with an active (non-terminal) job for the same release.
  • Skip / satisfy via dedupe: if a LibraryItem at/above cutoff already exists → set fulfilled.
  • For state = grabbed: if now - firstGrabbedAt >= upgradeWindowDays or currentQualityClass >= cutoff → set fulfilled and skip; else it is an upgrade candidate.
  • For due rows (lastSearchedAt null or older than retryIntervalHours): create Request(artist, album, monitoredReleaseId=id) + Job(state=requested); set lastSearchedAt = now. The slice 1 pipeline takes it from there.

reconcile(job_id) — feed outcomes back

After the pipeline finishes a job that has a monitoredReleaseId:

  • imported: set currentQualityClass from the winning candidate and firstGrabbedAt (if null); set state = fulfilled if qualityClass >= cutoff, else grabbed.
  • needs_attention: leave state = wanted (already has lastSearchedAt, so it waits out the retry interval before the next attempt).

Reconcile is a separate call in the main loop, so the pipeline code is untouched and the whole feature is inert when monitor.enabled = false.

MonitoredRelease state machine

discovery / follow
   ├─(autoMonitorFuture & new)─► monitored + wanted ─┐
   └─(default / back-cat)──────► unmonitored + wanted │  user toggle → monitored
                                                       ▼
   wanted ──sweep enqueues──► [job] ──imported < cutoff──► grabbed ──window expired / at cutoff──► fulfilled
      ▲                          │                            │
      └─────needs_attention──────┘        imported ≥ cutoff───┴──────────────────────────────► fulfilled

Unmonitor at any time → dormant (the sweep ignores it). A want is never dropped; it is retried on its interval until satisfied.

MusicBrainz access

The worker's existing MusicBrainzResolver handles pipeline-time album resolution. Monitoring adds a small browsing seam.

Worker — MbBrowser protocol (real impl via musicbrainzngs, lazily imported; faked in tests):

  • search_artist(name) -> list[ArtistHit]{mbid, name, disambiguation}.
  • browse_release_groups(artist_mbid) -> list[ReleaseGroupInfo]{rgMbid, title, primaryType, secondaryTypes, firstReleaseDate}.

Web — read-only MB read module for the live picker only: artist search and discography browse, returning the same shapes the web persists as MonitoredRelease rows on Follow. No credentials; read-only; metadata only.

Web UI & API

Three surfaces, thin over the two tables:

  • /artists — watched artists (name, #monitored / #have), Add artist (live search picker), remove. Per-artist autoMonitorFuture toggle.
  • /artists/[id] — the discography: every MonitoredRelease for the artist, each with a status badge (Have@quality / Wanted / Dormant), a monitor toggle, and a Search now button (enqueues immediately).
  • /wanted — cross-artist list of everything monitored and not fulfilled (includes standalone manual adds and needs-attention retries), with last-searched time and Search now. Plus Add to wanted (resolve an arbitrary artist+album → standalone monitored row).

API routes (Next.js):

  • GET /api/mb/artists?q= — live artist search (read-only MB).
  • GET /api/mb/artists/[mbid]/releases — discography browse (read-only MB).
  • POST /api/artists — follow: create WatchedArtist + persist discography as unmonitored MonitoredRelease rows.
  • GET /api/artists, GET /api/artists/[id], DELETE /api/artists/[id], PATCH /api/artists/[id] (toggle autoMonitorFuture).
  • GET /api/wanted — monitored, not-fulfilled rows across artists.
  • POST /api/wanted — standalone add (resolve + create row).
  • PATCH /api/releases/[id] — toggle monitored.
  • POST /api/releases/[id]/search — enqueue now (Request + Job).

"Search now" and standalone adds create the same Request + Job the monitor does (with monitoredReleaseId set), so reconcile handles their outcomes too.

Integration touch-point: upgrades must replace

Quality upgrades re-import an album that is already in the library. LibraryItem is @@unique([artist, album]), so the import stage must upsert the row and overwrite the on-disk files rather than error on a duplicate. Current behavior will be verified during planning; if it errors today, the minimal import-upsert change is included in this slice. This is the only pipeline-adjacent change; the six stages otherwise stand unchanged.

Error handling & edge cases

Situation Handling
New release not yet on any source Job → needs_attention; row stays wanted; retried each interval until a copy appears (natural pre-order behavior).
Artist name ambiguous The live picker shows MB disambiguation; the user chooses. No blind auto-match.
Duplicate discovery MonitoredRelease.rgMbid is unique; re-browsing never duplicates rows.
Piled-up jobs Sweep skips a release that already has a non-terminal job.
Already in library Dedupe against LibraryItem; at/above cutoff → fulfilled without enqueuing.
Upgrade never improves Upgrade window expires → fulfilled (frozen); no endless re-searching.
MB unreachable during sweep Discovery for that artist is skipped this tick; retried next poll. Web picker surfaces the error to the user.
Monitor disabled monitor.enabled = false → no sweep, no reconcile effects; worker behaves exactly as slice 1.

Testing strategy

Same philosophy as slice 1 — logic is testable with zero real credentials.

  • Monitor unit testsFakeMbBrowser (canned release-group lists) + seeded Postgres. Assert: discovery sets monitored correctly (default off; on only when autoMonitorFuture and new-vs-monitorFrom); sweep enqueues only due rows; respects poll/retry intervals; skips releases with active jobs; applies cutoff and upgrade-window transitions; dedupes against LibraryItem.
  • Reconcile tests — a finished job in each end state updates its MonitoredRelease correctly (imported below/above cutoff; needs_attention).
  • Web API route tests — artists/wanted/releases CRUD and toggles against seeded Postgres; the MB read routes tested against a stubbed MB read module.
  • Opt-in live test (LYRA_LIVE_TESTS) — browse a real artist's discography and search a real artist name.
  • Backward-compat — with monitor.enabled = false, the full existing worker suite passes unchanged and no monitoring side effects occur.

Goal: the entire library manager is built and tested on a laptop with fakes and a seeded database; real MusicBrainz enters only for the opt-in live test.

Notes: after this slice

Slice 2 makes Lyra a standing library manager: follow artists, monitor exactly the releases you want (or auto-monitor future ones per artist), and let the background sweep obtain and upgrade them through the slice 1 pipeline, with a wanted list that never forgets. Deferred to later: notifications, a release calendar, and slice 3 (discovery), each with its own spec → plan → implementation cycle.