The library carried titles, systems and genres but almost nothing else:
developer was 3.8% filled, publisher 2.9%, description 0%. These are columns
the app has always had and never been able to populate.
enrich_metadata.py reads them off the same articles the cover fetcher
locates. Only empty fields are touched unless --overwrite is given.
year 79% -> 96%
developer 3.8% -> 95%
publisher 2.9% -> 96%
description 0% -> 96%
Parsing infoboxes needed several guards, each found by checking output
rather than trusting the first pass:
* "Infobox video game" is a substring of "Infobox video game series", so
the loose test resolved Banjo-Kazooie to the series overview. Now
rejected, which also fixes the cover fetcher's article resolution.
* An article spans every release and its date block leads with the
original, so year is only filled when the article covers that platform.
Otherwise a DS port inherits the SNES original's year.
* A search hit that neither covers the platform nor closely matches the
title is discarded: "Dragon Ball Z Budokai" surfaces "Shin Budokai", a
different game on a different console. Left blank instead.
* Values are grouped under bold platform headings, tagged with region
codes, annotated with the platform in parentheses, and wrapped in
templates whose named parameters leak through. Each of those read as
the developer or publisher before being handled.
Developer, publisher and year are facts and written verbatim. Descriptions
are article summaries under CC BY-SA, stored with an attribution line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
11 KiB
LudosData
A personal video game library: catalogue what you own, what you've dumped, played and finished.
Originally built in 2018 on Angular 5 + PHP + MySQL. Rebuilt in 2026 on Angular 22 and ASP.NET Core 10 with SQLite, running in Docker.
Quick start
cp .env.example .env
# Generate a signing key and put it in .env as JWT_KEY:
openssl rand -base64 48
# Also set SEED_USERNAME / SEED_EMAIL / SEED_PASSWORD for the first account.
docker compose up --build
Then open http://localhost:8080 and sign in with the seed credentials.
On first run the API creates that account and imports the 105 games recovered from the 2018 database dump. Seeding only happens while the database has no users.
Password rules: 12+ characters, with an uppercase, a lowercase and a digit. The API refuses to start if
JWT_KEYis missing or shorter than 32 characters — that is deliberate, so a misconfigured deployment fails loudly instead of signing tokens with a guessable key.
Layout
backend/ ASP.NET Core 10 Web API (C#)
src/LudosData.Api/
Domain/ Game, AppUser
Data/ DbContext, migrations, seeder, games.json
Auth/ JWT options, token service
Controllers/ auth, games, images
Services/ image storage
frontend/ Angular 22 SPA
src/app/
core/ models, services, guard, HTTP interceptor
features/ login, register, game-grid, game-edit, account
shared/ toolbar, confirm dialog
archive/ the original 2018 MySQL dump, for provenance
Everything stateful lives in one Docker volume (ludos-data): the SQLite file,
uploaded box art, and the Data Protection keys. Back that volume up and you have
backed up the whole application.
Development
Host tooling (Node 24, .NET 10) is installed via Homebrew. dotnet-ef needs
~/.dotnet/tools on PATH, which ~/.bashrc.d/dotnet.sh sets up.
# API on http://localhost:5099
cd backend/src/LudosData.Api
Jwt__Key="a-dev-key-of-at-least-32-characters!!" dotnet run
# SPA on http://localhost:4200, proxying /api and /uploads to :5099
cd frontend
npm start
cd frontend && npm test # vitest
cd backend && dotnet build # 0 warnings expected
Cover art
tools/library/fetch_art.py fills in box art, pushing each image through the
app's own POST /api/images so it gets the same validation and WebP re-encoding
as a manual upload. Standard library only — no virtualenv, and neither source
needs an account.
It tries two sources in order:
- libretro-thumbnails — scanned retail
boxes, named to the No-Intro / Redump conventions. Best art where it has any,
but its coverage is the retro consoles. Its
Microsoft - Xbox 360set exists but holds about a dozen entries. - English Wikipedia — a cover on essentially every notable game article,
which is what fills the Xbox 360 shelf. The exact filename is read from the
article's infobox rather than guessed from file names, since filtering names
for "box" also matches
Xbox-360-Pro-wController.png. Calls are throttled to one per second and cached; the API returns 429 if pushed harder.
cd tools/library
python3 fetch_art.py --password '...' --dry-run # report matches, change nothing
python3 fetch_art.py --password '...' # download and attach
python3 fetch_art.py --password '...' --overwrite # also replace existing art
python3 fetch_art.py --password '...' --no-wikipedia # libretro only
Always dry-run first; it prints every match with a similarity score, marks the
source (W for Wikipedia), and flags anything below 0.95 for eyeballing.
Matching handles the gaps between a personal catalogue and a ROM-naming one:
accents (Pokemon → Pokémon), roman numerals (our SNES Final Fantasy 2 is
the catalogue's Final Fantasy II), trailing articles (Sims 2, The), missing
subtitles in either direction, and outright typos — Brett Hull Hocky 95 finds
Brett Hull Hockey 95. A sequel guard stops Donkey Kong Country 2 from
silently taking Donkey Kong Country's box.
All 105 games currently have art: 93 from libretro, 12 from Wikipedia.
Two rows got art of the right game but the wrong platform, because the
platform in the source data looks wrong — a Game Boy "Donkey Kong Country 2"
(never released on that system; the handheld sequels were Donkey Kong Land)
and a DS "Donkey Kong Country Returns" (a Wii game, later Returns 3D on 3DS).
Fix the system field and re-run with --overwrite to correct them.
Art is publisher copyright. Fetching it for a private collection is ordinary practice for library software; redistributing it is a different question.
Metadata enrichment
tools/library/enrich_metadata.py fills developer, publisher, year and
description from the same Wikipedia articles the cover fetcher locates. Only
empty fields are touched unless --overwrite is passed — anything typed by hand
outranks anything derived here.
cd tools/library
python3 enrich_metadata.py --password '...' --dry-run
python3 enrich_metadata.py --password '...'
python3 enrich_metadata.py --password '...' --fields developer,publisher
Coverage went from this to this:
| Field | Before | After |
|---|---|---|
| year | 79% | 96% |
| developer | 3.8% | 95% |
| publisher | 2.9% | 96% |
| description | 0% | 96% |
Developer, publisher and year are facts, written verbatim. Descriptions are article summaries, which are CC BY-SA, so each is stored with an attribution line naming the source article.
Parsing an infobox is messier than it looks, and the guards matter:
- Series articles are rejected. A substring test for
Infobox video gamealso matchesInfobox video game series, which resolved Banjo-Kazooie to the series overview instead of the 1998 game. - Year is only filled when the article covers that platform. An article spans every release, and its date block leads with the original — so a DS port would otherwise be dated to the SNES original.
- A search hit that neither covers the platform nor closely matches the title is discarded. Our "Dragon Ball Z Budokai" surfaces "Dragon Ball Z: Shin Budokai", a different game on a different console. A blank field beats a confidently wrong one.
- Platform headings (
'''PlayStation'''), region codes (JP,NA), trailing platform annotations (Rare (N64)) and named template parameters (title=) are all stripped, since each one otherwise reads as the value itself.
Four games have no usable article: a typo'd title (Brett Hull Hocky 95),
Dragon Ball Z Budokai, and two niche releases.
Database changes
cd backend/src/LudosData.Api
dotnet ef migrations add <Name> --output-dir Data/Migrations
Migrations are applied automatically at startup.
API
All /api/games and /api/images routes require Authorization: Bearer <token>.
| Method | Route | Notes |
|---|---|---|
POST |
/api/auth/register |
Returns a token; signs the new user straight in |
POST |
/api/auth/login |
Returns { token, expiresAt, user } |
GET |
/api/auth/me |
Current user |
GET |
/api/auth/available?userName= / ?email= |
Returns only a boolean |
GET |
/api/games |
search, system, genre, own, dumped, played, finished, page, pageSize, sort, dir |
GET |
/api/games/{id} |
|
GET |
/api/games/facets |
Distinct systems and genres, for filter dropdowns |
POST |
/api/games |
|
PUT |
/api/games/{id} |
|
DELETE |
/api/games/{id} |
|
POST |
/api/images |
multipart file; re-encodes to WebP |
GET |
/health |
Anonymous |
Ownership is always taken from the JWT subject, never from the request. A game
belonging to another user returns 404, not 403, so the response does not
confirm that the id exists.
Security notes
Rotate the old database password
The 2018 code committed live MySQL credentials to this repository
(interfaceServices/dbConfig.php, and again in four other files). They are in git
history. That password must be considered compromised and rotated, regardless
of this rewrite. Removing the files does not remove them from history.
The new stack keeps secrets in .env, which is gitignored.
What was fixed in the rewrite
The old backend was ~16,700 lines of PHP, of which ~16,400 were vendored
third-party code — four near-identical copies of php-crud-api plus
class.upload.php. Only ~150 lines were application logic. These problems were
not carried across:
| Old behaviour | Now |
|---|---|
| Two endpoints exposed unauthenticated CRUD over every table | Every data route requires a valid token |
Client chose whose rows to read (filter[]=userId,eq,N) |
Owner comes from the JWT subject, server-side |
| Login hardcoded to a single username | Any registered user can sign in |
crypt() with one global salt, silently truncating passwords to 8 chars |
ASP.NET Core Identity (PBKDF2, per-user salt) |
JWT secret was the literal string "testing", tokens never expired |
Key required from config, 12-hour expiry |
| Token passed in the query string | Authorization: Bearer header |
| Uploads anonymous, path built from the client filename | Authenticated, server-generated name, per-user folder, must decode as an image |
Access-Control-Allow-Origin: * |
Explicit origin allowlist |
Passwords could not be migrated — the old hashes are unrecoverable by design.
Known accepted risk
npm audit reports a moderate advisory in @hono/node-server, reached
transitively through @angular/cli's MCP server feature. It is:
- dev-only —
npm audit --omit=devreports 0 vulnerabilities, and it is not in the browser bundle - a Windows-only path traversal, on a Linux-only toolchain here
npm audit fix --force would downgrade Angular CLI to 21.0.4, a breaking change.
Overriding the dependency means forcing a major bump the MCP SDK does not accept
(^1.19.9). Left as-is deliberately; revisit when Angular CLI updates the SDK.
Notable version facts (as of 2026-08)
- Angular 22.1 is zoneless — there is no
zone.jsin the dependency tree. Component state must be signal-based for change detection to see it. @angular/animationsis deprecated in v22; Material 22 no longer depends on it. There is noprovideAnimations()inapp.config.ts, and a test pins that Material still renders without one.- Unit tests run on Vitest, not Karma/Jasmine.
- Fonts and Material icons are bundled from
node_modules, so the app makes no third-party requests at runtime. - The backend pins two transitive packages (
Microsoft.OpenApi,SQLitePCLRaw.lib.e_sqlite3) to clear high-severity advisories. See the comment inLudosData.Api.csproj. - Image processing uses SkiaSharp, not ImageSharp: ImageSharp v4 requires a paid licence key at build time.