The library carried titles, systems and genres but almost nothing else:
developer was 3.8% filled, publisher 2.9%, description 0%. These are columns
the app has always had and never been able to populate.
enrich_metadata.py reads them off the same articles the cover fetcher
locates. Only empty fields are touched unless --overwrite is given.
year 79% -> 96%
developer 3.8% -> 95%
publisher 2.9% -> 96%
description 0% -> 96%
Parsing infoboxes needed several guards, each found by checking output
rather than trusting the first pass:
* "Infobox video game" is a substring of "Infobox video game series", so
the loose test resolved Banjo-Kazooie to the series overview. Now
rejected, which also fixes the cover fetcher's article resolution.
* An article spans every release and its date block leads with the
original, so year is only filled when the article covers that platform.
Otherwise a DS port inherits the SNES original's year.
* A search hit that neither covers the platform nor closely matches the
title is discarded: "Dragon Ball Z Budokai" surfaces "Shin Budokai", a
different game on a different console. Left blank instead.
* Values are grouped under bold platform headings, tagged with region
codes, annotated with the platform in parentheses, and wrapped in
templates whose named parameters leak through. Each of those read as
the developer or publisher before being handled.
Developer, publisher and year are facts and written verbatim. Descriptions
are article summaries under CC BY-SA, stored with an attribution line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The API had no tests. These run against the real application through
WebApplicationFactory — same pipeline, same Identity configuration, same JWT
validation — with only the SQLite file, upload folder and signing key
swapped, so a passing test says something about what ships.
The ownership tests pin the rule the old PHP API got wrong: a second user
sees an empty library, gets 404 (not 403, which would confirm the id exists)
when reading, updating or deleting someone else's game, and cannot reassign
ownership by putting ownerId or userId in the request body.
Also covered: the password policy, that login is indistinguishable between a
wrong password and an absent user, that the availability endpoint leaks no
row data, that a token signed with an untrusted key is refused, that the
sort parameter is allow-listed rather than interpolated, and that uploads
must decode as an image regardless of extension or content type.
36 tests, ~1s. Program is now declared public partial so the test host can
reach it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
libretro-thumbnails stops at the retro consoles, leaving the ten Xbox 360
titles blank. English Wikipedia carries a cover on essentially every
notable game article and needs no account, so it now runs as a second pass
for anything libretro cannot match.
Correcting an earlier claim in this repo: libretro does publish a
"Microsoft - Xbox 360" set. I had reported it as absent after checking only
my own hardcoded system map, not the actual catalogue of 123 sets. The set
turns out to hold about a dozen entries, none of them ours, so the
conclusion held but the reason given was wrong. It is now mapped and
searched anyway, in case it fills out later.
The cover filename is read from the article's infobox rather than inferred
from file names: filtering names for "box" also matches
"Xbox-360-Pro-wController.png". Two details that cost a round each:
* the infobox writes the field both bare ("Halo 3 final boxshot.JPG") and
prefixed ("File:Lost-Planet-New.jpg"), so any prefix is stripped before
exactly one is added back
* the API returns 429 under an unthrottled loop, so calls are spaced one
second apart, retried with a longer backoff, and cached to disk
Coverage is now 105/105 — 93 from libretro, 12 from Wikipedia.
Two rows took art of the right game but the wrong platform, because their
system field looks wrong in the source data: a Game Boy "Donkey Kong
Country 2" and a DS "Donkey Kong Country Returns". Noted in the README
rather than silently corrected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tools/cover-art/fetch_art.py matches each game against libretro-thumbnails
by title + system and attaches the result through the app's own
POST /api/images, so fetched art goes through the same validation and WebP
re-encoding as a manual upload. Standard library only.
Matching bridges a personal catalogue and a ROM-naming one:
* accents stripped, so "Pokemon Yellow" reaches "Pokémon"
* roman numerals folded to digits, so the SNES "Final Fantasy 2" lands on
"Final Fantasy II" and the PS1 "Final Fantasy V" on its own entry
* trailing articles unwound ("Sims 2, The" -> "The Sims 2")
* subtitle containment in both directions, since our rows sometimes omit
what the catalogue carries ("Wave Race 64" vs "... - Kawasaki Jet Ski")
and sometimes carry what it omits ("Donkey Kong Country 2: Diddy's Kong
Quest" vs the GBA set's "Donkey Kong Country 2")
* a sequel guard, so containment cannot collapse "Donkey Kong Country 2"
onto "Donkey Kong Country"
* fuzzy enough to absorb typos: "Brett Hull Hocky 95" finds "Hockey 95"
93 of 105 games now have art. The remainder: 10 Xbox 360 titles, which
libretro has no thumbnail set for, and two rows whose platform looks wrong
in the source data (a Game Boy "Donkey Kong Country 2", which was never
released on that system, and a DS "Donkey Kong Country Returns", which was
Wii and later 3DS).
Real art also invalidated a layout assumption: the grid used object-fit:
cover, which was fine for uniform placeholders but crops actual boxes, whose
aspect ratios run from near-square SNES to tall N64. Switched the grid and
the editor preview to object-fit: contain so the whole cover is visible.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Angular's production build defers the main stylesheet with
`media="print" onload="this.media='all'"` and inlines a critical subset
ahead of it. The nginx CSP sets `script-src 'self'`, which blocks that
inline event handler — so the swap never ran and the stylesheet stayed
print-only. The app rendered from the ~23kB critical subset alone.
Most of the page still looked right, which is what made it easy to miss.
Material icons did not: `.material-icons` was not in the critical subset,
so every icon fell back to the body font and rendered its ligature name
("videogame_asset") clipped to the icon box.
`inlineCritical: false` emits a plain <link rel="stylesheet">. The
stylesheet is 24kB and same-origin, so the optimisation bought little and
cost correctness under a strict CSP.
Verified in headless Chromium: icons render as glyphs across the toolbar,
grid, editor and mobile layouts, and the console is now clean where it
previously logged four CSP violations per page load.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 2018 stack (Angular 5.2 / CLI 1.7, PHP, MySQL) had not been touched since
July 2018. Rebuilt rather than upgraded in place: the frontend was 17 major
versions behind, and of ~16,700 lines of PHP only ~150 were application logic —
the rest was four near-identical vendored copies of php-crud-api plus
class.upload.php.
Backend — ASP.NET Core 10, EF Core, SQLite
* ASP.NET Core Identity (PBKDF2) + JWT bearer auth
* Clean REST API replacing php-crud-api's filter[]/transform query syntax
* Box art uploads re-encoded to WebP via SkiaSharp
* Imports the 105 games recovered from the 2018 dump on first run
Frontend — Angular 22, zoneless, signals, Material 22
* Standalone components, lazy routes, functional guards and interceptor
* Vitest replaces Karma/Jasmine; fonts and icons bundled, no CDN calls
* No provideAnimations: @angular/animations is deprecated in v22 and
Material no longer depends on it (pinned by a test)
Docker
* Multi-stage builds for both services, non-root at runtime
* nginx serves the SPA and reverse-proxies the API, so everything is
same-origin; one volume holds the database, uploads and DP keys
Security issues in the old code, not carried across:
* Two endpoints exposed unauthenticated CRUD over every table
* The client chose whose rows to read (filter[]=userId,eq,N); ownership now
comes from the JWT subject server-side
* Login was hardcoded to a single username
* crypt() with one global salt, silently truncating passwords to 8 chars
* JWT secret was the literal string "testing", tokens never expired
* Token travelled in the query string rather than a header
* Uploads were anonymous with the path built from the client filename
* Access-Control-Allow-Origin: *
The live MySQL password committed in 2018 remains in git history and must be
rotated independently of this change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>