LIBRARY ORGANIZER - DOCKER / NAS EDITION
=========================================

QUICK START: read INSTALL.md - it is the full, battle-tested install guide
covering both Portainer and SSH, the NAS path rules, verification steps,
and a troubleshooting table of real errors with their fixes.

Compose files in this package:
  docker-compose.yml        normal install (Portainer or SSH; no build step)
  docker-compose.host.yml   host-networking variant - use if the page is
                            unreachable but curl localhost:8765 works
  docker-compose.build.yml  builds a self-contained image (SSH only)

USING IT
--------
Inside the container your library is always at /source and output at /dest
(those are pre-filled in the UI - normally you don't change them).

1. Scan            - reads tags, folder structure, and filenames
2. Review          - orange rows need attention; click any Author / Series /
                     # / Title cell to edit in place (saves on click-away).
                     Filter box + "needs review" filter help with big scans.
                     "Lookup all needs-review online" polishes names via
                     Google Books / Open Library.
3. Copy            - copies everything not yet copied into /dest.
                     Re-running skips books already copied.

Then point Audiobookshelf or Calibre at the output folder.

WHAT'S NEW IN v1.9 - TUNED AGAINST A REAL 1,200-BOOK LIBRARY
-----------------------------------------------------------
The scanner was dry-run against a complete real Audiobookshelf library
listing (1,209 books, 41,000 files) with the folder names as the only
information. Before tuning it produced 15,627 "books" and 10,587 false
duplicates; after tuning: 1,192 books, 1,159 exact matches, 0 books missed,
and every merge a genuine disc/part folder. What changed:
- Numbered chapter files are ONE book: "Title - 01 - Opening Credits.mp3",
  "Title - 02 - Chapter 1.mp3", "03 - Beauty.mp3"... are grouped by the
  text before the track number; chapter names after it are ignored.
- Series index vs part number: "Lee Child - Reacher 27 - Personal 12"
  (27 = series position, 12 = part), "Jack Reacher Series, Book 18",
  "MR 23 Capture or Kill" and "The Expanse 2.5" are all read correctly.
- More layouts: "(6of17)" and "D01"/"D12" disc folders, "The Last Juror
  01..10" sibling folders, "[The Expanse 1.0] Leviathan Wakes" and
  "[Series 01] - Title" bracket forms, "Author Trilogy Series 1-Title",
  "Series by Author" top folders, numbered prefix codes ("LNSR01-95",
  "01-10 DFOA"), "01 to 14" and "3 of 12" units, dotted names
  ("Long.Shadows.[Unabridged].-.006").
- A numbered folder ("01/") that merely CONTAINS a book folder is never
  treated as a disc of its parent.
- Cleanup: Amazon ASINs and bare ISBNs glued to folder names, "(10 MP3s -
  U)", "(4 Discs - A)", "Unb"/"Unabridged", "A Novel", "Collection-",
  author-name prefixes ("Larry Niven-Footfall"), double extensions
  ("Lasher.doc"), "word- word" colons, and subtitle echoes of the series
  ("If I Had a Nickel: Roy Ballard Mysteries, Book 3").
- The folder name arbitrates when a filename hint has title and series
  swapped ("Transfer of Power - Rapp 01").
Remember: on the NAS the sidecar metadata.opf/.json files win over all of
this - the tuning covers the books that have no sidecar.

WHAT'S NEW IN v1.10 - ONE APP, TWO LAUNCHERS
--------------------------------------------
The Windows edition is now the same web app as this Docker edition, started
by a small launcher (LibraryOrganizer.pyw) that serves it on 127.0.0.1 and
opens your browser. The old Tkinter desktop window is retired: it had only
the engine improvements but none of the interface features (layouts,
duplicate policies, skip, bulk edit, filters, Stop, persistence...), and it
drifted. Now both editions are byte-identical app files - a fix in one is
a fix in both. Windows specifics: the folder picker starts at "My Computer"
(drive letters, mapped shares), the plan auto-saves under
%LOCALAPPDATA%\LibraryOrganizer, and the server binds to localhost only.

WHAT'S NEW IN v1.9.1 - BOX SETS, BAD metadata.json
--------------------------------------------------
- Box sets: a number range on the folder ("Box Set (5-8)", "Books 17-20")
  is carried into a generic sidecar/tag title, so ten box sets no longer
  collapse into one title. Ranges and volume markers inside (...) are
  content and survive even "Strip all (...)". If a series has no position
  yet, the range becomes it ("5-8") - Audiobookshelf sorts that correctly
  among the single books.
- Duplicate detection is series-aware: the same title at different series
  positions (box sets 02 and 03) is NOT a duplicate.
- Rows whose sidecar disagrees with the folder path are marked
  "(path differs)" - e.g. a metadata.json claiming "Addison Cain" inside
  "Wilbur Smith/The Dark of the Sun".
- "Use only as last resort" scan options (v1.10): tick any of .opf,
  metadata.json, or embedded tags to demote that source behind folders and
  filenames - it then only fills what nothing else could, and the Meta
  column says e.g. "tags (last resort)". Use these for whichever source a
  bad import polluted in YOUR library (json from an early Audiobookshelf
  scan, tags from a sloppy ripper, opf from a mismatched Calibre import).
  "Trust folder names over embedded tags" still decides the folders-vs-tags
  order when tags are not demoted.

WHAT'S NEW IN v1.9 - METADATA SYNC (fix bad embedded tags for good)
--------------------------------------------------------------------
The reviewed table is now the single source of truth, and at copy time it
is written EVERYWHERE, so all three metadata layers agree in the output:
  - metadata.opf   (Calibre)          - title, author, series, narrator
  - metadata.json  (Audiobookshelf)   - title, authors, narrators, series #N
  - embedded tags  in the COPIED files - originals are never touched:
      MP3   album/artist/albumartist = title/author, composer = narrator,
            SERIES + SERIES-PART (and MVNM/MVIN), track numbers 1/N..N/N in
            playback order, junk "grouping" tag cleared
      M4B/M4A  same via iTunes atoms (movement name/index, SERIES atoms)
      FLAC/OGG/OPUS  vorbis comments
      EPUB  the OPF inside the epub is rewritten (title, author, series)
Toggles next to Copy: "Write metadata.opf", "Write metadata.json",
"Fix embedded tags in copies" - all on by default.
Reading is bidirectional too: a rescan of the OUTPUT folder (with "Trust
folder names" unticked) reproduces the reviewed data exactly from the tags.

Also new: "Trust folder names over embedded tags" (on by default). When a
structured Author\Series\Book path and the embedded tags disagree, the
folders win and the row's Meta column shows "(tags differ)" so you can eye
it. Untick to trust the tags instead. Sidecar files always win over both.

The Windows edition writes all three layers as well.

WHAT'S NEW IN v1.8 - PARTS WITH ROMAN NUMERALS, TITLE CLEANUP, NARRATOR
-----------------------------------------------------------------------
- Part folders like "1_ Part I - The World of Jeremy Walker", "7_ Part VII",
  "Disc 1 of 5", or plain "01"/"02" are recognized as parts of the parent
  book (roman numerals, numeric prefixes and subtitles all handled), kept
  in order I, II ... VII.
- Leading years on book folders ("1986 - Belinda") are dropped.
- Title/series cleanup: "(read by X)", "(Unabridged)", format, bitrate and
  year tags are removed automatically; an author-fragment tag such as
  "(Jyr)" on "Star Force Universe (Jyr)" is removed when it matches the
  author. Anything else in (...) is kept - unless you tick
  "Strip all (...) from titles" before scanning.
- Narrators found in "(read by ...)" are captured (see the CSV export) and
  written into metadata.opf as a contributor, so Audiobookshelf/Calibre
  get the narrator without it polluting the title.
- Colons in titles become " - " in folder names ("Star Force - Origin"),
  instead of a jammed "Star Force- Origin".
- Tag hygiene: the ID3 "content group"/vorbis "grouping" fields are no
  longer treated as a series (they usually hold chapter/part names). A
  series identical to the title is discarded.
- Windows edition: fixed a latent crash on Copy and brought its copy
  engine to full parity (metadata.opf, layouts, part renaming).

WHAT'S NEW IN v1.7 - DISC / PART AWARENESS
------------------------------------------
Multi-disc and multi-part books are now recognized as ONE book instead of
several (which previously showed up as false "duplicates"):
  - Nested disc folders:   Title\Disc 1\, Title\CD 2\, Title\Part 3\ ...
  - Sibling disc folders:  "Title - Disc 1", "Title - Disc 2" side by side
  - Numbered files sharing one folder: "Dune Part 1.mp3", "Dune Part 2.mp3"
    next to "Martian CD1.mp3", "Martian CD2.mp3" -> two books, not one
Recognized markers: disc, disk, cd, part, pt, side, tape (with or without a
number separator, "3 of 12" forms too), plus numbered prefix codes such as
"LNSR01-95 Larry Niven - Saturn's Race.mp3" (v1.7.1). When the file names
carry a clean "Author - Title", that is used in preference to a sloppy
folder name like "Larry Niven-Saturns Race". "Book 1" / "Vol 2" are NOT disc
markers - they stay series positions. Titles containing numbers
("Fahrenheit 451", "1984") are left intact.

Parts are kept in playback order (disc 1 first, natural sort so disc 10
follows disc 9). When a book spans several disc folders, copied parts are
prefixed with the disc folder ("Disc 1 - 01.mp3") so nothing clashes and
everything sorts right; with "Rename multi-part files" on, they become
"Title - 01.mp3" ... numbered straight through.

WHAT'S NEW IN v1.6 - DUPLICATE POLICIES
---------------------------------------
How duplicates work: a "duplicate" is a book whose author + title + type
(audio/ebook) matches another book found elsewhere in the scan. They are
highlighted purple, counted at the top, and have their own filter.

Resolve them with the "Duplicates" dropdown + "Apply to duplicates":
  decide manually     - the default: nothing automatic; skip by hand.
  keep best format    - m4b > m4a > flac > opus > ogg > mp3 (audio);
                        epub > azw3 > azw > mobi > pdf (ebooks).
  keep largest        - most total bytes (usually the best-quality rip).
  keep newest files   - most recently modified files win.
  keep oldest files   - oldest files win.
  keep fewest files   - single-file over multi-part (ties -> best format).
  keep ALL            - nothing is skipped; at copy time the extra copies
                        get numbered folders: "Title", "Title (2)", ...

Applying a policy keeps ONE book per duplicate group and marks the rest
"skipped" (grey). Nothing is deleted or copied by this step - review the
"skipped" filter, un-skip anything you disagree with, then Copy. Applying
a different policy later re-resolves every group from scratch.

Also in v1.6: single-letter or alphabetical-bin folders ("A", "B", "M-Z")
are no longer mistaken for author names.

WHAT'S NEW IN v1.5
------------------
- Sidecar metadata files are now the TOP source: a metadata.json
  (Audiobookshelf) or metadata.opf / any .opf file (Calibre) sitting in the
  book's folder beats everything else - so curated metadata wins even when
  folder and file names are garbage. The Meta column shows "opf" or
  "metadata.json" when this happened.
- Multi-file tag fix: for multi-part books the per-track TITLE tag is a
  chapter name; only the ALBUM tag is trusted as the book title now (no
  more books called "Chapter 1").
- New "Rename multi-part files" option (off by default): parts of
  multi-file audiobooks are renamed to "Title - 01", "Title - 02"...,
  numbered in their original sort order.
- A book's own sidecar .opf is preserved as-is in the output; the tool only
  writes its own metadata.opf when the book didn't bring one.

WHAT'S NEW IN v1.4 (review pass)
--------------------------------
- Big-library performance fix: the status endpoint was recomputing the
  duplicate map per book on every poll (O(n^2)); with 100k books the UI
  would have crawled. Now computed once per poll.
- Collision guard: before copying, the plan is checked for two different
  books resolving to the same output folder (easy with the flat/Title-only
  structures, or unresolved duplicates). You're warned and can review,
  skip, or force.
- "Last, First" author flipping now handles two-word surnames:
  "Le Guin, Ursula K." becomes "Ursula K. Le Guin" (was truncated before).
- Serves through Waitress, a production WSGI server (no more dev-server
  warning; better under concurrent use).
- Diagnostics now also print to `docker logs`, not just the web UI's log
  pane - the startup mount report is visible from the terminal.

NEWER FEATURES
--------------
- Your plan auto-saves to /dest/.organizer-plan.json - a container restart
  never loses review work. Delete that file for a fresh start.
- Check rows (checkbox column) to bulk-set author / series / # across many
  books at once - fixing a whole author's worth of rows takes seconds.
- "Export CSV" downloads the whole plan as a spreadsheet for offline review.
- Audiobook online lookups now hit Audible's catalog first, which returns
  real series names AND positions (Google Books rarely knows series).
- "Write metadata.opf" (on by default) drops a small metadata file into each
  book folder so Calibre and Audiobookshelf pick up author/series/title
  directly, even if a folder is ever renamed.
- Copy pre-checks free space on the destination and warns you before
  starting a copy that can't finish.
- Possible duplicates (same author + title + type found in two places) are
  highlighted purple and counted at the top - use the "possible duplicates"
  filter to review them, then Skip the format you don't want.
- Skip (the small circle button on each row, or "Skip checked" in the bulk
  bar) parks a book: it stays in the plan, greyed out, and is never copied
  until you un-skip it.
- Every copied file is size-verified against the original; any mismatch is
  flagged as an ERROR row instead of silently leaving a truncated file.
- Long scans and copies show a time-remaining estimate above the table.
- Output structure is your choice - a Structure dropdown offers:
    Author\Series\01 - Title   (default; best for Audiobookshelf)
    Author\Title
    Author - Series 01 - Title  (flat, one folder per book)
    Author - Title              (flat)
    Title only                  (flat)
  The "Will be filed under" column previews your chosen layout live.
- A Stop button appears during any scan, lookup, or copy. Stopping is safe:
  it finishes the current book cleanly, and a stopped copy resumes right
  where it left off on the next run.
- Browse... buttons next to Source and Destination open a folder picker, so
  you can scan just one subfolder (e.g. only /source/NewStuff) instead of
  the whole library. To reach other shares on the NAS, add more volume
  lines to the compose file (e.g. - /volume2/OtherShare:/extra:ro) and
  they'll appear in the picker.
- Smarter series detection from folders: top-level series folders like
  "Dresden Files", "Stormlight Archive" or "Discworld Saga" are now
  recognized as series (not mistaken for authors), and numbered book
  folders like "Book 2 - Title" fill in the series position.

NOTES
-----
- No login/auth: it's meant for your LAN only. Don't port-forward it.
- Big libraries: scanning ~100k books takes a while - the log shows progress
  every 100 books; the table is paginated and filterable so it stays fast.
- Online lookup needs the container to have internet access, and on a very
  large library "lookup during scan" will be slow - better to scan first,
  filter to "needs review", and lookup just those.
- If Copy fails with a permission error, the /dest mount is read-only or
  owned by another user - check the folder's permissions on the NAS.
- Stop/remove:  docker compose down     (your files are untouched; only the
  in-memory scan plan is lost)
