The ScholarAIO Library WebUI is a local, metadata-read-only workspace for browsing records, copying citations, reading PDFs, checking metadata quality, and running ranked retrieval without leaving the browser. On WSL, PDF saves made through the Windows default viewer are the deliberate exception: ScholarAIO validates and synchronizes those PDF bytes back to the canonical library file.
scholaraio gui
By default, ScholarAIO listens on 127.0.0.1:8765 and opens the page in your browser. To choose another local port or avoid opening a browser automatically:
scholaraio gui --port 9000 --no-open
The page loads only assets packaged with ScholarAIO. It does not load remote JavaScript, fonts, analytics, or CDNs.
The main library offers four explicit modes.
| Mode | What it searches | Index required |
|---|---|---|
| Metadata | Title, author, journal/source, DOI, IDs, and directory name across the whole library | None |
| Keyword | FTS5 title, abstract, and conclusion text | Keyword index |
| Semantic | Vector similarity over embedded paper metadata/content | Embeddings and the configured embedding provider |
| Unified | Reciprocal-rank fusion of Keyword and Semantic results | Keyword index; embeddings recommended |
Metadata mode applies server-side filters across the whole library after a short typing delay. Keyword, Semantic, and Unified modes run ranked retrieval on the server when you choose Search or press Enter in the query field. Ranked rows show their rank, retrieval leg, and score. Click a column heading to sort the current result set by that field; run the search again to restore relevance order.
Proceedings child papers currently support Metadata mode only. The mode selector is disabled on the Proceedings tab and the limitation is shown directly in the search status area.
Build or refresh the keyword index after library metadata or full text changes:
scholaraio index --rebuild
Build semantic embeddings after the keyword index exists:
scholaraio embed
Unified search remains honest about degradation. If vectors are unavailable, it labels the result as keyword-only and shows scholaraio embed; if the keyword index is unavailable, it shows scholaraio index --rebuild. Semantic mode never substitutes keyword results while claiming semantic retrieval occurred.
Filters combine with AND semantics. You can compose:
Clear all removes every filter, cancels stale ranked responses, returns to Metadata mode, and restores the default year-descending sort. The status at the bottom of the filter card reports the active-filter count.
Select a row to open its Inspector. The action row uses the selected record’s stable paper ID and canonical metadata.
Choose Copy BibTeX to generate a complete entry on the server with ScholarAIO’s canonical BibTeX formatter and copy it to the clipboard. The browser clipboard API is used first. If it is unavailable or denied, the WebUI uses a temporary local textarea fallback and removes it immediately afterward. A status toast reports success or failure.
Main-library and proceedings child records are both supported. Proceedings entries include their proceedings title as booktitle when available.
Choose Preview PDF or the table’s PDF pill to keep reading inside the WebUI. The existing fullscreen and back-to-records controls remain available.
When ScholarAIO can reach a desktop safely, choose Open in default viewer to launch the PDF in the operating system’s configured PDF application. This supports opening several papers in independent native windows while keeping the WebUI available for searching.
On native Windows, macOS, and Linux, the viewer opens the canonical local PDF directly. On WSL, ScholarAIO instead maintains one stable edit mirror per library PDF under %LOCALAPPDATA%\ScholarAIO\editable-pdfs. The readable filename no longer receives a new random prefix on every launch. The Windows default association is respected, including applications such as Foxit Reader.
When the reader saves an embedded annotation, highlight, comment, form value, or other PDF change, ScholarAIO waits for a stable complete file, validates it, and automatically writes a valid one-sided edit back to the canonical WSL library PDF. Canonical refetches flow in the other direction. If both sides changed to different contents, synchronization pauses with a conflict and keeps both files. Use Review PDF versions to inspect, download and explicitly select the desired version; automatic synchronization then resumes. A conflict does not automatically launch or download either version. Reconciliation uses guarded publication, a recovery backup, and retained reader-save files; conflict recovery archives are not automatically deleted.
The synchronization monitor survives WebUI and machine restarts because mirror mappings and hashes are stored in data/state/pdf-edit-mirror/sync.db. A reader may keep saving after the WebUI stops; the next WebUI startup detects and reconciles that edit. Old random files under the legacy Windows temporary ScholarAIO folder are cleanup-only and are never adopted or written back.
Normal successful synchronization is silent. The Inspector shows a non-modal message only while synchronization is pending or has failed. A malformed, partial, locked, or unvalidated mirror is never launched or copied over a valid canonical PDF; the existing browser-download fallback remains available.
Automatic persistence covers changes written into the PDF or saved by atomically replacing that PDF. It cannot recover annotations stored only in a reader-specific cloud account, application database, or unrelated sidecar while the PDF itself remains unchanged.
When the WebUI is bound to a non-loopback host, or the server has no compatible desktop launcher, the action automatically becomes Download PDF. The browser receives an attachment and opens or saves it on the client machine according to its own settings. This is the correct behavior for a WebUI hosted on another computer: a server cannot directly launch an application on the browser’s computer.
The action is intentionally restricted:
If an advertised native launch fails at runtime, or a safe current WSL mirror cannot be established, the WebUI automatically starts the browser download and reports the fallback. The action is disabled only when the selected record has no PDF. Opening or downloading a PDF does not edit bibliographic metadata.
The record list continues to refresh from the library. Expensive semantic or unified queries are not automatically rerun on every poll. If filters change, the search status asks you to run Search again. Stale detail, refresh, and ranked-search responses are ignored so an older request cannot overwrite the current source or query.
Unchanged rows and detail text retain their existing browser elements. Automatic list/detail refresh pauses while you drag or retain a text selection, read an inline PDF, or leave the tab in the background, so you can select and copy titles, authors, and abstracts without interruptions. Clear the selection to resume automatic updates, or use Refresh for an explicit update. Slow automatic refresh requests do not overlap.
| Message or state | What to do |
|---|---|
Keyword search index is unavailable |
Run scholaraio index --rebuild. |
Semantic search index or embedding provider is unavailable |
Configure an embedding provider if needed, then run scholaraio embed. |
Unified search is degraded |
Results are still valid for the available retrieval leg; follow the displayed rebuild command to restore both legs. |
| The action says Download PDF | This is expected for remote/non-loopback deployments or hosts without a desktop launcher; the file is delivered to the browser computer. |
| Native viewer launch fails on WSL | Confirm Windows interoperability provides powershell.exe and wslpath, and that Windows has a default PDF application. The WebUI falls back to a browser download if launch still fails. |
| PDF synchronization remains pending | Finish or close the reader’s Save operation, then check free disk space and permissions for both the library and %LOCALAPPDATA%\ScholarAIO\editable-pdfs. ScholarAIO retries automatically. |
| Native viewer launch fails on Linux | Install/configure xdg-open and a default PDF application in the desktop session. The WebUI falls back to a browser download if launch still fails. |
| BibTeX copy fails | Allow clipboard access or use a browser that supports the local textarea copy fallback. |
| No rows remain after filtering | Choose Clear all, then add filters one at a time. |
For CLI-level search details, see Search & Browse. For the full command surface, see the CLI Reference.
Metadata quality checks run in the background. The list can appear before the first audit finishes; the update indicator shows that checks are running, and any previous audit results remain available during refresh. PDF lookup and citation copying do not wait for a full-library audit.
List requests use ETags so unchanged results return without the full payload. PDF delivery supports single byte ranges and validation with ETags, including after edits. Large first-time Windows opens still need a validated mirror copy; the UI reports preparation while this happens. Server-Timing on successful native-open requests separates lookup, mirror preparation, and launcher time. The settle/lock budget does not impose a hard deadline on file copying, and launcher completion does not mean the external viewer has finished rendering.
Use Review PDF versions (described below) to resolve conflicts with version checks and retained copies. Close PDF readers before confirming a selection.
The table loads 100 records at a time. Previous/Next move between pages; metadata filters and sorting apply to the whole library. Type and volume choices also come from the whole library. A changed library revision resets an obsolete page to the first page. The Refresh button forces a new filesystem scan; automatic scans normally start about five seconds apart. Slow disks can extend this delay. Text selection, inline reading and conflict review suspend background DOM refresh.
Open in default viewer remains the normal editable-PDF action. Read in
browser uses the streaming PDF route without preparing a Windows edit mirror.
Browser reading does not save changes into the library; download a copy if needed.
The first editable open of a large PDF still needs a safe copy. The interface
shows preparation status, and server logs/Server-Timing distinguish lookup,
preparation and viewer-launch dispatch. Dispatch timing does not measure when
Windows finishes rendering the document.
When synchronization reports a conflict, choose Review PDF versions. Preview or download the canonical, mirror and retained reader-save versions. Download both if you want to keep independent copies. Damaged/partial versions can be downloaded for manual recovery, but cannot be previewed or selected. To synchronize a chosen version:
The server checks that the inspected versions have not changed. A concurrent
save rejects the old decision; inspect again. All valid versions are archived
under data/state/pdf-edit-mirror/resolutions/ (or the configured state root),
and displaced files remain beside their original PDF in
.scholaraio-pdf-recovery/. Nothing is automatically deleted by conflict
resolution. A late save through an old file handle reopens the conflict.
Agents can use the same recovery service:
python -m scholaraio.cli pdf-recovery inspect PAPER_ID
python -m scholaraio.cli pdf-recovery export PAPER_ID --token TOKEN --version canonical --output copy.pdf
python -m scholaraio.cli pdf-recovery resolve PAPER_ID --token TOKEN --version mirror --readers-closed
TOKEN and version IDs come from inspect. Add --source proceedings for a
proceedings paper. Export refuses to overwrite an existing output file. Keep
recovery files until you have verified the intended annotations and closed every
reader; storage cleanup remains an explicit manual action.
See Library projection contracts for index freshness, pagination and performance budgets.