scholaraio

Agent Reference

This document is the deeper reference for agents and maintainers. The root entry docs such as AGENTS.md, CLAUDE.md, and .qwen/QWEN.md are intentionally kept lighter and should stay focused on durable project facts, hard constraints, and navigation. For the full repository knowledge map, start at docs/DESIGN.md.

Instruction Layering

Use the instruction stack in this order:

  1. Root wrapper for your host tool
    • CLAUDE.md
    • AGENTS.md
    • .qwen/QWEN.md
    • .cursor/rules/scholaraio.mdc
    • .clinerules
    • .windsurfrules
    • .github/copilot-instructions.md
  2. Matching project skill under .claude/skills/<name>/SKILL.md
  3. Focused reference docs such as CLI, setup, writing, or migration specs
  4. Source code and tests

Practical rule:

Repository Knowledge System

ScholarAIO treats repository-local Markdown as the system of record for agent context. AGENTS.md is the map injected early in agent sessions; it should not become the encyclopedia.

Key indexes:

Internal plans, validation records, and audits are intentionally excluded from the published documentation site.

How Skills Are Organized

The canonical project skill source is:

Cross-agent discovery wrappers expose the same skill set through:

For reuse from another project, prefer the automated registration command:

scholaraio setup agent
scholaraio setup agent --apply
scholaraio setup agent check

It previews and applies shell runtime wiring, Codex/OpenClaw global skill discovery, project-local wrappers for supported hosts, and Claude Code plugin instructions where automation is not possible.

Project-local wrappers are local machine integration blocks. They may contain absolute paths to the active ScholarAIO checkout and config, so review them before committing target-project files.

Project guidance for maintaining skills:

Capability-Based Routing

Shared skills route by the task’s required capability and output contract, not by the host or Agent brand:

Route Use it when
Current-session native capability The capability is actually exposed in this session and can complete the one-off reading, reasoning, writing, browsing, or visual task
ScholarAIO core CLI The task needs library access, provenance, persistent notes, reproducible IR, deterministic Office files, or another tested project contract
Optional sidecar or external extension The user explicitly requests it, or the native/core route cannot meet a specialized rendering or benchmark contract

Never infer tool availability from an Agent name. Check the capabilities that are actually available, select the smallest route that satisfies the output, and state any verification boundary. This keeps the canonical skills portable while still allowing host-specific setup commands in dedicated integration documentation.

Representative skills:

Repo And Module Map

ScholarAIO’s canonical implementation namespaces are:

High-signal mental model:

The breaking cleanup generation removed legacy public facades such as scholaraio.index, scholaraio.workspace, scholaraio.translate, and scholaraio.ingest.pipeline.

Current import rules:

Current Runtime Layout

Fresh-layout runtime:

Breaking cleanup behavior:

Workspace rules:

Migration rules:

Agent Operating Model

ScholarAIO is meant to be used through an agent, not only through direct shell scripting.

Agents should:

Notes And Cross-Session Analysis

When analysis should persist across sessions, use paper-level notes.md.

Conventions:

Useful mental model:

Use the smallest doc that answers the question:

The maintenance rule for this repo is simple:

Concurrent record updates

Use scholaraio.stores.papers.update_meta for field changes, or modify_meta for short in-memory edits such as updating a nested translation entry. These helpers lock a persistent sidecar for the complete read–modify–write and atomically replace the JSON through a unique temporary file. Do network calls and extraction before entering the transaction. write_meta is a whole-record replacement, not a merge of a previously read snapshot. Workspace reference create/add/remove/show-refresh and workspace rename share one collection lock. Add/remove/show require an existing workspace; use create explicitly to initialize one, and never recreate a renamed-away path from a waiting write. Hidden .scholaraio-<hash>.lock files live at the paper-library root for metadata, and beside the shared workspace collection directory for references. Keeping the handle outside the directory allows Windows to rename that directory while holding the lock. Metadata transactions in one library root share a single lock, including directory rename and registry commit. This deliberately serializes short metadata writes across papers so a rename cannot change lock identity; network calls and extraction must remain outside it. These are persistent coordination files; do not delete them while writers are active. Restart all writer processes when upgrading from the older per-directory metadata/reference lock protocol.

Library rename/refetch/repair callers pass cfg.index_db explicitly so fresh and custom runtime layouts update registered directory names and search paths together. An index update failure rolls the directory name back and surfaces the error. Standalone file tools may omit the database, but must not guess an old runtime path from a file’s parent.