Files
Files are documents attached to references, typically PDFs of the referenced work.
Each reference can have one primary file and any number of supplementary files. The primary field distinguishes the
main PDF (e.g., the article itself) from supplementary materials, and is the only file field you can change after
upload — set primary: true with PATCH /files/{file_id}. The filename is fixed at upload time.
Upload a file by POSTing a multipart/form-data body (with the binary in the file part) to
/references/{reference_id}/files. Files always belong to a reference, so there is no non-scoped upload endpoint;
POST /files is not available.
File lists are not paginated: GET /files (and its reference- and library-scoped variants) returns every matching
file in a single response, and the cursor field is always null. Filter the global list with reference to get
the files for one reference, or use the dedicated GET /references/{reference_id}/files alias.
To read normalized searchable PDF text, use GET /files/{file_id}/text_content with page, or
page_start and page_end for an inclusive span. Optional start/end are page-local UTF-8 byte
bounds, inclusive/exclusive. For spans they affect only the first/last page. start must be zero or
more; a negative end counts from that page's end, and omitting end includes its final character.
Omit selectors and bounds for all pages.
Fulltext search results' fileId, pageNumber, and charOffsets map directly to the file, page,
and byte bounds. These offsets enclose matched tokens, not necessarily the entire displayed snippet.
Widen with start: Math.max(0, charOffsets[0] - context) and end: charOffsets[1] + context.
A negative start is rejected. For annotation hits, request their whole page.
| Limit | Value | Behaviour |
|---|---|---|
max_bytes |
default 20,000; maximum 100,000 | Shared text budget, spent in page order; invalid budgets rejected |
| Page numbers | Positive safe integers | Exact pages and span starts must exist; span ends clamp to the last available page at or before page_end |
| Output entries | At most 100 | Independent of the byte budget |
start / end |
Safe integers; start zero or more |
Resolve a negative end from the page end, clamp, then align to UTF-8 characters |
Responses contain file_id, offset_unit: "utf8_bytes", source, total_pages, and a content array
with one entry per page. Each entry includes range_index, page, actual start/end, page_bytes,
text, and truncated. Resume a truncated page from its returned end. Blank pages are retained.
Extraction covers at most 100 pages; total_pages is the available extraction, not the PDF's full length.
The reader uses stored extraction and the same normalization as search. Missing extraction returns
404 unless the host has an explicitly configured compatible Poppler fallback. Normalization can alter
accents and ligatures; raw exported text and PDF text need not have matching coordinates. For durable
locations, cite the page or highlight the passage with POST /files/{file_id}/highlight.
To work in a shared library, or to address a file through its reference, see Resource and endpoint conventions.
See the File object for the full schema.