Loading page…
Loading page…
Extract text and line locations from PNG and JPEG images with a browser worker. This guide describes the published alpha API. Local inference runs in browsers; Node.js supports SSR imports and the explicit remote client.
@diakrio/ocrThis scoped package has no verified public npm release as of 2026-09-18. You can try the locally built runtime on this site while the release is prepared.
Free for personal and commercial application use. License and notices.
The worker downloads developer-configured runtime and model assets, then decodes and processes the image locally. Asset downloads are separate from document uploads. The OCR local mode does not transmit image data to Diakrio. A separate authenticated server integration can upload images for CPU processing; the public tool currently runs locally.
The preview uses a PP-OCRv6 detector and recognizer with a pinned model dictionary. A quality preset selects the model files: Fast uses Tiny models, Balanced uses a Tiny detector and Small recognizer, and High accuracy uses Medium models. Downloads are about 6, 23 and 139 MB respectively, plus a 3.6 MB runtime. A model’s dictionary does not establish accuracy for every language it contains.
Balanced is the browser default. High accuracy is experimental in browsers: our desktop Chromium sample took about a minute per page. Model files are verified and cached on your device; use Clear downloaded models to remove them. Images and OCR results are not cached.
The alpha wrapper creates an instance, recognizes one image at a time, and disposes its worker. Overlapping calls are rejected. Aborting inference or reaching its deadline terminates the worker; create a new instance to continue. Malformed images return an error.
JPEG orientation is applied and transparency is composited on white. Results contain text, lines, locations, scores, and stage timings. Scores are model confidence values, not calibrated correctness probabilities.
The OCR tool downloads the selected models only when you run recognition. Results include corrected orientation and unresolved regions for review. The published alpha has Chromium, Firefox, and WebKit consumer checks. This does not establish mobile performance or accuracy for every language.