ctrl+f for your screenshots folder
gyotaku watches your screenshots folder, reads every word inside every image with OCR, and makes it all searchable in milliseconds — fully offline, open source, native Linux.
For this job: yes. PaddleOCR's PP-OCRv5 is the best accuracy-per-megabyte OCR you can run on a CPU in 2026 — roughly 4.5% character error rate against ~18% for Tesseract on the same benchmark, a 7.5 MB English recognizer, 109 languages, Apache 2.0, no GPU, no cloud. It dethroned Tesseract as the most-starred OCR project on GitHub in March 2026. For keyword search over screenshots, nothing else is close on the accuracy/size/speed triangle.
Scroll for the full pipeline walkthrough, a simulated live demo, and the 2026 OCR showdown table.
You know exactly which screenshot has the thing. You have no idea which file it is.
The author's words: "i screenshot everything and can never find anything." Receipts, error messages, wifi passwords, that one tweet, the terminal output you swore you'd remember — they pile up in ~/Pictures/Screenshots as shot_4821.png, shot_4822.png, shot_4823.png. Your OS can search filenames. It cannot search inside the pixels.
gyotaku fixes that with the oldest trick in computing: build an index. Every screenshot gets read once by OCR; every word goes into a searchable database. After that, finding anything is a database lookup — milliseconds.
Walk the pipeline one step at a time. The screenshots below are simulated drawings, not real OCR output.
Eight fake screenshots, drawn live on canvas. Type any word that appears in one — the library filters instantly, exactly like gyotaku's search box. All content below is simulated.
Seven engines, one table. Numbers are directional — drawn from 2026 benchmarks and project docs, linked below — not lab-grade measurements.
| Engine | Accuracy | Speed (CPU) | Model size | Languages | Fully offline | License |
|---|---|---|---|---|---|---|
| PaddleOCR PP-OCRv5 ★ gyotaku's pick |
~4.5% char error rate; 92.86 OmniDocBench; matches GPT-4o on some OCR tasks at 5M params | ~1.5–2.5 s per A4 page | 7.5 MB recognizer (EN) + 84 MB detector | 109 | Yes | Apache 2.0 |
| Tesseract 5.x | ~18% char error rate on same benchmark; fine on clean print | ~2–5 s per page | ~30 MB | 100+ | Yes | Apache 2.0 |
| EasyOCR | Good on print; weak on the hard stuff (0.26 edit distance vs 0.07 PaddleOCR on OmniDocBench EN) | 3–5× slower than PaddleOCR | ~2 GB install (drags PyTorch) | 80+ | Yes | Apache 2.0 |
| RapidOCR (ONNX) | Same as PaddleOCR — it runs PP-OCR weights | Fast on CPU | ~40–150 MB total, no torch | Same as PaddleOCR | Yes | Apache 2.0 |
| Surya 2 | Best-in-class layout + tables; strong multilingual | Slow per page; wants a GPU | 0.65B params (~1.3 GB+) | 90+ | Yes, but needs GPU | OpenRAIL (not fully OSS) |
| Apple Vision | Excellent on Latin scripts | Real-time on-device | Built into the OS | Broad | Yes | Proprietary, Apple-only |
| Google ML Kit | Very good, mobile-tuned | Real-time on-device | ~10s of MB | Broad | Yes | Proprietary, mobile-only |
Sources: character-error figures from a 2026 community benchmark (github.com/nickrivers1983/opencode-vision); PP-OCRv5 specs, 60k+ stars and GPT-4o-matching claim via abit.ee, Mar 2026; model sizes via huggingface.co/monkt/paddleocr-onnx; Tesseract standing via koncile.ai; Surya/OpenRAIL and VLM rankings via 2026 community research docs; EasyOCR weight from multiple migration notes (2 GB → 150 MB switching to RapidOCR).
There is no "best OCR". There is best OCR for a constraint.
ort crate — the RapidOCR playbook, not pip install paddleocr.The index lives in a single file. FTS5 full-text search turns "find the word" into a B-tree lookup — that's the milliseconds.
Zed editor's Rust GPU UI framework. Native widgets, GPU-composited, no Electron — search-as-you-type stays at 60fps.
No server, no API key, no per-page pricing. Your screenshots never leave the disk — the privacy story is the feature.