A visual breakdown

shotgrep

ctrl+f for your screenshots folder

00The verdict, first

What it is

gyotaku watches your screenshots folder, reads every word inside every image with OCR, and makes it all searchable in milliseconds — fully offline, open source, native Linux.

Is PaddleOCR the best OCR?

For this job: yes. PaddleOCR's PP-OCRv5 is the best accuracy-per-megabyte OCR you can run on a CPU in 2026 — roughly 4.5% character error rate against ~18% for Tesseract on the same benchmark, a 7.5 MB English recognizer, 109 languages, Apache 2.0, no GPU, no cloud. It dethroned Tesseract as the most-starred OCR project on GitHub in March 2026. For keyword search over screenshots, nothing else is close on the accuracy/size/speed triangle.

One asterisk: PaddleOCR's official runtime is Python/C++ — heavy for a Rust app. The sharp move is running PP-OCRv5's models through ONNX Runtime (the RapidOCR trick): same weights, a fraction of the footprint, no Python. Models: yes. Runtime: go ONNX.

Scroll for the full pipeline walkthrough, a simulated live demo, and the 2026 OCR showdown table.

01The problem: screenshot hoarding

You know exactly which screenshot has the thing. You have no idea which file it is.

The author's words: "i screenshot everything and can never find anything." Receipts, error messages, wifi passwords, that one tweet, the terminal output you swore you'd remember — they pile up in ~/Pictures/Screenshots as shot_4821.png, shot_4822.png, shot_4823.png. Your OS can search filenames. It cannot search inside the pixels.

gyotaku fixes that with the oldest trick in computing: build an index. Every screenshot gets read once by OCR; every word goes into a searchable database. After that, finding anything is a database lookup — milliseconds.

02How it works

Walk the pipeline one step at a time. The screenshots below are simulated drawings, not real OCR output.

SIMULATED WALKTHROUGH

1 / 5WATCH

03Try it: the simulated demo

Eight fake screenshots, drawn live on canvas. Type any word that appears in one — the library filters instantly, exactly like gyotaku's search box. All content below is simulated.

8 of 8 screenshots
No screenshots contain that. Try "wifi", "banana", "deploy"…

04The 2026 OCR showdown

Seven engines, one table. Numbers are directional — drawn from 2026 benchmarks and project docs, linked below — not lab-grade measurements.

EngineAccuracySpeed (CPU)Model sizeLanguagesFully offlineLicense
PaddleOCR PP-OCRv5
★ gyotaku's pick
~4.5% char error rate; 92.86 OmniDocBench; matches GPT-4o on some OCR tasks at 5M params ~1.5–2.5 s per A4 page 7.5 MB recognizer (EN) + 84 MB detector 109 Yes Apache 2.0
Tesseract 5.x ~18% char error rate on same benchmark; fine on clean print ~2–5 s per page ~30 MB 100+ Yes Apache 2.0
EasyOCR Good on print; weak on the hard stuff (0.26 edit distance vs 0.07 PaddleOCR on OmniDocBench EN) 3–5× slower than PaddleOCR ~2 GB install (drags PyTorch) 80+ Yes Apache 2.0
RapidOCR (ONNX) Same as PaddleOCR — it runs PP-OCR weights Fast on CPU ~40–150 MB total, no torch Same as PaddleOCR Yes Apache 2.0
Surya 2 Best-in-class layout + tables; strong multilingual Slow per page; wants a GPU 0.65B params (~1.3 GB+) 90+ Yes, but needs GPU OpenRAIL (not fully OSS)
Apple Vision Excellent on Latin scripts Real-time on-device Built into the OS Broad Yes Proprietary, Apple-only
Google ML Kit Very good, mobile-tuned Real-time on-device ~10s of MB Broad Yes Proprietary, mobile-only

Sources: character-error figures from a 2026 community benchmark (github.com/nickrivers1983/opencode-vision); PP-OCRv5 specs, 60k+ stars and GPT-4o-matching claim via abit.ee, Mar 2026; model sizes via huggingface.co/monkt/paddleocr-onnx; Tesseract standing via koncile.ai; Surya/OpenRAIL and VLM rankings via 2026 community research docs; EasyOCR weight from multiple migration notes (2 GB → 150 MB switching to RapidOCR).

05So what should you pick?

There is no "best OCR". There is best OCR for a constraint.

The honest caveat on gyotaku's stack: PaddleOCR's models are the right call, but its official runtime is Python. A native Rust/gpui app would most sensibly run the ONNX-exported PP-OCRv5 weights through the ort crate — the RapidOCR playbook, not pip install paddleocr.

06The rest of the stack

SQLite FTS5

The index lives in a single file. FTS5 full-text search turns "find the word" into a B-tree lookup — that's the milliseconds.

gpui

Zed editor's Rust GPU UI framework. Native widgets, GPU-composited, no Electron — search-as-you-type stays at 60fps.

fully offline

No server, no API key, no per-page pricing. Your screenshots never leave the disk — the privacy story is the feature.