# AGENTS.md

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

## Setup

```
pip install -r requirements.txt
```

## Commands

```
pytest   # run the test suite
```

## Layout

- `deploy/`       262 source files
- `ppocr/`        254 source files
- `paddleocr-js/` 81 source files
- `paddleocr/`    72 source files
- `benchmark/`    63 source files
- `tests/`        43 source files

## Conventions

- Tests live alongside the code they cover, following `api_sdk/go/client_test.go`.
- CI defines what passing means. See `.github/workflows/build_publish_develop_docs.yml`, and keep it green.
- TODO: add the two or three conventions a newcomer always gets wrong here.

## Gotchas

- `deploy/android_demo/app/src/main/cpp/ocr_clipper.cpp` is 135KB. It will not fit comfortably in context, so read it in parts.
- No lockfile is committed, so an install here may not match what CI produced.

---

Drafted by agentreadme.com from what is in this repository. Everything marked TODO
needs a human. Check it in as AGENTS.md at the root.
