agentreadme

Report · marked 9 Oct 2026

unclecode/crawl4ai

An agent can work here, but it will waste turns finding its footing.

Python · 85,040 stars · 948 files · branch main

Fix these first

01
Agent instruction file
Add an AGENTS.md at the root: what the project is, how to install, how to run, how to test, and the two or three conventions a newcomer always gets wrong.
02
Instruction quality
Instructions earn their keep by naming exact commands. 'Run the tests with `pnpm test`' beats three paragraphs of philosophy.
03
File sizes fit in context
A single file that fills the context window forces an agent to work from fragments, and it will confidently edit code it never saw. Splitting the worst offenders pays for itself immediately.

The full marking

Every deduction below names the file or setting it came from.

Instructions 3/27 F

Whether the repo tells an agent how to behave before it starts guessing.

✗Agent instruction file 0/12
None found. Every agent that opens this repo starts from zero.
Add an AGENTS.md at the root: what the project is, how to install, how to run, how to test, and the two or three conventions a newcomer always gets wrong.
Looked for AGENTS.md, CLAUDE.md, .github/copilot-instructions.md, .cursorrules
✗Instruction quality 0/12
Nothing to judge, since there are no instructions.
Instructions earn their keep by naming exact commands. 'Run the tests with `pnpm test`' beats three paragraphs of philosophy.
✓README as an entry point 3/3
README has a clear getting-started section.
Setup 16/20 A-

Whether an agent can install the project and get it running without a human.

✓Deterministic install 6/6
uv.lock pins the dependency tree.
Found uv.lock
✓Discoverable commands 6/6
Commands are declared where an agent will look for them.
pyproject.toml scripts
scripts/ directory
✗Environment config 2/4
Env vars are mentioned in the README but there's no example file to copy.
Add a .env.example listing every variable with a safe placeholder value. It's the cheapest possible fix and it unblocks the whole first run.
✗Pinned runtime version 0/2
No pinned runtime version.
Add a .nvmrc, .python-version, or equivalent so the agent's toolchain matches yours.
✓Reproducible environment 2/2
dockerfile gives a known-good environment.
Verification loop 19/25 B+

Whether an agent can check its own work. This is the category that most decides whether agent output is trustworthy.

✓Tests exist 8/8
276 test files against 548 source files.
e.g. deploy/docker/tests/conftest.py, deploy/docker/tests/demo_monitor_dashboard.py, deploy/docker/tests/requirements.txt
✓Test command is discoverable 7/7
An agent can find and run `pytest`.
✓Continuous integration 4/4
4 GitHub Actions workflows define what "passing" means.
✗Lint and format rules 0/3
No linter or formatter config. Style is tribal knowledge, so agent output will drift from yours.
A formatter config is the cheapest way to stop reviewing whitespace in agent diffs. It moves style from opinion to a command.
✗Static type checking 0/3
No static type checking.
A type checker gives an agent an error message instead of a runtime surprise. It's the second-fastest feedback loop after the compiler.
Context economy 13/20 B

Whether the repo fits in a context window, or fights it.

✓No committed build output 6/6
No generated directories committed.
✓.gitignore hygiene 3/3
.gitignore covers 186 patterns.
✗File sizes fit in context 2/6
4 source files are over 100KB.
A single file that fills the context window forces an agent to work from fragments, and it will confidently edit code it never saw. Splitting the worst offenders pays for itself immediately.
crawl4ai/utils.py — 127KB
crawl4ai/async_configs.py — 121KB
crawl4ai/async_crawler_strategy.py — 119KB
✗Repository weight 2/5
About 151MB checked out. Large enough that cloning and searching are both slow.
Large binaries and vendored trees slow every operation an agent performs. Git LFS or a separate assets repo keeps the working tree navigable.

Here is your AGENTS.md

Drafted from what is actually in this repository: the install command from your lockfile, the commands you already declare, your real directory layout. Anything marked TODO needs a person. Save it at the root as AGENTS.md.

AGENTS.md — drafted for unclecode/crawl4ai
# AGENTS.md

Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.

## Setup

```
uv sync
```

## Commands

```
# TODO: nothing declares a build or test command, so an agent has to guess.
# This is the single most valuable section of this file. Fill it in.
```

## Layout

- `tests/`    230 source files
- `docs/`     172 source files
- `crawl4ai/` 86 source files
- `deploy/`   55 source files
- `scripts/`  2 source files

## Conventions

- Tests live alongside the code they cover, following `deploy/docker/tests/conftest.py`.
- CI defines what passing means. See `.github/workflows/docker-release.yml`, and keep it green.
- TODO: add the two or three conventions a newcomer always gets wrong here.

## Gotchas

- `crawl4ai/utils.py` is 127KB. It will not fit comfortably in context, so read it in parts.

---

Drafted by agentreadme.com from what is in this repository. Everything marked TODO
needs a human. Check it in as AGENTS.md at the root.

Open the raw markdown  or  curl -o AGENTS.md agentreadme.com/draft/unclecode/crawl4ai.md

A draft gets the commands right and stops at the things only a maintainer knows. What a good Python AGENTS.md looks like covers what to add by hand, and what AGENTS.md is explains the format itself.

Show the mark

The badge re-checks daily, so it keeps up as the repository changes. Use mark again to force it now.

agent ready 60 out of 100

[![agent ready](https://agentreadme.com/badge/unclecode/crawl4ai.svg)](https://agentreadme.com/unclecode/crawl4ai)

Other Python repositories, marked

vinta/awesome-python 81 A- NousResearch/hermes-agent 84 A- TheAlgorithms/Python 89 A yt-dlp/yt-dlp 54 C+ microsoft/markitdown 48 C+ Significant-Gravitas/AutoGPT 65 B

All Python repositories

Think this mark is wrong?

Every deduction above names the file it came from, so this can be settled by looking. If a check missed something, that is a rule worth fixing.

Open an issue, already filled in