agentreadme

Report · marked 25 Aug 2026

apache/spark

An agent can work here, but it will waste turns finding its footing.

Scala · 43,875 stars · 27,310 files · branch master

Fix these first

01
Test command is discoverable
Add a `test` script or a `test` target that runs the whole suite with no arguments. The convention is the interface — an agent tries `npm test` and `make test` before it tries reading your CI config.
02
Deterministic install
Commit the lockfile your package manager generates. Without it, an agent's install and your install are different builds, and 'works on my machine' becomes unfalsifiable.
03
Discoverable commands
Declare the handful of commands that matter in package.json scripts, a Makefile, or a justfile. Naming them turns guesswork into a lookup.

The full marking

Every deduction below names the file or setting it came from.

Instructions 24/27 A

Whether the repo tells an agent how to behave before it starts guessing.

Agent instruction file 12/12
AGENTS.md is present, which is the format the most tools read.
Also found: CLAUDE.md
Instruction quality 10/12
Specific enough that an agent can act on it.
Worth fixing: has no code blocks.
19,783 characters, a workable length
names the actual commands to run
uses headings, so an agent can skim it
README as an entry point 2/3
README exists but never explains how to get the thing running.
Give the README an Install and a Usage heading with real commands under each. Agents pattern-match on those headings.
Setup 2/20 F

Whether an agent can install the project and get it running without a human.

Deterministic install 0/6
No lockfile. An agent installing dependencies may not get what CI got.
Commit the lockfile your package manager generates. Without it, an agent's install and your install are different builds, and 'works on my machine' becomes unfalsifiable.
Discoverable commands 0/6
No declared commands. An agent has to infer how to build and run this from the file tree.
Declare the handful of commands that matter in package.json scripts, a Makefile, or a justfile. Naming them turns guesswork into a lookup.
Environment config 2/4
Env vars are mentioned in the README but there's no example file to copy.
Add a .env.example listing every variable with a safe placeholder value. It's the cheapest possible fix and it unblocks the whole first run.
Pinned runtime version 0/2
No pinned runtime version.
Add a .nvmrc, .python-version, or equivalent so the agent's toolchain matches yours.
Reproducible environment 0/2
No container or devcontainer definition.
A Dockerfile or devcontainer removes an entire class of 'it won't install' failures.
Verification loop 15/22 B

Whether an agent can check its own work. This is the category that most decides whether agent output is trustworthy.

Tests exist 8/8
20700 test files against 9968 source files.
e.g. .github/workflows/build_and_test.yml, .github/workflows/maven_test.yml, .github/workflows/python_hosted_runner_test.yml
Test command is discoverable 0/7
There's no obvious way to run the tests. Even if tests exist, an agent won't reliably find the entry point.
Add a `test` script or a `test` target that runs the whole suite with no arguments. The convention is the interface — an agent tries `npm test` and `make test` before it tries reading your CI config.
Continuous integration 4/4
47 GitHub Actions workflows define what "passing" means.
Lint and format rules 3/3
pyproject.toml encodes the house style.
Static type checking n/a
Not applicable — the Java compiler type checks every build.
Context economy 9/20 C

Whether the repo fits in a context window, or fights it.

No committed build output 6/6
No generated directories committed.
.gitignore hygiene 3/3
.gitignore covers 129 patterns.
File sizes fit in context 0/6
104 source files are over 100KB.
A single file that fills the context window forces an agent to work from fragments, and it will confidently edit code it never saw. Splitting the worst offenders pays for itself immediately.
python/pyspark/sql/functions/builtin.py — 1135KB
sql/api/src/main/scala/org/apache/spark/sql/functions.scala — 660KB
python/pyspark/pandas/frame.py — 523KB
Repository weight 0/5
About 601MB checked out. Large enough that cloning and searching are both slow.
Large binaries and vendored trees slow every operation an agent performs. Git LFS or a separate assets repo keeps the working tree navigable.

Here is your AGENTS.md

Drafted from what is actually in this repository: the install command from your lockfile, the commands you already declare, your real directory layout. Anything marked TODO needs a person. Save it at the root as AGENTS.md.

AGENTS.md — drafted for apache/spark
# AGENTS.md

Apache Spark - A unified analytics engine for large-scale data processing

## Setup

```
# TODO: the install command. No lockfile was found, so this could not be inferred.
```

## Commands

```
ruff check .   # lint
```

## Layout

- `sql/`      5039 source files
- `python/`   1307 source files
- `core/`     1179 source files
- `mllib/`    610 source files
- `examples/` 523 source files
- `common/`   439 source files

## Conventions

- Tests live alongside the code they cover, following `.github/workflows/build_and_test.yml`.
- CI defines what passing means. See `.github/workflows/benchmark.yml`, and keep it green.
- TODO: add the two or three conventions a newcomer always gets wrong here.

## Gotchas

- Generated output is committed (`build/`). Search will return the compiled copy, so edit the source, not that.
- `python/pyspark/sql/functions/builtin.py` is 1135KB. It will not fit comfortably in context, so read it in parts.
- No lockfile is committed, so an install here may not match what CI produced.

---

Drafted by agentreadme.com from what is in this repository. Everything marked TODO
needs a human. Check it in as AGENTS.md at the root.

Open the raw markdown  or  curl -o AGENTS.md agentreadme.com/draft/apache/spark.md

Show the mark

The badge re-checks daily, so it keeps up as the repository changes. Use mark again to force it now.

agent ready 61 out of 100

[![agent ready](https://agentreadme.com/badge/apache/spark.svg)](https://agentreadme.com/apache/spark)

Other Scala repositories, marked

twitter/the-algorithm 31 D

All Scala repositories

Think this mark is wrong?

Every deduction above names the file it came from, so this can be settled by looking. If a check missed something, that is a rule worth fixing.

Open an issue, already filled in