Hands-on with the OpenAI Codex App · Challenge 02 · 中文版

Lesson 2: Write the task contract, let AGENTS.md carry the rules

Checkpoint l0210 ptsself-reported — code from verify.pyOpenAI Codex desktop app

Turn "make it better" into a bounded, testable contract — then make it durable so every Codex thread starts with the same rules.

Six agent-tool courses share this challenge. OpenAI Codex desktop app-specific steps are still being written — until they land and pass review, this page defers its canonical to the primary course's copy.

Doing this with your agent? One sentence starts the whole course:

Read https://flypython.com/skills/flypython/SKILL.md and start the FlyPython course `hands-on-with-openai-codex`.

Lesson 2: Write the task contract, let AGENTS.md carry the rules

Objective

You can read TASK.md as a set of testable statements, trace each statement to a test in tests/test_report_tool.py, and write durable rules a Codex thread will pick up automatically.

Why this lesson exists

Vague requests produce vague code. “Handle bad rows better” gives an agent permission to guess; “invalid rows are collected in errors with index and reason, valid rows still produce a report” gives it a target and gives you a way to check. In the Codex app there is a second failure mode: a rule you typed into one thread does not exist in the next. The contract fixes the task; AGENTS.md fixes the working rules.

The lesson

Open TASK.md. Notice what every line has in common: it names an observable behavior, not an implementation. Four statements from the contract, and the tests that pin them:

Contract lineTest
“a JSON file loads into the same record list as CSV”test_load_json_records_returns_list_of_dicts
“invalid rows land in errors with index and reason; valid rows still aggregate”test_invalid_records_are_isolated_with_reasons
“group totals are rounded to two decimals”test_group_totals_are_rounded_to_two_decimals
“the report write is atomic and creates missing parents”test_write_report_creates_missing_parent_directories

Now the durable half. Codex reads AGENTS.md files before doing any work — a global file in ~/.codex, then project files from the repo root down to your directory; closer files win. This course’s working rules are exactly the kind of thing that belongs there. Open a scratch file and draft three rules for this project, for example:

- Change only starter/report_tool.py. tests/, solution/, scenario/ are read-only.
- One failing test group per turn; run `python verify.py starter` after each.
- Standard library only — no new dependencies.

Tell the thread: “Here are the rules I want for this project — write them to AGENTS.md at the folder root so every future thread starts with them.” Then start a new thread and ask it to summarize its working rules. If it quotes your three lines back, the instruction chain is working; if not, check where the file landed.

Exercise

Write one contract line for a script you actually own, using the same shape: inputs, outputs, error cases, and “done means <command> exits 0”. Then write the two rules you would put in that project’s AGENTS.md.

Checkpoint

Run python verify.py — this checkpoint’s claim code prints when you can answer:

  1. Which TASK.md line does test_invalid_records_are_isolated_with_reasons pin, in your own words?
  2. Which file carries rules across threads — the prompt, or AGENTS.md?
  3. Where does a global ~/.codex/AGENTS.md sit in precedence versus the project file?

Expected evidence

Your drafted contract line, your two rules, and the thread’s summary of its own working rules from the new thread.

Hints

Stuck? Open one at a time.

Hint 1 — What this checkpoint tests

That you can read TASK.md as testable statements, not prose. Every sentence that starts with a function name is a contract line the suite can assert.

Hint 2 — AGENTS.md is the instruction chain

Codex reads AGENTS.md before doing work: a global file in ~/.codex plus project files from the repo root down. Rules you want every thread to follow live there — not retyped in each prompt.

Hint 3 — The testability check

If you cannot tell whether a statement is testable, ask: could a suite assert it without reading your mind? Numbers, exit codes, files on disk — never vibes.

Submit

Done? Record it.

The 8-character code verify.py progress printed for this checkpoint — from the browser, or straight from your agent.

Youbrowser

Submit the claim code

Your agentauto-submit

Let the agent solve and submit

Hand it this checkpoint (copy button on hover) — it solves, verifies, and submits on its own:

Work on the FlyPython challenge "Hands-on with the OpenAI Codex App" (course id course-codex-cli), checkpoint l02.
Machine-readable brief: https://flypython.com/api/challenges/hands-on-with-openai-codex
Open the course folder and read TASK.md first — it is the contract.
Rules: smallest change, no new dependencies, never edit tests/ or solution/.
Check with python verify.py until its gates pass, then report each claim code to me.
Target: solve only checkpoint l02, and submit it as soon as it passes.

With a token from your agent page (env FLYPYTHON_TOKEN), it records the result directly:

curl -X POST https://flypython.com/api/claims \
  -H "Authorization: Bearer $FLYPYTHON_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"claims":[{"course":"hands-on-with-openai-codex","checkpoint":"l02","code":"<8-char code>"}]}'

Or install the full FlyPython agent skill once and skip the paste.

Files

The starter and verifier

Solving runs on your machine — your agent fetches the course files itself via /api/challenges/hands-on-with-openai-codex/files; you download nothing by hand. Prefer reading the source? Browse the folder on GitHub ↗