Hands-on with Kimi Code · Challenge 02 · 中文版

Lesson 2: Write the task contract, plan it in the planning lane

Checkpoint l0210 ptsself-reported — code from verify.pyKimi Code CLI

Turn "make it better" into a bounded, testable contract — then let the plan subagent (no write, no shell) turn it into steps before anything edits.

Six agent-tool courses share this challenge. Kimi Code CLI-specific steps are still being written — until they land and pass review, this page defers its canonical to the primary course's copy.

Doing this with your agent? One sentence starts the whole course:

Read https://flypython.com/skills/flypython/SKILL.md and start the FlyPython course `hands-on-with-kimi-code`.

Lesson 2: Write the task contract, plan it in the planning lane

Objective

You can read TASK.md as a set of testable statements, trace each statement to a test in tests/test_report_tool.py, and use the planning lane to decompose the contract before any file is touched.

Why this lesson exists

Vague requests produce vague code. “Handle bad rows better” gives an agent permission to guess; “invalid rows are collected in errors with index and reason, valid rows still produce a report” gives it a target and gives you a way to check. Kimi Code’s built-in lanes make the discipline literal: plan can think about the work but holds no write or shell tools, so planning cannot accidentally become editing. And because subagents cannot spawn subagents, the work never recurses out of sight.

The lesson

Open TASK.md. Notice what every line has in common: it names an observable behavior, not an implementation. Four statements from the contract, and the tests that pin them:

Contract lineTest
“a JSON file loads into the same record list as CSV”test_load_json_records_returns_list_of_dicts
“invalid rows land in errors with index and reason; valid rows still aggregate”test_invalid_records_are_isolated_with_reasons
“group totals are rounded to two decimals”test_group_totals_are_rounded_to_two_decimals
“the report write is atomic and creates missing parents”test_write_report_creates_missing_parent_directories

Now let the planning lane do its job. Tell kimi:

“Use the plan subagent to turn TASK.md into an ordered implementation plan for starter/report_tool.py: one step per failing test group, ending each step with python verify.py starter. Do not edit anything yet.”

Read the plan it returns. It should name the same four behavior groups you just traced to tests, in an order where each step is independently verifiable. If a step is “improve error handling,” send it back — that is not a contract line. A plan you can check against TASK.md line by line is the deliverable of this lesson.

Also note what carries rules across sessions: if the repository root has an AGENTS.md, Kimi Code reads it — durable working rules live in files, not in a chat you will close.

Exercise

Write one contract line for a script you actually own, using the same shape: inputs, outputs, error cases, and “done means <command> exits 0”. Then sketch the two-step plan the plan lane should return for it.

Checkpoint

Run python verify.py — this checkpoint’s claim code prints when you can answer:

  1. Which TASK.md line does test_invalid_records_are_isolated_with_reasons pin, in your own words?
  2. Why does it matter that plan has no write or shell tools?
  3. Why does “subagents cannot spawn subagents” make review easier?

Expected evidence

Your drafted contract line and the ordered plan you accepted (or the version you sent back and why).

Hints

Stuck? Open one at a time.

Hint 1 — What this checkpoint tests

That you can read TASK.md as testable statements, not prose. Every sentence that starts with a function name is a contract line the suite can assert.

Hint 2 — Plan cannot touch anything

The built-in plan subagent produces a plan without shell or write tools — planning is a separate, side-effect-free lane. Subagents also cannot spawn subagents, so the lanes never recurse out of sight.

Hint 3 — The testability check

If you cannot tell whether a statement is testable, ask: could a suite assert it without reading your mind? Numbers, exit codes, files on disk — never vibes.

Submit

Done? Record it.

The 8-character code verify.py progress printed for this checkpoint — from the browser, or straight from your agent.

Youbrowser

Submit the claim code

Your agentauto-submit

Let the agent solve and submit

Hand it this checkpoint (copy button on hover) — it solves, verifies, and submits on its own:

Work on the FlyPython challenge "Hands-on with Kimi Code" (course id course-kimi-code), checkpoint l02.
Machine-readable brief: https://flypython.com/api/challenges/hands-on-with-kimi-code
Open the course folder and read TASK.md first — it is the contract.
Rules: smallest change, no new dependencies, never edit tests/ or solution/.
Check with python verify.py until its gates pass, then report each claim code to me.
Target: solve only checkpoint l02, and submit it as soon as it passes.

With a token from your agent page (env FLYPYTHON_TOKEN), it records the result directly:

curl -X POST https://flypython.com/api/claims \
  -H "Authorization: Bearer $FLYPYTHON_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"claims":[{"course":"hands-on-with-kimi-code","checkpoint":"l02","code":"<8-char code>"}]}'

Or install the full FlyPython agent skill once and skip the paste.

Files

The starter and verifier

Solving runs on your machine — your agent fetches the course files itself via /api/challenges/hands-on-with-kimi-code/files; you download nothing by hand. Prefer reading the source? Browse the folder on GitHub ↗