Hands-on with DeepSeek Harness · Challenge 02 · 中文版

Lesson 2: Write the task contract, let cordis.yml carry the setup

Checkpoint l0210 ptsself-reported — code from verify.pyDeepSeek Harness

Turn "make it better" into a bounded, testable contract — then make the working rules part of the composed agent, not a forgotten prompt.

Six agent-tool courses share this challenge. DeepSeek Harness-specific steps are still being written — until they land and pass review, this page defers its canonical to the primary course's copy.

Doing this with your agent? One sentence starts the whole course:

Read https://flypython.com/skills/flypython/SKILL.md and start the FlyPython course `hands-on-with-deepseek-harness`.

Lesson 2: Write the task contract, let cordis.yml carry the setup

Objective

You can read TASK.md as a set of testable statements, trace each statement to a test in tests/test_report_tool.py, and make working rules part of the composed agent instead of a prompt you have to retype.

Why this lesson exists

Vague requests produce vague code. “Handle bad rows better” gives an agent permission to guess; “invalid rows are collected in errors with index and reason, valid rows still produce a report” gives it a target and gives you a way to check. With a composed agent there is a second failure mode: the rules you typed into one session are not in the next — unless they live in the composition.

The lesson

Open TASK.md. Notice what every line has in common: it names an observable behavior, not an implementation. Four statements from the contract, and the tests that pin them:

Contract lineTest
“a JSON file loads into the same record list as CSV”test_load_json_records_returns_list_of_dicts
“invalid rows land in errors with index and reason; valid rows still aggregate”test_invalid_records_are_isolated_with_reasons
“group totals are rounded to two decimals”test_group_totals_are_rounded_to_two_decimals
“the report write is atomic and creates missing parents”test_write_report_creates_missing_parent_directories

Now the durable half. In DeepSeek Harness the agent is declared in cordis.yml: the model plugin, the tool plugins it may use, the session storage where trajectories land, and the instructions it starts with. Put this course’s working rules where they survive restarts — in the agent’s configured instructions, for example:

- Change only starter/report_tool.py. tests/, solution/, scenario/ are read-only.
- One failing test group per step; run `python verify.py starter` after each.
- Standard library only — no new dependencies.

Start a new session and ask the agent to state its working rules. If it quotes yours back, the composition is carrying them; if not, check whether the instructions actually landed in the configuration the session loaded.

Exercise

Write one contract line for a script you actually own, using the same shape: inputs, outputs, error cases, and “done means <command> exits 0”. Then write the two rules you would put in that project’s agent composition.

Checkpoint

Run python verify.py — this checkpoint’s claim code prints when you can answer:

  1. Which TASK.md line does test_invalid_records_are_isolated_with_reasons pin, in your own words?
  2. Which artifact carries rules across sessions — the prompt, or the cordis.yml composition?
  3. Why does an append-only trajectory make a composed agent easier to audit than a chat transcript?

Expected evidence

Your drafted contract line, your two rules, and the new session’s statement of its working rules.

Hints

Stuck? Open one at a time.

Hint 1 — What this checkpoint tests

That you can read TASK.md as testable statements, not prose. Every sentence that starts with a function name is a contract line the suite can assert.

Hint 2 — Composition is the instruction chain

A harness agent is assembled in cordis.yml: which model plugin, which tool plugins, what system instructions, what session storage. Rules you want every run to follow belong in that composition — not retyped per session.

Hint 3 — The testability check

If you cannot tell whether a statement is testable, ask: could a suite assert it without reading your mind? Numbers, exit codes, files on disk — never vibes.

Submit

Done? Record it.

The 8-character code verify.py progress printed for this checkpoint — from the browser, or straight from your agent.

Youbrowser

Submit the claim code

Your agentauto-submit

Let the agent solve and submit

Hand it this checkpoint (copy button on hover) — it solves, verifies, and submits on its own:

Work on the FlyPython challenge "Hands-on with DeepSeek Harness" (course id course-deepseek-harness), checkpoint l02.
Machine-readable brief: https://flypython.com/api/challenges/hands-on-with-deepseek-harness
Open the course folder and read TASK.md first — it is the contract.
Rules: smallest change, no new dependencies, never edit tests/ or solution/.
Check with python verify.py until its gates pass, then report each claim code to me.
Target: solve only checkpoint l02, and submit it as soon as it passes.

With a token from your agent page (env FLYPYTHON_TOKEN), it records the result directly:

curl -X POST https://flypython.com/api/claims \
  -H "Authorization: Bearer $FLYPYTHON_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"claims":[{"course":"hands-on-with-deepseek-harness","checkpoint":"l02","code":"<8-char code>"}]}'

Or install the full FlyPython agent skill once and skip the paste.

Files

The starter and verifier

Solving runs on your machine — your agent fetches the course files itself via /api/challenges/hands-on-with-deepseek-harness/files; you download nothing by hand. Prefer reading the source? Browse the folder on GitHub ↗