an agent writes a review

Optional Week 4 exercise for BIOEE 7600-103

← back to the course page

Optional. Give Claude Code one prompt about a topic you know well. It searches the literature and writes a critical review, Annual Reviews style, with every reference checked. You judge it against what you know. Bring what you found on Friday 2 October.

How

  1. Pick a topic you know well, and what to compare against: a published review you trust, or your own knowledge. For a review, set the cutoff date to just before it came out, and do not name it in the prompt, since the model may have read it.
  2. Make an empty folder and open Claude Code there, set to Opus 5.5 (/model to check).
  3. Paste the prompt below with its four lines filled in. Here are the four lines John used for the run we will discuss on Friday:

    TOPIC: adaptation at geographic range edges
    SCOPE: plants and animals; theory and empirical work on whether and how populations adapt at geographic range edges. Out: purely ecological explanations of range limits, and species distribution modelling.
    CUTOFF DATE: 2020-08-31
    FULL TEXT VIA CHROME: no
    
  4. When it finishes, it gives you a link to your review as an Artifact, a page on claude.ai that only you can see.
  5. (Optional) Mark it up. Select a claim you disagree with, leave a comment, and choose Send to Claude. Claude replies and revises the page. Does it defend the claim, fix it, or give way? Your untouched first version stays in the folder as review_original.html, and the Artifact keeps every version (under Share).

By default it reads abstracts, plus the full text of open-access papers. From an abstract it cannot see methods or effect sizes, so it is weakest at judging evidence. The review says how many papers it read in full. See the “read full text” block below for instructions on how to let Claude read paywalled papers.

The prompt
# Fill in these four lines, then paste everything into Claude Code

TOPIC: [e.g. "adaptation at geographic range edges"]
SCOPE: [what is in and out: taxa, systems, scales, kinds of study]
CUTOFF DATE: [if you will compare against a published review: the date just before it came out, as YYYY-MM-DD. Otherwise "none". Do not name the review here.]
FULL TEXT VIA CHROME: [yes / no]

---

You are writing a critical review of the literature on the TOPIC above, of the kind published in
*Annual Review of Ecology, Evolution, and Systematics* or *Trends in Ecology & Evolution*. It is
for a researcher who knows the field and wants synthesis and judgement, not a list of papers.

## Rules

- **Cutoff.** If a CUTOFF DATE is given, use only work published on or before it. Do not read or
  cite anything later, including later reviews of this topic, even if a search turns them up.
- **Every reference must be real.** Cite only papers you found through a search in this session,
  each with a DOI. Never cite from memory.
- **Say what you read.** Each paper is marked as *full text* or *abstract only*, and the review's
  claims about a paper must not go beyond what you read of it.
- Work in the current folder. Do not read files outside it.

## Budget

This run should fit comfortably inside a standard usage allowance. Stay within these caps:
at most 5 search subagents, each reading at most 20 abstracts, about 25 papers kept for the
evidence table, and (if Chrome is on) at most 10 papers read in full through the browser.

## Steps

### 1. Plan (you, the main agent)
Split the TOPIC into up to 5 sub-questions that together cover it. Write them to `methods_log.md` with
one line on why this split. Choose the split by the mechanisms or hypotheses the field argues
about, not by taxon.

### 2. Search (one subagent per sub-question, in parallel, on Sonnet)
Launch one subagent per sub-question using the Agent tool with `model: "sonnet"`. Each one:
- Searches **OpenAlex** (`https://api.openalex.org/works?search=...&per-page=50`), adding
  `&filter=to_publication_date:CUTOFF` if there is a cutoff. Abstracts come back as
  `abstract_inverted_index`, which it must rebuild into text. It may also follow references and
  citations of key papers through OpenAlex. OpenAlex sometimes refuses searches or returns 503
  under heavy load: retry a few times with a pause, and if it keeps failing, search Crossref
  instead (`https://api.crossref.org/works?query=...&rows=50`, adding
  `&filter=until-pub-date:CUTOFF`) and fetch abstracts from OpenAlex by DOI or from Semantic
  Scholar. If the environment variable `OPENALEX_API_KEY` is set, add `&api_key=` with it.
- Screens titles and abstracts against the SCOPE, reading at most 20 abstracts, and keeps its
  share of the ~25 papers, favouring primary studies with strong designs, key theory, and prior
  syntheses.
- For each kept paper, fills one row: `doi, year, authors, title, sub_question, system,
  study_type (theory / observational / experiment / genomic / synthesis), design, main_result,
  evidence_strength (strong / moderate / weak, with one reason), read_as (abstract / full text)`.
- If OpenAlex lists an open-access copy (`best_oa_location`), reads that full text and marks the
  row *full text*. If the copy will not load as readable text, the row stays *abstract only*.
- Returns its rows, its search strings, and counts screened and kept.

Wait until all subagents have returned before going on.

### 3. Full text via Chrome (only if FULL TEXT VIA CHROME is yes; you, the main agent)
Pick the ≤10 kept papers most important to the argument that are still *abstract only*. For each,
open `https://login.proxy.library.cornell.edu/login?url=https://doi.org/<DOI>` in Chrome and read
the methods, results and discussion. Update its row. If you hit a login page, stop and ask me to
log in; never type a password. If a publisher blocks the page or a PDF won't render as text, note
it and move on.

### 4. Evidence table and first draft (you, on Opus)
Merge the rows into `evidence_table.csv`. Then write the draft as `review.md` (your working copy),
about 3,000 words:

1. **Introduction:** why the question matters and how the review is organised.
2. **What the evidence shows**, one section per sub-question. Weigh evidence by design and
   strength, not by number of papers. Say where findings rest on weak designs or on abstracts only.
3. **Where the field disagrees.** Name the real disagreements, who holds which position, and why
   the evidence so far has not settled them. Do not invent balance where there is consensus.
4. **What would settle it:** for each disagreement, the specific study or analysis that would.
5. **Ways forward:** concrete, not "more research is needed".
6. **Limits of this review:** how many papers were read in full vs. abstract only, and what that
   means for the conclusions.

At the top, one line: date, model, cutoff, papers screened / kept / read in full.

### 5. Critic (1 subagent, on Opus)
Give a fresh subagent (Agent tool, `model: "opus"`) `review.md` and `evidence_table.csv` only. Its job: act as a demanding
editor at *Annual Reviews*. It lists the five most serious problems: claims the table does not
support, missing counter-evidence, false balance, vague future directions, and any section that
summarises papers instead of synthesising them. Revise the review once in response. Record what
you changed and what you declined to change, with reasons, in `methods_log.md`.

### 6. Check every reference
Write a short script that looks up each DOI at Crossref (`https://api.crossref.org/works/<DOI>`)
and confirms the title and year match. Remove from the review anything that fails or falls after
the cutoff, and list the removals in `methods_log.md`. Save the cleaned list as
`references.csv`.

### 7. Publish it
Make the final review a readable page: light background, clear headings, each reference linked to
its DOI, and each paper marked *full text* or *abstract only*. Keep the wording of the review
unchanged. Save it as `review_original.html`, a copy that stays as it is if the review is revised
later. Then publish the same page as an Artifact, private to me, and give me the link. If you
cannot publish Artifacts in this session, tell me; the file is the review.

### 8. Finish
`methods_log.md` should end with: sub-questions, search strings, counts (screened, kept, full
text, removed at checking), and the critic's points. Then tell me in three lines what you are
least confident about.

Optional: read full text through the library

Set FULL TEXT VIA CHROME: yes, then:

  1. Install Claude in Chrome and sign in with your Claude account.
  2. In Chrome, log in to the library proxy by opening any paper through it, e.g. https://login.proxy.library.cornell.edu/login?url=https://doi.org/10.1086/703187.
  3. Start Claude Code with claude --chrome.

It reads up to 10 key papers this way. It is slow, uses more of your allowance, and some publishers block it. It never types a password; a login page stops it and it asks you. Outside Cornell, put your own library’s proxy in the prompt.

Usage limits, and running it overnight

If you hit your limit partway, Claude Code says when it resets; then type continue. The allowance resets every five hours, so a run started at the end of the day uses allowance you would not otherwise use. For an unattended run: stay for the first few minutes and choose “don’t ask again” when it asks permission for its searches and scripts, keep the computer plugged in and awake, and log in to the library proxy first if Chrome is on.

For Friday

Nothing to hand in. Bring your laptop with the review, and think about:

  1. What did it miss? The paper you would cite first? A line of work?
  2. Do the references say what the review says they say? Check two or three.
  3. Are its disagreements real, or balance it made up?
  4. Does it judge evidence, or count it?
  5. Would it be useful to you? Think about what may be lost with this approach, as well.