This is what you get after a candidate takes a code review test. The candidate is fictional; the task, the planted problems and the grading are real. Send your first test free.

Sample report

Sam Lindqvist (fictional)

Password reset (request + confirm) · Senior backend · reviewed a pull request written by an AI coding agent, against a written spec

Score band

9 to 10 out of 10

Caught nearly every planted problem and explained why it matters, including one that needed a careful read of the spec.

For comparison: a plain AI review scores 7 to 8 on this test.

Found 1 of the 2 problems a plain AI review usually misses.

Sam caught 5 of 6 planted problems, including the most serious security holes, and explained the impact of most of them in plain terms. They missed "Re-implements normalizeEmail with different behaviour". Their verdict to request changes was right. One comment flagged an import as unused when it is used, a minor mistake.

Verdict
Request changes (right call)
Time spent
About 27 minutes
Spec reading
Found 1 of 1

Found 1 of 1 problems that needed reading the requirements: the code looks fine on its own, but doesn't do what the spec asks.

Caught (5 of 6)

  • Reset link origin taken from the request Host header

    Serious

    The link in the reset email is built from whatever website name the requester claims to be on. An attacker can trigger a reset for a victim with a fake host, so the victim's email contains a link to the attacker's site that hands over the secret reset token, leading to account takeover.

    In their words, on src/routes/passwordReset.ts, line 30
    “The reset link is built from req.protocol and the Host header. Anyone can send Host: evil.com when requesting a reset for someone else, and the victim gets a link that leaks the token to that domain. Use config.appUrl like the spec says.”
  • Existing sessions are not revoked after the reset

    SeriousNeeded the spec

    After a password reset, anyone who was already logged in (for example an attacker who stole the account) stays logged in. The spec requires logging the user out everywhere, which is the main reason people reset a password after a compromise.

    In their words, on src/services/passwordReset.ts, line 108
    “revokeAllSessions is imported but never called. After a reset an attacker's existing session would survive. Spec says log out everywhere.”
  • Confirm handler not wrapped in asyncHandler

    This app's web framework does not handle errors from this kind of code on its own, and the project has a helper for that which this one endpoint skips. Every failed confirmation (wrong token, short password, database hiccup) leaves the user's request hanging until it times out instead of showing an error.

    In their words, on src/routes/passwordReset.ts, line 57
    “router.post('/confirm', confirm) passes the async function directly. On Express 4 the rejected promise is never passed to next(), so HttpError(400)/429 and DB errors become unhandled rejections and the request hangs with no response. Wrap it in asyncHandler like the other routes.”
  • Token lookups can't use the index (user_id leads it), and hashes aren't unique

    Every reset confirmation looks a token up in a table that keeps growing, but the index the code adds is organised by user first, so the database can't use it for this lookup. Confirmations get slower with every reset ever issued.

    In their words, on migrations/20261004_password_reset_tokens.sql, lines 10 to 11
    “The migration replaces the unique index on token_hash with a composite index on (user_id, token_hash). confirm looks tokens up by token_hash alone, and a btree can't be used without its leading column, so each confirm is a sequential scan on a table the spec says grows to millions of rows. It also drops the uniqueness guarantee. Index token_hash on its own, UNIQUE.”
  • Single-use test cannot fail

    The test meant to prove a reset link only works once is written so it passes even if the link works twice. It gives false confidence about one of the feature's key security rules.

    In their words, on test/passwordReset.test.ts, lines 69 to 72
    “The 'can only be used once' test swallows the second confirm's error and replays the same NEW_PASSWORD, so the final password check passes whether or not the token was reusable. Assert that the second call rejects with invalid_or_expired_token (and use a different password).”

Missed (1)

  • Re-implements normalizeEmail with different behaviour

    Instead of reusing the project's existing helper for cleaning up email addresses, the code adds its own slightly different copy. Two versions of the same rule drift apart, and here the copy forgets to strip spaces, so the per-address limit can be dodged by adding a space.

Also noticed

Real issues they raised that weren't on our list. They count in the candidate's favour.

  • migrations/20261004_password_reset_tokens.sql, line 1

    “Nit: I'd split the request and confirm handlers into separate files once this grows.”

    A fair maintainability point outside the planted problems.

False alarms

Problems they claimed that aren't real. A small, capped deduction.

  • migrations/20261004_password_reset_tokens.sql, line 1

    “Is this import unused? Looks like dead code.”

    Why it isn't a problem: The import is used further down the file.

AI tools they said they used: ChatGPT to double-check the race condition

Three questions for the interview

  1. Walk me through src/routes/passwordReset.ts around line 5. Does it do what the spec asks? What would you change?

    Listen for: The route defines its own normalizeEmail (lower-case only) instead of importing the one from models/users (trim + lower-case). Duplicated logic that already diverges: ' ada@example.com' gets a separate rate-limit bucket but still resolves to the same user. Import the shared helper.

  2. You flagged the rate limit key. How would you test that the per-IP and per-address limits work behind the load balancer?

    Listen for: Uses req.ip with trust proxy, tests both limits separately, and checks the 202 is still returned when the per-address limit silently skips sending.

  3. If you had 10 more minutes on this review, what would you check next, and why?

    Listen for: Prioritises by impact: token handling and session revocation before style, and reads the spec again for rules not yet checked.

What this doesn't measure: how they write code from scratch, system design, how they work with a team, or how they do on a codebase they know. It's one 25-minute review of one pull request.

The score comes from a fixed answer key. An AI grader checks each planted problem against the candidate's own words, and you can flag any decision you disagree with.

A person on your side makes every hiring decision.