When your agent can't be the judge, ask people who can.
Send a brief and 2 to 4 options, images or text. People with a measured track record in the category pick the strongest and say why. Your agent keeps working and reads the verdict when it's ready.
- Two or more options are all competent and the choice depends on human perception: which reads as premium, trustworthy, calm, or on-brand.
- The audience's reaction matters and you can't observe it: what a first-time visitor, a senior engineer, or a hurried shopper would pick.
- Tone and wording choices: headlines, button text, names, error messages, subject lines.
- The decision is costly to undo (a launch page, a brand direction) and a few dollars of human judgment is cheap insurance.
- You have revised several times without converging, or your own confidence is low.
- Your user needs evidence for a choice: a vote split and written reasons are something you can show them.
- Objective checks a tool can verify: contrast ratios, broken layouts, typos, spec compliance, accessibility rules.
- Anything that needs an answer in seconds. Human answers take minutes; call quote_human_judgment first for a timing estimate.
People with a tier in the job's category earned by agreeing with other people on paid jobs. Choose the minimum: Rater from $0.10, Senior from $0.40, Expert from $1.50 per answer. Their pay is held until their account has a record, and accounts that answer like an AI model forfeit it and the money comes back to you. Each result also shows what a frontier model picked on jobs with a clear majority, so you can see where people saw it differently.
Sign in to get an API key
Every developer starts with $25.00 in test funds.
Signing in also sets up a payout account, so paid work can reach you in US dollars the day it passes review. No crypto knowledge needed.