Quality & Scoring
Quality Score Basics: Why Tasks Get Rejected (and How to Fix It)
Every data-labeling and RLHF platform scores your work differently under the hood, but the reasons submissions get rejected are surprisingly consistent across platforms. Most rejections come down to a handful of avoidable mistakes rather than genuinely ambiguous edge cases.
The most common rejection reasons
- Not re-reading the task-specific guidelines. General instructions change per project or per batch. Skimming once and relying on memory from a previous task is one of the biggest causes of avoidable errors.
- Rushing on "obvious" tasks. Reviewers flag inconsistency more than difficulty — a labeler who is careful on hard tasks but sloppy on easy ones often scores worse than one who is consistently careful.
- Missing edge-case instructions. Guidelines often bury exceptions ("unless the image contains X, in which case...") a few paragraphs in. These are exactly the rules reviewers check first.
- Formatting and structure errors. Especially in RLHF ranking or text-generation tasks, an otherwise-correct answer can be rejected for not following the required output format.
- Not using reference examples. Most projects provide a handful of "gold" examples. Skipping them and going straight to live tasks means you're guessing at a standard that's already been shown to you.
A pre-submit routine that catches most of it
Before you submit a batch, run through this in under a minute:
- Re-read the guideline summary or last update note for this specific project — not just the task itself.
- Check your answer against at least one gold/reference example, if provided.
- Confirm formatting: required fields filled, no placeholder text left in, correct structure for the answer type.
- For ambiguous cases, check whether the guidelines mention a tiebreaker rule before you guess.
- If a task feels genuinely unclear even after this, flag it instead of guessing — most platforms treat a flagged task more favorably than a wrong answer submitted with confidence.
Why this matters for your score, not just individual tasks
Quality scores on most platforms are trailing averages, which means a short streak of careless mistakes can outweigh a much longer streak of good work. Protecting consistency — especially on tasks that feel easy or repetitive — tends to move your score more than trying to be perfect on the hardest tasks alone.
Want more guides like this?
Ask us on WhatsApp →