Quality control is usually described as a filter at the end of a pipeline. We treat it as the pipeline itself: a sequence of gates that starts before a contributor writes a single word and continues until a batch is delivered. This note describes each gate and how we measure whether it is working.
Credential verification
Every applicant submits a professional registry number appropriate to their field: OAB for lawyers, CRM for physicians, CRC for accountants, CREA for engineers. We check each registration directly against the issuing professional board, confirm it is active and in good standing, and revalidate it periodically for as long as the contributor works with us.
A registry check establishes identity and licensure, not judgment. That is what the exam is for.
The practical entry exam
Applicants complete a practical exam in their own specialty: a small set of tasks built to the same standard as production work, graded against a rubric by senior reviewers. Most applicants do not pass. That selectivity is intentional, and it is the single largest quality lever we have, because no downstream review can fully compensate for a weak contributor pool.
We measure pass rates by area and over time, and we investigate any sudden shift in either direction, since it usually signals a change in the exam, the grading or the applicant pool rather than a change in talent.
Layered review by senior practitioners
Rubrics are written by the senior tier of our network, the same people who review production work. Review is layered: a trajectory is read by at least one senior reviewer who did not author it, against a rubric that scores each step, each decision point and the final artifact separately. Reviewers annotate rather than merely score, so a rejected trajectory carries a written explanation the contributor can learn from.
Because rubrics are written before tasks are distributed, contributors know the standard in advance, and review measures execution against a public yardstick rather than a private one.
Agreement between reviewers
A rubric that different reviewers apply differently is not a rubric yet. We measure agreement between reviewers continuously, using overlapping assignments in which the same work is graded independently by more than one reviewer.
When agreement on a task family falls, we do not average the disagreement away. We treat low agreement as a signal about the task or the rubric, not about the reviewers: the rubric is revised, the task is rewritten, and the affected work is graded again under the corrected standard. Only when agreement recovers does the task family return to production.
Sampling and audit
Beyond full review of flagged work, we audit a random sample of accepted trajectories from every batch. Audits regrade the sample blind, without access to the original scores, and we track the rate at which audit disagrees with initial acceptance. That rate is one of the core metrics we report internally on every delivery, because it measures the health of the whole pipeline rather than of individual contributors.
Detecting model generated submissions
Contributors are forbidden from submitting model generated work as their own, and we enforce this rather than relying on the rule. Detection combines signals: metadata about how the work was produced inside our instruments, statistical analysis of the text, and review by seniors who know what authentic practitioner reasoning looks like in their field. Suspected violations trigger a manual investigation, and confirmed ones end the relationship.
We are explicit that detection is probabilistic, which is why it never acts alone. No contributor is sanctioned on a score without human review of the underlying work.
Rejection as a feature
The most important property of this system is cultural rather than technical: we prefer to reject a batch over shipping an uncertain one. A rejected batch costs us time and margin. An uncertain batch that ships costs the customer training signal and costs us the only asset that matters in this business, which is trust in the label.
Every metric above is something we measure continuously, not a number we claim in advance. If your team wants to inspect how this applies to a specific domain, book a data consultation and we will walk through the quality design for your use case.