All articles

Who reviews your listings before they go live: what each option actually costs when it breaks

9 min read

The question you cannot shop your way out of

At some point after launch, every classifieds operator hits the same wall: listings are coming in faster than one person can look at each one before it goes live, and the site needs a real answer to who reviews what, and how fast. The instinct is to go looking for a price, a vendor, a plan that makes the problem go away. That instinct runs into a wall of its own almost immediately, because the pricing that outsourced moderation providers publish is close to useless for planning a real budget. Rates quoted on vendor marketing pages vary by an order of magnitude from one company's blog post to the next, they rarely say what a review decision actually covers, and the few numbers that get repeated across the industry mostly trace back to the same handful of self-published sources restating each other rather than any independent benchmark. An operator who budgets from those numbers is budgeting from a guess.

That does not mean the decision is unknowable, only that it has to be made a different way: not by finding the right price, but by understanding which architecture fits the site's actual volume and risk, and what each one costs in ways that never show up on an invoice. The three options, in-house human review, outsourced human review, and AI-first filtering, do not fail in the same places, and knowing where each one breaks is worth more than any per-item rate a sales call will offer.

This matters for a directory specifically because the content is not neutral. A photo, a headline, or a stated availability window can be a completely legitimate advertisement or a listing that has to be pulled and, in some cases, reported rather than simply deleted. Getting review wrong in the direction of being too slow lets bad listings sit live; getting it wrong in the direction of being too aggressive kills legitimate advertisers' trust in the site. Both failure modes cost money, and neither shows up in a vendor's demo.

Why the three architectures fail differently

Pure AI-first filtering, an automated classifier scoring every listing and auto-approving or auto-rejecting based on that score, is the cheapest option to run at volume and the fastest to approve a clean listing. Its failure mode is context. A classifier trained to catch obvious violations is good at catching obvious violations, and comparatively weak at anything that depends on context it was never shown: a phrase that reads as a violation out of context but is fine in the actual listing, or the reverse, a listing that clears every keyword filter while violating the site's actual policy in a way no filter was built to catch. An operator running AI-only review finds out about both failure modes only after an advertiser complains publicly about a wrongful rejection, or after something that should never have gone live already has.

Pure human review, whether in-house or outsourced, is the most accurate option per decision and the one that scales worst. A human reviewer catches context a classifier misses, but a human reviewer also has a fixed number of decisions they can make competently in an hour before quality drops, and that number does not move no matter how urgently the queue is growing. A traffic spike, a marketing push that works better than expected, a seasonal surge, turns a review queue that was keeping pace into a backlog within days, and a backlog in listing review is not a neutral wait, it is either listings going live unreviewed to clear the queue or legitimate advertisers waiting long enough that they take their listing, and their money, somewhere else.

The hybrid model, an automated first pass that clears the unambiguous cases and routes anything it is not confident about to a human, is what nearly every platform operating at real scale actually runs, for a straightforward reason: it is the only structure where the machine's speed and the human's judgment cover for each other's weak point instead of both being exposed to it. The design decision that actually matters here is not whether to use a hybrid model, almost everyone should, it is where the confidence threshold sits. Set it too low and the AI defers almost everything to a human, which reproduces the backlog problem the hybrid model was supposed to solve. Set it too high and the AI is quietly auto-approving or auto-rejecting listings a human would have judged differently, and nobody finds out until one of those listings becomes a problem.

The cost that never appears on an invoice

Whoever does the human review, in-house staff or an outsourced team, is looking at the worst material a site receives before anyone else does: the listings that get flagged, appealed, or escalated precisely because something about them is disturbing enough to need a second look. Treating that exposure as a routine data-entry task, something any interchangeable reviewer can be swapped into without consequence, is not just an unkind way to run a team. It is a documented source of liability, and the documentation exists because a much larger platform already tested the alternative and lost.

In May 2020, Meta agreed to a 52 million dollar settlement in Scola v. Facebook, a lawsuit brought by a former content moderator who developed post-traumatic stress disorder from sustained exposure to graphic material in the course of the job. The settlement covered moderators who worked for a Facebook contractor in California, Arizona, Texas, or Florida between September 2015 and August 2020: a guaranteed payment to every eligible moderator, with an additional award of up to 50,000 dollars available to those diagnosed with PTSD or a related condition, plus an ongoing commitment to make mental health counseling available to moderators doing the job. The case did not turn on Facebook's moderation policy being wrong. It turned on the working conditions under which that policy was carried out being treated as someone else's problem for years before anyone had to answer for it.

A small directory is not staffing a review operation at Facebook's scale, and none of this means outsourcing or hiring for the role is itself a mistake. It means the operational habits that created that exposure, no rotation off the worst content, no cap on how long a session of graphic review runs, no access to support built into the job rather than offered as an afterthought, are exactly as available to a two-person moderation setup as to a fifteen-thousand-person one. The fix costs far less than the failure: rotate who handles escalated content rather than parking it permanently with whoever happens to be fastest at it, set a real limit on how long a single review session runs before a break is required, and make sure whoever does this work knows, before their first day, what kind of material they will see and where to go if it affects them.

What a vendor contract needs to say, and what a quote from one cannot tell you

Since the published per-item and monthly rates are not reliable enough to budget from, the only number worth trusting is the one an operator measures directly: a pilot, run on the site's own listings at real volume, timed against the site's own turnaround target, not a vendor's demo reel or a reference client in an unrelated industry. A pilot that cannot be timed against real listings, because a vendor prefers to quote a rate first and demonstrate performance later, is not a pilot worth running.

What the contract needs to specify, regardless of price, is the operational detail vendors leave vague by default: which languages are covered by native or fluent reviewers rather than machine translation, what the actual response-time commitment is measured against, not "usually same day" but a stated number of hours, what happens to the identity documents and photos a reviewer sees in the course of the job, who has access to them, and how long they are retained, and what the appeals path looks like when an advertiser disputes a rejection. A vendor that cannot answer all four of those in writing before a contract is signed is not a vendor an operator has actually evaluated, whatever the sales deck says about their pricing.

The other clause worth insisting on is an exit: a contract term short enough, or a termination clause loose enough, that switching providers after a bad first quarter does not mean waiting out a year of paid-for underperformance. Vendors that resist this are signaling something about their own confidence in the numbers on the pitch deck.

Building the review process before shopping for one

The sequence that actually works starts before any vendor conversation. Write the decision policy first: a short document, not a legal brief, that states plainly what gets an automatic pass, what gets an automatic reject, and what goes to a human, tied directly to the site's own terms of service rather than left to a reviewer's individual judgment call each time. Without that document, every reviewer, in-house or outsourced, is inventing the policy listing by listing, and two reviewers will disagree on the same listing in ways that show up as inconsistent enforcement the moment an advertiser compares notes with another.

Log every decision from day one, even as a spreadsheet: what was reviewed, what was decided, and when. This is not extra work invented for its own sake; it is the same record the EU's notice-and-action rule already requires an operator to keep for any listing flagged as illegal content, and building the habit before it is a compliance obligation means it is already running by the time it becomes one.

Start in-house, even if the plan is to outsource eventually. A month of an owner or a single hired reviewer handling every decision by hand produces the only volume and complexity numbers worth budgeting from, because they come from the site's actual listings rather than a vendor's estimate of a typical customer. That month also surfaces the edge cases, the listing type that is genuinely hard to judge, the phrase that keeps needing a second look, that a written policy has to account for before a larger team or an outside vendor is asked to apply it consistently.

Only after that baseline exists does a vendor conversation produce a useful answer, because at that point there is a real number to hold a pilot's performance against, a real list of edge cases to test a vendor with, and a real basis for judging whether a quoted rate reflects what the job actually requires or just what the vendor thinks the market will accept. Shopping for moderation before that groundwork is done means shopping blind, no matter how confident the number on the quote looks.

Try the DEMO

Escort directory software, ready to go