After It Works: Trusting and Teaching Alchemy, DoorDash's AI Moderation Platform

Most AI talks end at "it works." This one starts there. At DoorDash we built Alchemy, a content-agnostic moderation platform that any team can point at any user-generated content, such as text, images, voice, to run AI moderation and get back a decision. Under the hood, Alchemy agents follow a pattern: a cheap in-house classifier gates an expensive LLM that scores severity, so the ~90% of obviously-fine content never touches expensive layers. But shipping an AI that makes real decisions is the easy part. The hard part is everything after.

How do you know a moderation decision is right when there's no clean ground truth and the judge is non-deterministic? How do you change a prompt or swap a model without flying blind? And how do you keep an expensive LLM from staying expensive forever? In this talk I'll go one layer deeper than the architecture: the evaluation harness built into Alchemy that lets teams trust, and safely change, an AI decision-maker in production (shadow mode, backtesting on real historical data, labeling, and tying metrics to incident reduction rather than model accuracy), and the retraining flywheel that turns the LLM's own judgments into training data for the cheap classifier, shrinking latency and cost over time. Real ML and engineering trade-offs, the failure modes we hit, and a couple of patterns you can take home to any high-volume AI system.


Speaker

Bruna Pereira

Bruna Pereira

Software Engineer @DoorDash, 10+ Years in Software Engineering

Bruna Pereira is a software engineer at DoorDash with 10+ years of experience in software engineering. She enjoys solving hard problems, working on systems that scale, and learning from building things in production.

When she's not coding, you'll likely find her at the beach or playing beach tennis.

Read more →
Find Bruna Pereira at: