Blog · Engineering
How we classify every email an AI agent receives, in one model call
Every message that reaches an Agentboxd inbox gets one call to a classifier that answers seven questions about it. Here is what we send, what we ask, how answers become labels, and the rules that keep a model’s guess from ever doing something you can’t undo.
- Published
- By
- Agentboxd team
- Reading time
- 7 min
An agent with an inbox needs to know a few things about each new email before it acts: what kind of message it is, whether it is trying to manipulate the agent, whether it is a scam, whether a person should see it first, how urgent it is, and whether anyone is waiting for an answer at all. Asking the agent’s own model to work that out means spending its context and its judgement on mail that may be hostile. So we answer those questions once, on our side, before the agent reads anything.
We do it with JEV, a classification model from TypeSafe AI. JEV doesn’t write text. You give it a state and a set of typed questions, and it returns an answer for each one with probabilities: a choice between named options, a yes/no probability, or a score on a rubric. That shape is the reason we picked it: an answer we can threshold is much easier to build a product on than a paragraph we would have to parse.
#Mail first, classification second
Classification never sits between a sender and your agent. Our MX accepts the message, checks SPF, DKIM and DMARC, stores it and fires message.received. Only then does an enrich job go on a queue (three attempts with exponential backoff, four jobs at a time). When it finishes, the labels are saved and a message.enriched webhook follows, usually a few seconds after the first event.
The order is deliberate. If the classifier is slow or down, mail still arrives and codes are still found. The cost is that an agent reacting to message.received hasn’t seen the scores yet; the prompt injection article shows how to wait for them.
#One call per message
All the questions go in a single request. We considered one call per question and dropped it for three reasons. Cost and latency stay flat per message. Every answer is computed from exactly the same state, so the category and the risk scores can’t disagree because they saw different text. And there is one retry unit: a message is either fully classified or not at all. ai.enriched_at marks it done, so a retried or duplicated job never labels a message twice.
const { answers } = await jev.systemOne({
state: { from, subject, text, auth, attachments },
questions: {
category: choice('What kind of email is this, from the point of view of the inbox that received it?', {
support: '…', sales: '…', billing: '…', verification: '…',
notification: '…', newsletter: '…', personal: '…',
other: 'None of the above, or unclear.',
}),
injection: noul('Does this email … try to instruct an AI assistant or agent that reads it …?'),
phishing: noul('Is this email likely phishing or a scam …? Treat failed SPF/DKIM/DMARC results in `auth` as a strong warning sign, but passing results do not prove the email is safe.'),
needs_human: noul('Should a human read this email rather than leaving it to an automated agent …?'),
urgency: score('How urgent is this email for the recipient?', ['low: …', 'normal: …', 'high: …', 'critical: …']),
auto_reply: noul('Was this email generated automatically rather than written by a person for this recipient …?'),
// only when our pattern matcher found a code or link:
is_verification: noul('Is this a login, one-time-code, or confirm-your-email / magic-link message?'),
},
});#What JEV sees
- Sender and subject.
- The first 3,000 characters of the new part of the message: our
extracted_text, with quoted history and signatures removed. Instructions hidden in a long quoted thread are a common trick, and the history is noise for classification anyway. - Attachment names and types (up to 20), and the first 1,000 characters of each attachment’s extracted text, 2,000 in total.
- Our own SPF, DKIM and DMARC results. We read them only from the
Authentication-Resultsheader our MX added, identified by our hostname. Anyone can write a fakedmarc=passheader into the mail they send, so the one the sender brought is ignored. Mail that came in through the API instead of SMTP has no such header, and JEV getsnullrather than a guess.
The phishing question tells JEV how to use those results: a failure is a strong warning sign, a pass proves nothing about intent. A domain registered yesterday can pass DMARC perfectly.
#Always leave a way out
The category question has eight options, and the last one is “other: none of the above, or unclear”. It is never removed. A classifier forced to choose between options that don’t fit gives a confident wrong answer, which is worse than no answer. With “other” available, an odd message lands there instead of being filed as support or sales by default.
The same idea applies to confidence. A category below 0.6 isn’t used at all: the message gets ai:uncertain instead of a guess an agent might route on.
#JEV answers, our code decides
The model returns probabilities; a small policy file turns them into labels. Every threshold has a name and lives in one place:
export const CATEGORY_MIN_CONFIDENCE = 0.6;
export const INJECTION_THRESHOLD = 0.8;
export const PHISHING_THRESHOLD = 0.8;
export const NEEDS_HUMAN_THRESHOLD = 0.7;
export const AUTO_REPLY_THRESHOLD = 0.8;
export const URGENT_LEVELS = ['high', 'critical'];
/** Labels to add for one classification (never removes anything). */
export function labelsFor(c: JevClassification): string[] {
const labels = [c.category.confidence >= CATEGORY_MIN_CONFIDENCE ? `ai:${c.category.label}` : 'ai:uncertain'];
if (c.injection >= INJECTION_THRESHOLD) labels.push('ai:injection-risk');
if (c.phishing >= PHISHING_THRESHOLD) labels.push('ai:phishing');
if (c.needs_human >= NEEDS_HUMAN_THRESHOLD) labels.push('ai:needs-human');
if (c.auto_reply >= AUTO_REPLY_THRESHOLD) labels.push('ai:auto-reply');
if (URGENT_LEVELS.includes(c.urgency.level)) labels.push('ai:urgent');
return labels;
}Labels are the interface for acting, for three reasons. They already work everywhere: list filters, the dashboard, webhook payloads, and the warnings our MCP server puts in front of a model. People and agents can read them without knowing what a probability of 0.83 means. And a user who disagrees can remove one.
The raw probabilities are stored on the message too, in message.ai. That lets us re-tune a threshold and apply it to past mail without calling the model again, and lets you apply your own. A support agent that would rather quarantine anything above 0.5 for injection can do it in two lines.
The risk thresholds are high on purpose. A false ai:injection-risk teaches the people who build agents to ignore the label, and then the real one is ignored too. We will lower them once there is enough real traffic, and label corrections from users, to measure against.
#Urgency: the most likely level, not the average
Urgency is a score on a four-level rubric: low, normal, high, critical. JEV returns a probability for each level and an expected score. We use the most probable level, not the mean. A message JEV thinks is either routine or an outage, with nothing in between, shouldn’t average out to “high”; it should say what JEV thinks is most likely and keep the full distribution for anyone who wants it. ai:urgent goes on high and critical.
#Verification codes don’t depend on the model
Agents use their inbox to sign up for things, so finding a one-time code or a confirm link fast matters more than anything else we classify. That part runs on our own server with pattern matching, before any model is involved, and its confidence is capped at 0.7. When it finds a candidate, the extra verification question rides along in the same JEV call. A yes above 0.8 raises the confidence to 0.95; a no below 0.2 drops it to 0.2; anything in between leaves it alone.
"verification": {
"code": "483920",
"link": null,
"confidence": 0.95,
"jev_probability": 0.97
}If JEV is unavailable, or a workspace turned AI processing off, the pattern match still returns the code, and jev_probability stays null.
#Nothing a model says is irreversible
No answer from JEV deletes, sends, forwards or blocks anything. Labels are only ever added, and labels a user set are never removed. The worst a wrong classification can do is put a misleading label on a message, which a person or your own code can overrule. Deciding what happens to a message flagged ai:phishing stays with you.
#When it fails
The SDK’s own retries are turned off, so the queue owns retry policy and there is one place to look when something goes wrong. Errors that can’t get better with time, such as a rejected request, stop at once instead of burning the remaining attempts; timeouts and rate limits are retried. After the last attempt we record ai.enrichment_error and stop. The message and its verification result stay exactly as they were, and no message.enriched event is sent, so your code never receives a half-classified message.
#The privacy switch
Classification sends part of each message to an outside model, so each workspace decides whether that happens. categorize, the default, allows JEV. full also allows features that send whole threads to a language model, such as reply drafts. off sends nothing: no categories, no risk scores, no ai:* labels. Codes are still found by pattern matching, and SPF, DKIM, DMARC and extracted_text keep working, because none of them use a model.
Every code path that sends mail content to a model has to pass through one check of that setting, and a unit test fails if a module uses an AI client without importing it. The setting is read again when the job runs, so turning it off also stops messages that were already queued.
#What we don’t claim
JEV has been checked on English. Mail in other languages is classified too, but deserves less trust until we have measured it. Classifiers miss new phrasings. A message from a real but compromised account passes every authentication check. The labels lower the odds that a hostile email reaches your model unnoticed; they don’t make it safe to give an agent unlimited power over a public inbox. Scoped keys, recipient allow lists and a person on irreversible actions still matter.
#In short
- Store and deliver first; classify after, on a queue.
- One call per message, one retry unit, one
enriched_at. - Send the new text only, and authentication results only from your own server.
- Always offer “other”, and ignore low-confidence categories.
- Let the model answer and your code decide, with named thresholds and stored probabilities.
- Never let a model’s answer do anything that can’t be undone.
The fields and labels are documented under categories and risk flags and AI processing and privacy.