Posted in Guardrails · 8 min read

AI reply assistants: what to automate, and what to keep in your own hands

Drafting help earns its place on the routine reply and loses money on the important one. Here is the test that sorts them — and the categories worth keeping in your own hands permanently.

Farhad

Founder, Reply Pilots ·

An AI chat interface glowing on a dark screen

In short

AI drafting is worth using where a reply is routine, low-stakes and about information you have already written down, and worth avoiding where the message is the relationship, the facts are contested, or being slightly wrong is expensive. The sorting test is four questions — is it routine, are the facts already in your context, what does a wrong version cost, and would the recipient mind knowing it was drafted. Volume is what makes the distinction matter, because a tool that is right ninety percent of the time produces one bad message in every ten, and at ten messages a day that is one a day.

Key takeaways

  • Drafting helps most where the reply is routine and the facts are already written down somewhere.
  • The cost of a wrong draft is not evenly distributed. One reply in a complaint thread outweighs fifty routine ones.
  • If the recipient would mind knowing it was drafted, that is the answer — not a reason to hide it.
  • Volume converts a small error rate into a daily occurrence. Ninety percent accurate means one bad message per ten.
  • Reading every draft is not a temporary precaution. It is the design.

The reply that causes a problem is almost never the one you worried about. It is the ordinary one — the fourth message in an unremarkable thread, sent between two other things, that agreed to something you did not mean to agree to. You did not review it carefully because there was nothing about it that seemed to need reviewing.

That is the actual risk profile of AI drafting, and it is not the one the argument usually has. The public debate is about whether using AI to write is honest, which is mostly a question about personal messages. The operational question is narrower and more useful: which of the messages you send this week benefit from a draft, and which ones cost you something if you use one.

The two categories are separable, and the line is not where people expect. It has little to do with how important the recipient is and a lot to do with whether the facts needed to answer are already written down somewhere.

This article draws the line. Guardrails are what stop a draft claiming things you cannot deliver; this is the prior question of whether to draft the message at all.

What does drafting actually do well?

Three things, and they are narrower than the marketing suggests.

It removes the blank box. The expensive part of a reply is usually the first sentence, and a draft to react to is faster than a page to start. This holds even when you rewrite most of it — the editing mode is cheaper than the composing mode.

It reads what is on the screen. For a thread that has run to forty comments, a draft generated from the visible conversation knows what has already been said, which is something you would otherwise have to reconstruct by scrolling. That is where the time actually goes on long threads.

It keeps a consistent register at volume. The tenth reply of the morning is worse than the first — shorter, blunter, more likely to skip the thing that needed saying. A draft does not degrade across a session, and consistency is most of what people mean by professional.

None of those is "it writes for you". They are all versions of removing friction from work you were going to do anyway.

Where does it quietly cost you?

Four categories, and they share a property: the cost is invisible at the moment of sending.

The message that is the relationship. A long-standing client asks how you have been. The reply that matters is the one that shows you remember something. A drafted reply will be warm, appropriate and empty, and the emptiness is the message.

Anything where being slightly wrong is expensive. Money owed, a disputed job, a safety question, anything with a legal or medical dimension. A draft is confident by default, and confidence in a situation you have not fully established is exactly the wrong register.

A complaint that has become a thread. The draft will be conciliatory, because conciliatory is the safe default, and conciliatory before you know what happened concedes a position you may need. It also reads as processed to the person complaining, which escalates.

Anything that commits an outcome. Dates, prices, capacity, what a colleague will do. Guardrails stop most of it, and the residual risk is that the draft is fluent enough that you skim it.

The cost of a bad draft is not spread evenly across your messages. One reply in a complaint thread outweighs fifty routine ones, and it is the one you were least likely to read carefully.

What is the four-question test?

Run these in order. A no on question two or a high answer on question three means write it yourself.

#QuestionDraft itWrite it yourself
1Is this reply routine for me?I send versions of this weeklyI have not had this one before
2Are the facts already in my written context?Services, process, how I workCapacity, stock, what we agreed last month
3What does a wrong version cost?An awkward correctionMoney, a client, a public retraction
4Would the recipient mind knowing?No — it is an information replyYes — the attention was the point

Question two is the one that does most of the sorting, and it is the one people skip. A drafting tool knows what you told it. It does not know your diary, your stock, or the phone call you had on Tuesday — so any reply whose correctness depends on those facts is a reply it can only guess at, and it will guess fluently.

Question four is the honesty test, and it is worth taking at face value rather than arguing with. If the answer is that they would mind, the message was one where the personal attention was the substance. That is not a disclosure problem to be managed; it is a signal about which pile the message belongs in.

Why does volume change the calculation?

Because a per-message error rate becomes a schedule.

A tool that produces an acceptable draft nine times out of ten sounds good. At ten replies a day it means one unacceptable draft per day — and unacceptable does not mean obviously broken, it means subtly over-committed or slightly off-key, which is the kind that gets sent.

This is why "read every draft" is not a temporary precaution for the first week. It is the permanent design, and it is the reason a product that inserts drafts into a composer is a different category from one that sends them. The review step is not friction that better models will eventually remove. It is the thing standing between a small error rate and a daily incident.

It is also why the volume you are aiming for should be honest. Ten well-chosen comments a day is a sustainable routine. Fifty is a number that only works if you stop reading, and stopping reading is where the arithmetic above turns against you.

How do you sort your own messages?

Do it empirically rather than in principle. Open your last twenty sent messages and put each one in a pile using the four questions.

Most people find the same shape: about three quarters are routine information replies where a draft saves real time, and a small identifiable set — usually two or three — are the ones they would want to have written personally. The useful part is that the second set is identifiable. It is not a vague risk spread across everything; it is a short list of recognisable situations.

Write that list down. Complaints. Anything about money. The three clients you have known for years. Safety questions. Once it exists, the sorting stops being a judgement call made under time pressure and becomes a rule you already made.

Sort once, in advance, when you are not in a hurry. The decision made while a notification is waiting is the one that goes wrong.

What about auto-reply?

Worth addressing directly, because it is the obvious next step and it is a different product.

An unreviewed reply posted publicly is a published commitment with nothing between the model and your audience. Every mitigation discussed here — guardrails, reading the draft, sorting the message — assumes a person in the loop. Remove the person and the residual failures are no longer awkward, they are public and permanent, and they arrive at the moment you are least able to respond.

There is also a platform dimension. Rules across the major networks target automated action rather than assisted writing, and the distinction is exactly whether something is sent on your behalf. A tool that drafts into your composer sits on one side of that line. A tool that replies for you sits on the other, and the account at risk is yours.

This is why Reply Pilots has no auto-reply, no scheduler and no bulk send. Not as a roadmap gap — as the boundary the product is drawn around.

What should you deliberately keep in your own hands?

The sorting decision itself, permanently. Everything else here can be made faster; deciding which pile a message belongs in is the judgement that makes the rest safe.

The specific detail in any reply that is doing real work. A drafting tool has your context, not your experience, and the sentence that could only have come from having done the job is the one that makes a reply worth reading.

The first message in anything important. Openings set the register for everything that follows, and the twelve routine replies after it are where the time was actually going anyway.

And the ones on your written list. That list exists precisely so this is not decided in the moment.

What does Reply Pilots do here, and what does it not?

Reply Pilots is built for the drafting side of that line and nothing else. It drafts in the composer on Facebook, Instagram, LinkedIn, X, Reddit and Gmail, from the post or conversation visible on your screen, grounded in the business context you wrote and inside your guardrails. It summarises long threads so the reading step is shorter, and it rewrites your own rough draft when you would rather start from your words than its own.

What it does not do — and this is the part worth reading twice: it never posts, sends, likes, follows or messages on your behalf. There is no auto-reply, no scheduling, no bulk send, and no outbound sequencing. It does not scrape or scroll a page. It reads what is already rendered on a page you opened yourself, and it holds no AI provider key — the requests are made by our server, or by your own provider if you have added your own key.

The honest limitation: it knows what you wrote down. Not your diary, not your stock, not the call you had on Tuesday. Every reply whose correctness depends on those stays yours, and no amount of configuration changes that.

Your next step

Open your sent messages and take the last twenty. Sort them with the four questions and write down the categories that landed in the second pile — the ones you would want to have written yourself.

That list is the useful artefact. It is short, it is specific to your business, and having it written down before the next busy morning is the difference between a drafting tool that saves you an hour a day and one that eventually sends something you have to apologise for.

Related reading

See how Reply Pilots works for the product this article is about, end to end.

Frequently asked questions

Is it dishonest to use AI to draft replies?

Not when you read, edit and send it yourself, any more than using a template or a saved reply is dishonest — the message is still your commitment. It becomes a problem when the message is one where the personal attention itself is the point, such as condolences or a note to a long-standing client. The test is whether the recipient would feel differently if they knew.

What kinds of messages should never be drafted?

Anything where being slightly wrong is expensive or public, and anything where the message is the relationship. In practice that means complaints, safety questions, legal or medical situations, anything about money owed, and personal messages to people who know you well. Everything else is usually fine with a read-through.

How accurate are AI drafts, really?

Accurate enough on routine, well-scoped replies that most people stop editing them heavily after a week, and unreliable on anything requiring facts that were never written down — capacity, stock, what you agreed with someone last month. The useful mental model is a fast assistant who has read your briefing document and nothing else.

Does using AI to reply violate platform rules?

Drafting a message that you then read and send yourself is not what platform rules target. What they target is automated action — bulk sending, auto-replying, scripted engagement — which is why the line between a drafting tool and an automation tool matters. If a tool sends on your behalf, that is the part worth checking against the platform's rules.

Can people tell a reply was AI-drafted?

Often, and increasingly so, but the tells are mostly removable — filler adjectives, even sentence rhythm, an absence of anything specific. What people actually detect is the absence of a concrete detail rather than the presence of AI, which is why adding one real specific before sending does more than any amount of rephrasing.

Should I tell people I use AI to draft replies?

There is no obligation to disclose a drafting aid you review and send yourself, and most businesses do not. What matters more is not using it in the categories where disclosure would change how the message lands — which is a good sign you should have written that one yourself.

What about auto-replying to comments overnight?

That is a different product category and a meaningfully different risk. An unreviewed public reply is a published commitment with nobody between the model and your audience, and it is the pattern most likely to produce the incident that ends up screenshotted. Reply Pilots does not do it, deliberately.

How do I decide for my own business?

Run the four-question test on the last twenty messages you actually sent. The split is usually clearer than expected, and it tends to show that the bulk of your volume is genuinely routine while the messages you worry about are a small, identifiable set.

Stop reading, start replying

Your next comment is one click away.

Reply Pilots reads the post and everything already said under it, then drafts a reply in your voice — right in the box you were already about to type into. You read it, tweak a word if you need to, and send it yourself.

Free to start · You approve every reply · It never posts for you