Every team we talk to has the same shape of story. Two years ago the first-touch reply rate was somewhere around 22 percent. Now it is nine. The sequences did not get worse. The volume went up, everyone's volume went up, and the inbox on the other end did the arithmetic.
The standard response is to send more, which is the one response guaranteed to make the underlying condition worse.
The arithmetic of the volume trap
Consider a candidate pool of a hundred senior engineers in a specific niche, and thirty companies hiring into that niche. If each company sends one sequence of four touches per quarter, each engineer receives 120 messages a quarter. That is more than one per working day, all of them about jobs, most of them opening with a compliment about a GitHub profile.
Reply rate is not a property of your email. It is a property of your email divided by everything else in the inbox that quarter. Doubling your send doubles the denominator for everyone including yourself, and the equilibrium it moves toward is one where nobody replies to anybody and the only winners are the email providers.
We are contributing to this. So is every other tool in the category. Pretending otherwise would be strange, and the honest position is that a sequencing product should be measured on replies per thousand messages, not messages per hour.
What actually moves a reply
We looked at 2.4 million first-touch messages sent through Arbi over eighteen months and modelled reply against everything we could measure. Three things came out with an effect size worth caring about. Most of what gets written about outreach did not.
Specificity that could not be templated
Not personalisation tokens. The model does not care that you interpolated a first name, and neither does the recipient. What moves the number is a sentence that could only have been written about that person: a reference to a specific project, a talk, a design decision, an unusual path between two roles.
Messages containing at least one such sentence replied at 2.7x the rate of messages without one. The effect held after controlling for sender, seniority, and company.
A first message that asks for one thing
Messages with a single explicit ask outperformed messages with two or more by 61 percent. The common failure is a message that asks for a reply, a call, a CV, and a referral, which reads as a form rather than a conversation and gets processed accordingly.
The best-performing single ask was not "are you open to a call". It was a closed question about the person's situation that could be answered in a sentence. Low cost to answer, and answering it starts a thread.
Length, but not the way people think
Short messages do better up to a point and then stop. The curve bottoms out somewhere around 60 words and rises again past 220, and the long tail is real: detailed messages about a specific technical problem the company is facing reply well. What performs badly is the middle, where 120 words of generalities are long enough to demand attention and short enough to say nothing.
Timing matters, and it is not a hack
One finding did surprise us. Reply rate against the recipient's own tenure is far from flat. Messages landing between month 20 and month 34 in a role reply at roughly double the rate of messages landing in the first year.
This is not a trick to schedule around. It is a reason to build a pipeline you can wait with. The candidate who says no in March because she started in January is a strong yes eighteen months later, and the only teams that capture that are the ones where "no, not now" writes a date into a system rather than closing a tab.
Most sequencing tools are built for a campaign that ends. The valuable thing is the one that does not.
The uncomfortable conclusion
If specificity is the thing that works, and specificity is expensive, then the honest version of outreach at scale is not "send more, personalised automatically". It is: send fewer, to a list that has been narrowed properly, with something real in the first paragraph.
That puts the weight back on the stage before outreach. A sequence sent to 400 loosely-matched people at 4 percent yields 16 replies and burns the list. The same effort spent narrowing to 80 genuinely strong matches, with a real sentence each, yields more replies, more conversations worth having, and a pool that will still take your email next year.
The volume trap is not primarily a writing problem. It is a targeting problem that shows up in the writing, because you cannot say anything specific about a person you had no real reason to contact.




