The first batch is almost always too big.
A team decides physical mail is worth trying, exports the full account list from the CRM, and puts four hundred notes in the mail in week one. Six weeks later, two deals have closed. Nobody can say whether either one touched a note. The rep who addressed all four hundred has quietly stopped doing it. The program gets written up as inconclusive and shelved.
The order of operations killed that program, not the channel. We have already made the case that the data on handwritten mail holds up, so this is the other half: what you actually do on Monday morning if you believe it.
Why run a channel that costs more per touch?
Because the cheap channel stopped converting, and the numbers on that are worse than most teams realize.
Belkins measured 7.5 million cold emails sent across 2025 and reported an average reply rate of 0.45 percent. The denominator is the interesting part. Belkins changed how it counts this year: previous studies reported replies as a share of people who opened, and this one reports replies as a share of total sends. Those two methods can describe the identical campaign and produce a 5 percent reply rate or a 0.45 percent reply rate depending on which you pick. The stricter number is the one your finance team is implicitly using when they divide pipeline by list size.
The trend inside that year is worse than the average. The first half of 2025 averaged 0.50 percent and the second half dropped to 0.40 percent, a 20 percent decline in twelve months, bottoming out at 0.35 percent in December.
Direct mail, over the same kind of measurement, averages a 4.4 percent response rate in ANA’s Response Rate Report. Physical media also requires 21 percent less effort to understand and produces 70 percent higher brand recall than digital, according to Canada Post’s neuromarketing research. Budgets have followed: Lob’s State of Direct Mail found 82 percent of respondents increasing direct mail spend in 2024, up from 58 percent in 2023.
That is the case for the channel. The rest of this is the plan.
Who should get the first batch?
Segment selection comes before volume, and the Belkins data tells you exactly where to point.
Reply rates in that study fall on an almost linear gradient by company size. Companies with fewer than 10 employees reply to cold email at 0.72 percent. Enterprises with more than 10,000 employees reply at 0.22 percent, roughly a third of the rate. Seniority splits the same way: founders and owners reply at 0.57 percent, C-level executives at 0.42 percent, and VPs at 0.32 percent, the hardest group in the dataset to reach.
Invert that. The first batch should go to the segment where email is measurably worst, because that is where a physical touch has the most room to add. A VP at a 12,000-person company is the single least responsive contact type in the study. That person is also the one your reps have been emailing seven times.
Two versions of what this looks like in practice.
An enterprise field team. Sixty named accounts, one note each, addressed to VP-and-above contacts at companies over 5,000 employees where a rep has already had a real conversation. That cohort is a deliberate bet rather than a segment the data hands you: Belkins measures company size and seniority separately, so the two lowest-reply groups are known and their overlap is inferred. Not a cold list. Sixty is small enough that one person can own it and every reply can be traced back to a specific note.
A mid-market team with a stalled pipeline. Pull every opportunity that reached proposal stage and then went quiet for more than 21 days. That list is usually longer than anyone expects, and it is a better first cohort than any prospecting list because the recipients already know who you are. The mechanics of restarting those conversations are covered in what to do when a deal goes dark.
Both cohorts share the trait that matters: the recipient can tell why they specifically received the note. A note that could have gone to anyone performs like a printed mailer.
How many notes should you send in the first month?
Fewer than you want to. Somewhere between 50 and 75 is right for a first cohort, and the reason is attribution.
At a 4.4 percent response rate, 50 notes produce about two responses. Two responses prove nothing statistically, and you should not pretend otherwise in the readout. What 50 notes do prove is operational: that someone owned the list, that the notes went out inside the window they were supposed to, that the messages referenced something real, and that when a reply arrived, somebody knew which note caused it. Those four things are what fail in month one. Response rate is a month-four question.
The four-hundred-note version fails on all four at once, which is why it produces a shrug instead of a decision.
When you are ready to actually send, the constraint most teams hit is that writing 60 genuinely specific notes by hand costs a rep the better part of two days. That is the point where the choice is between shrinking the program and automating the handwriting rather than the message. See what it costs to run the first batch before you scope the cohort, because the per-note number changes how large a first cohort you can defend.
What does a handwritten note program cost?
Budget on cost per response, not cost per piece. Per piece, physical mail loses to email by a factor of a thousand and always will. Skip that comparison.
Work an example. Assume 5 dollars all-in per note, which covers materials, postage, and fulfillment for most B2B programs. Sixty notes is 300 dollars. At the 4.4 percent ANA response rate, that is roughly 2.6 responses, or about 115 dollars per response.
Run the same 300 dollars through cold email and the arithmetic is not the win it looks like. Email’s marginal cost per send rounds to zero, so cost per response is dominated by list, tooling, and rep time rather than postage. The honest comparison is not 5 dollars against a fraction of a cent. It is 115 dollars per response against what your team currently spends in salary and tooling to generate one reply at 0.45 percent. For most B2B teams carrying a loaded rep cost, running that number the first time is uncomfortable.
We walk through that framing in detail in how to get a sales VP to approve a direct mail budget. Bring cost per response. Leave the 42-to-1 ROI statistics at home, because a VP who has seen one marketing deck has seen that number and discounted it.
What do you measure in the first 90 days?
Measure these in order. Response rate does not join the list until month three.
Did the notes go out on schedule? Track the gap between the trigger event and the postmark. A note referencing a demo that arrives 19 days later reads as an afterthought. Under seven days is the target, and missing it is the most common way a first program quietly dies.
Did each note reference something checkable? Sample ten notes a month and read them. If a note would work equally well sent to a different person, it is a mailer. This check takes fifteen minutes a month and catches more damage than any other measure here.
Can you attribute a reply? Decide before the first send how a response gets logged. A CRM field, a task, an activity type, anything, as long as it exists before the replies arrive. Retroactive attribution is how programs end up unable to defend themselves at the 90-day review.
Only after those three hold do response rate and influenced pipeline mean anything. And notes should sit inside an existing sequence rather than replacing it, which is why the touch structure in the five-touch B2B follow-up sequence matters more than the channel choice. The gap between physical and digital response rates exists partly because so few teams send physical mail at all, and that advantage is a function of scarcity, not of paper.
What breaks when you skip a step?
Each of the failures below shows up in a different month, which is part of why they are hard to catch from inside the program.
Sending before segmenting produces volume nobody can attribute. This is the four-hundred-note opening. The fix costs nothing: pick 60 accounts with a reason attached to each one.
Scaling before the 90-day checkpoint locks in whatever was broken at 60 notes and multiplies it. If notes were arriving 19 days late in month two, they will arrive 19 days late at 600 notes, at ten times the cost.
Automating the message instead of the handwriting is the failure that looks most like success. A generated paragraph in real handwriting still reads as generated, because the recipient is checking whether the sender knew anything specific about them, not whether the letters are round. Handwriting activates broader neural networks than typed text, according to EEG research from the Norwegian University of Science and Technology, but that effect is about how the message is received, not about whether it was worth sending. The words still have to be yours.
The first 90 days, week by week
Weeks 1 and 2. Pick the cohort. Sixty contacts, each with a written reason. Decide the trigger event, usually a demo, a proposal, a renewal, or a stall past 21 days. Create the CRM field that will hold the attribution. Nothing gets mailed in this window.
Weeks 3 through 6. First batch goes out, staged rather than all at once, roughly 15 notes a week. Staging matters because it lets you fix the message in week four instead of discovering the problem after all 60 have shipped.
Weeks 7 through 10. Read ten notes. Check the trigger-to-postmark gap. Fix whichever of the three measures is failing, and expect one of them to be failing.
Weeks 11 and 12. The readout. Report the operational measures first and the response numbers second, framed as directional. Then make one decision: double the cohort or fix the process. Not both.
The takeaway
The order is segment, then send, then measure, then scale, and every failed program reverses two of those. Holding the order buys you one specific thing: a cohort of 60 gives you something to defend in week 12. A batch of 400 gives you an anecdote.
Cold email reply rates fell 20 percent inside a single year in the Belkins data. Whatever margin physical mail has right now, it comes from being difficult enough that most teams do not bother. That advantage rewards the team that gets the operation right at 60 notes before anyone else gets it right at 600.
FAQ
How do you start a handwritten note program for a business?
Start with 50 to 75 contacts rather than a full list. Pick a cohort where you can attach a specific reason to every name, usually accounts a rep has already spoken with. Define the trigger event that causes a note to be sent, create the CRM field that will record replies before the first note goes out, and stage the first batch over three to four weeks so the message can be corrected mid-flight. Measure the operational side first: whether notes went out within seven days of the trigger, whether each note referenced something checkable about the recipient, and whether replies can be traced to a specific note. Response rate becomes a meaningful metric around month three.
How much does a handwritten note program cost for a B2B team?
Budget on cost per response rather than cost per piece, because per piece physical mail will always lose to email. At roughly 5 dollars all-in per note and the 4.4 percent response rate ANA reports for direct mail, a 60-note cohort costs about 300 dollars and produces two to three responses, near 115 dollars per response. The relevant comparison is not against email’s per-send cost, which rounds to zero, but against the loaded rep and tooling cost of generating one reply at cold email’s 0.45 percent average reply rate.
Who should receive the first handwritten notes?
The segment where email performs worst. Belkins found cold email reply rates falling on a near-linear gradient by company size, from 0.72 percent at companies under 10 employees down to 0.22 percent at enterprises over 10,000, and VPs replying at just 0.32 percent, the lowest of any seniority group. Belkins reports those two gradients separately and never crosses them, so treat the overlap as a working assumption rather than a finding: the segments email reaches worst are large companies and senior titles, and a list that is both is where a physical touch has the most room to win. A stalled-pipeline list, meaning opportunities that reached proposal stage and went quiet for more than 21 days, is usually a better first cohort than any prospecting list, because those recipients already know who you are.