
AI answering service for HVAC: The 30-day pilot scorecard
Baseline the line you already have
A pilot without a baseline produces a number with nothing to compare it to. Before any new service takes a call, pull 30 days of after-hours data from your phone system and your CRM: How many calls rang the after-hours line, how many reached a person, how many became a booked job, and how many of those jobs were kept. The first two come from the phone system. The last two take a CRM report matched by caller phone number, and that match is the point.
ServiceTitan's 2022 booking-rate report across 3,000 plus businesses found HVAC shops book 38% of calls on average, the top bracket books 59%, and rates drop after 6 PM. Your after-hours baseline is probably lower. The pilot is not judged against a vendor's promise. It is judged against the line you were running last month.
Score booked-and-kept jobs, nothing else
"Calls handled" is the metric each service will offer you, and it is the one to refuse. A message taken is handled. A caller told to call back at 8 AM is handled. Neither is revenue. The scorecard has one primary number: Jobs booked on after-hours calls that were still on the schedule when the tech arrived. A second tier explains it: Answer rate within three rings, share of calls reaching a disposition, transfer rate to a human, and bookings with the right job type and arrival window.
Metric | Baseline (prior 30 days) | Pilot (30 days) | Source |
|---|---|---|---|
After-hours calls received | Phone system | ||
Calls that reached a disposition | Disposition export | ||
Jobs booked | CRM, matched by phone number | ||
Jobs kept (tech arrived, work done) | CRM | ||
Same-night rolls the on-call tech agreed were needed | On-call tech log |
Demand recordings and a per-call disposition export
Ask for two deliverables in writing before the pilot starts. All calls recorded and downloadable, and a weekly export with one row per call: Caller number, time, disposition (booked, message, transferred, hung up, self-resolved), job type if booked, and the appointment it created. If a service cannot produce that file, the pilot cannot be scored.
Listen to a sample weekly, not at the end. Twenty calls chosen at random plus each call marked transferred or hung up. You are listening for three things: Did the caller get an answer, did the service collect what your dispatcher would have collected, and would you put your company name on that call.
Place your own test calls for the failure cases
Thirty days of real traffic may not include the calls that matter most, so create them. Have people outside the office call the pilot line from personal phones, once on a weekday evening and once on a weekend.
Scenario | Pass looks like |
|---|---|
Caller reports a gas smell | Told to leave the house and call the utility or 911; no booking attempted |
Existing member calls about a covered tune-up | Recognized as a customer, booked under the membership, not quoted a diagnostic fee |
Spanish-speaking caller | Full intake completed in Spanish and booked, not asked to call back tomorrow |
Angry customer whose tech no-showed | Rescheduled to a firm window, apology recorded, flagged for a manager callback |
No heat at 1 AM with an infant in the house | Rolled to the on-call tech with the after-hours rate disclosed first |
Price shopper asking for a replacement quote | Booked a free estimate; no price given over the phone |
The service that passes the gas smell call and the Spanish caller is the one that can be trusted with the calls you never hear.
Write the kill criteria and contract terms first
Decide before day one what ends the pilot early: A gas-smell test call handled wrong, a week without the disposition export, or a booked-and-kept rate below baseline at day 30. Put it in the agreement. On terms: A 30-day exit with no penalty, no auto-renew during the pilot, your recordings and disposition file are yours to keep, the phone number stays yours, and the service confirms in writing which system it books into (ServiceTitan, Salesforce, or your CRM) and whether that booking is live or a message someone re-keys in the morning. On price, stay qualitative until the scorecard is in, and ask how the invoice behaves in your busiest week.
Compare against the two options you already have
The pilot answers one question: Does this beat what you run today. Against the traditional answering service, score the same 30 days of its messages: How many became jobs, how long until a CSR called back, how many callers had already booked elsewhere. Against the in-house baseline, be honest about what the after-hours CSR or the on-call tech books, and what the rotation costs in overtime and turnover. An AI intake agent that books into your schedule live should beat the message-taker on booked jobs by a wide margin and approach the daytime CSR's booking rate.
Reading the scorecard at day 30
Put the two columns side by side. If booked-and-kept jobs went up and the test calls passed, extend to 90 days on the same terms. If the numbers are flat but the recordings are clean, check whether the schedule had open slots to book into. If the recordings are bad, stop, regardless of the numbers. Thirty days is long enough to know, and the scorecard is what keeps the decision yours.
Is AI better than a live answering service?
Judge it on booked-and-kept jobs, not on how the call sounds. A live answering service takes accurate messages but rarely books into the schedule, so most after-hours callers wait for a morning callback. An AI intake service that books live can beat that on bookings, but only a 30-day pilot with recordings, a per-call disposition export, and your own test calls will show whether it does for your shop.
What booking rate should an answering service deliver?
Start from your own numbers. ServiceTitan's 2022 booking-rate report found HVAC shops book 38% of calls on average and the top bracket books 59%, with rates dropping after 6 PM. Measure your current after-hours line for 30 days, then expect any replacement to beat that baseline on jobs booked and kept. A service that only reports calls handled is not reporting a booking rate at all.
Can an AI answering service dispatch emergency calls?
Some can page the on-call technician when the caller's answers meet the criteria you set, such as no heat with an infant in the house or an indoor temperature under 50 degrees. Test it yourself during the pilot with scripted calls, including a gas smell, which should end with the caller told to leave the house and call the utility or 911, not with a booking.
How do you test an answering service before signing?
Baseline your current after-hours line for 30 days first, then run the new service for 30 days with the same measurements: Calls received, dispositions, jobs booked, jobs kept. Require call recordings and a per-call export, listen to a sample weekly, and place your own scripted test calls for the hard cases. Write the kill criteria and a penalty-free exit into the agreement before the first call.
Does the answering service book into your scheduling software?
Ask, and get the answer in writing. Many services take a message and email it, and someone in your office re-keys the appointment the next morning, which is where callers are lost. A service that books live into ServiceTitan, Salesforce, or your CRM should be able to show you the appointment it created for each booked call in its disposition export, matched by the caller's phone number.

About Revin
Run a 30-day pilot on your after-hours line
Revin answers every HVAC call after hours, books into ServiceTitan live, and hands you the recordings and dispositions to score it.







