OttoflowRun Free Website Audit

Case studies

Real builds, measured results

Five AI agent engagements, end to end: the baseline that was measured before anything was built, what was built, and what the numbers said afterward. Client names are withheld by agreement.

Case study · Voice agent

The after-hours calls a contractor never heard

A residential HVAC and plumbing contractor — $12M revenue, 34 employees, nine trucks.

The problem, measured first

The office answered the phone from 8am to 5pm on weekdays. Everything else went to voicemail. Instead of guessing what that cost, we pulled six months of call records: roughly 90 calls a month arrived outside office hours, about 22 of them genuine new-job inquiries. Next-morning callbacks recovered around five. The rest had already reached the next contractor on Google.

At the company's average job value and margin, that gap was worth roughly $18,400 a month in gross profit — going to voicemail.

What we built

A voice agent that answers every out-of-hours call, qualifies the job, checks live technician availability, offers two appointment slots, confirms by SMS, and writes the job into the field service system before anyone wakes up. It was scripted from transcripts of the office's own best coordinator, so it sounds like the company — not like a chatbot.

  • Any mention of a gas odor, flooding, or no heat with an infant or elderly resident routes straight to the on-call technician's phone.
  • It never quotes a price beyond the published diagnostic call-out fee.
  • Addresses outside the service area get a polite decline and a texted referral.

How it went live

Before launch, the agent replayed thirty real recorded calls and its booking decisions were compared against what the office actually did — then the client's own staff spent a round of live calls actively trying to break it. It ran a week in shadow mode with every transcript emailed to the owner, then weekends only, then all out-of-hours. Nothing was trusted before it was measured.

After-hours jobs booked

5 / month

15–17 / month

Recovered gross profit

$0

~$13,000 / month

Annualized value

—

~$156,000

Running cost

—

$60–80 / month

The change, drawn to scale

After-hours jobs booked per month

Before
5
After
15–17

Recovered gross profit per month

Before
$0
After
~$13,000

What keeps it honest: availability is always a live read — never a cached schedule — so the agent can't book work the trucks can't service. Some callers do hang up on an AI voice. The honest comparison is against voicemail, not against a human.

Case study · Operations agents

600 hours of check calls, and quotes that arrived too late

A freight brokerage — eighteen brokers, roughly 2,400 loads a month.

The problem, measured first

Two workflows consumed the working day. Brokers phoned and emailed carriers three to five times per load for status updates — about 600 hours a month spent chasing status. Meanwhile, inbound spot-quote requests took an average of 45 minutes to answer, in a market where the first credible quote usually wins the load.

What we built

A track-and-trace agent that contacts carriers at scheduled checkpoints — SMS first, a voice call if there's no reply — parses whatever comes back, updates the transport management system, and involves a broker only when something is off-plan. And a quoting agent that reads inbound requests, extracts the lane and requirements, prices against live market data and the firm's own margin rules, and returns a quote within minutes. The language model extracts and drafts; deterministic code does the pricing.

  • A load projected more than an hour late, silence after two attempts, or any mention of a breakdown alerts the broker immediately — with full context.
  • It never commits capacity that hasn't been booked.
  • Quotes outside the approved margin band, for a new customer, or above a value threshold always go to a human for one-click approval.

How it went live

Five hundred historical quote emails were replayed and extraction was measured field by field before a single live quote went out. A month of past loads ran through the tracking logic to confirm it caught every exception the brokers had caught. Tracking launched on one broker's book first; quoting ran in approval-required mode for a month before any threshold was loosened.

Check-call hours

~600 / month

~90 / month

Quote response time

45 minutes

Under 5 minutes

Quote volume

Baseline

+40%, same headcount

Win rate

Baseline

+8–12%

The change, drawn to scale

Hours spent chasing load status per month

Before
~600 h
After
~90 h

Time to answer a spot-quote request

Before
45 min
After
< 5 min

Annualized value measured at over $700,000. What keeps it honest: approval bands stay tight permanently on non-standard loads, and the exception alerts belong to the brokers — not to a manager's dashboard.

Case study · Back-office agents

The 190 hours a month spent chasing documents

An accounting and bookkeeping practice — fourteen staff, 240 monthly bookkeeping clients.

The problem, measured first

Every month, staff emailed and called clients for bank statements, receipts and missing invoices — roughly 190 hours a month, and the work everyone deferred longest. Behind it, around 40,000 transactions a month needed categorizing, with about 15% requiring a manual touch or a query back to the client. Month-end close averaged eleven business days.

What we built

A document chaser that knows precisely what is outstanding for each client, chases on an escalating schedule across email and SMS, accepts uploads by reply — photographed receipts included — validates them and files them to the right client and period. And a reconciliation agent that categorizes transactions against each client's own history, never a generic chart of accounts, auto-posts only above a measured confidence threshold, and drafts plain-English queries for anything genuinely ambiguous.

  • Chasing stops entirely after the fourth touch and hands off to a human — no client relationship is worth a deadline.
  • It never posts to a tax-sensitive account without human review, and never alters a reconciled period.
  • No client query is sent without staff approval.

How it went live

Categorization was replayed against a full quarter of already-reconciled transactions, and the auto-post threshold was set where measured accuracy exceeded 98% — with a senior accountant reviewing 200 sample categorizations and every drafted query. The chaser launched on thirty clients with staff approving every message for the first week; reconciliation ran a full close cycle in review-everything mode before any threshold moved.

Document chase hours

~190 / month

~35 / month

Docs in by day 10

45%

80%

Manual-touch transactions

15%

4%

Month-end close

11 days

6 days

The change, drawn to scale

Document chase hours per month

Before
~190 h
After
~35 h

Clients with documents in by day 10

Before
45%
After
80%

Business days to close the month

Before
11 days
After
6 days

Annualized value measured at over $190,000 — plus capacity for roughly forty more clients without hiring. What keeps it honest: the auto-post threshold stays conservative permanently, because a wrong posting surfaced in an audit costs far more than the minutes it saved, and per-client data isolation is verified, not assumed.

Case study · Intake agents

530 hours a month screening cases the firm would never take

A personal injury law firm — six attorneys, roughly $4M in annual fees.

The problem, measured first

Paid advertising delivered roughly 400 inquiries a month. About 45 were viable. Paralegals spent around 90 minutes screening each one — approximately 530 hours a month, most of it spent ruling out the 355 cases the firm would never take. Worse, viable leads waited an average of four hours for a response, in a practice area where the firm that responds first usually signs the client.

What we built

An intake screening agent that engages every inquiry within sixty seconds by SMS, gathers the facts conversationally, scores viability against the firm's own criteria — with a written reason for every factor, not a bare number — and books qualified prospects directly into an attorney's calendar. Alongside it, a case summary agent that turns the intake record and uploaded documents into a one-page brief before the consultation, with every factual claim cited back to its source document. No citation, no claim.

  • It never gives legal advice, never says whether someone has a case, and never estimates value — every conversation opens by stating that no attorney-client relationship is formed.
  • Filing deadlines are computed by deterministic code; the language model is never permitted to do that arithmetic.
  • An expiring deadline, a near-threshold score, or a distressed caller escalates to a human immediately.

How it went live

Two hundred historical inquiries — one hundred accepted, one hundred declined, each with the recorded reason — became the gold-standard test set. The only score that truly mattered was false negatives: viable cases wrongly declined. The target was zero, and the threshold is deliberately tuned to over-refer to humans. The managing attorney reviewed and signed off every line of the script, two attorneys verified every citation in twenty sample briefs, and the agent ran two weeks in shadow mode with human review of every decision before going live on a single advertising source, then all inbound.

Response time

4 hours

Under 60 seconds

Screening hours

~530 / month

~120 / month

Signed cases

Baseline

+15–25%

Consultation prep

~40 minutes

~10 minutes

The change, drawn to scale

Paralegal screening hours per month

Before
~530 h
After
~120 h

Response time to a new inquiry

Before
4 hours
After
< 60 sec

Attorney prep per consultation

Before
~40 min
After
~10 min

Annualized value measured at over $250,000. What keeps it honest: a distressed caller reaching a bot is the failure mode that matters most, so sentiment detection with immediate human handoff is mandatory — and transcripts are reviewed on a schedule, because without tight output constraints the language drifts toward sounding like advice.

Case study · Reporting & revenue agents

630 hours of reporting by the most expensive people in the building

A full-service marketing agency — 60 staff, roughly $14M revenue, 45 retained clients.

The problem, measured first

Three structural drains. Account managers spent about fourteen hours per client per month assembling reports — pulling numbers from analytics, paid social, search and the CRM, then writing commentary — around 630 hours a month, performed by the most expensive client-facing people in the building. The fourteen-hour figure came from observation, not estimate. Meanwhile, client sales teams took hours to days to respond to the leads the agency generated — and renewals were judged on cost per acquisition, a step the agency didn't control. And approved creative briefs queued three to five days for a first draft.

What we built

A reporting agent that pulls platform data nightly, compares it against prior periods, cross-references the campaign changes logged in the project system, and drafts commentary explaining why numbers moved — producing a review-ready report for the account manager. A speed-to-lead agent deployed into client funnels that responds to form submissions within sixty seconds, qualifies, and books warm prospects for the client's sales team. And a brief-to-draft agent grounded in each client's approved past copy, so brand voice is inherited rather than invented.

  • Every figure is retrieved by an API tool — the model writes commentary but never produces a number, and unexplained movement is flagged, not explained away.
  • Nothing is sent or published without human review, and the lead agent never negotiates or quotes.
  • Messaging consent, quiet hours, opt-outs, and each client's own AI-disclosure policy are verified per client before launch — asked about, never assumed.

How it went live

The last three months of reports were regenerated for ten clients and put side by side with what was actually sent; account managers scored the commentary for accuracy, and anywhere the agent asserted a cause without evidence, the instructions were tightened. Three senior writers blind-reviewed drafts against a brief for voice consistency. Reporting launched on ten clients with full review, speed-to-lead with two willing clients, and drafting stayed internal until the writers trusted the output.

Reporting hours

~630 / month

~90 / month

Client lead response

Hours to days

Under 60 seconds

Brief to first draft

3–5 days

Same day

Reporting cost per client

~$1,050 / month

~$150 / month

The change, drawn to scale

Client reporting hours per month

Before
~630 h
After
~90 h

Reporting cost per client per month

Before
~$1,050
After
~$150

Approved brief to first draft

Before
3–5 days
After
Same day

Annualized value measured at over $500,000 in recovered senior time, plus the retention benefit of faster client results. What keeps it honest: the single most important guardrail is admitting uncertainty — a report that confidently invents a cause is worse than one that says a movement is unexplained. And the drafting corpus contains only genuinely approved work, because writers reject drafts that miss the voice.

Every engagement follows the same discipline we apply everywhere: measure the baseline first, build against it, prove the change with a re-measurement, then control it so the gain holds. Results above were measured on these specific engagements — they are not a guarantee of outcomes.

What would this look like in your business?

Every engagement starts the same way these did: with a measured baseline. Book a call and we'll find the workflow where the numbers say automation pays back first.