← Back to blog
Aug 10, 2026·9 min read

How to Audit Your Firm's Tool Stack (And Why Most Firms Guess Low)

How to Audit Your Firm's Tool Stack (And Why Most Firms Guess Low)

The short answer: most firms underestimate their own tool count, because nobody has ever traced a single engagement end to end and written down every system someone opened. This is the method we use to do that - what to count, what to exclude, which four numbers matter, and how to read the result. It takes about two hours for one engagement and it is the cheapest diagnostic available before any consolidation or AI project.

We kept running into the same conversation. A firm would tell us they were pretty consolidated, and then we would walk through an actual client engagement from first contact to final invoice and count eleven tools. Not because anyone made a bad decision, but because each tool was the right answer to the problem in front of someone at the time.

The method

Pick one representative client engagement that has already closed. Trace it from first inbound contact through final invoice, and record every tool that touched the work. Count a tool if it held data required to complete the engagement, or if someone had to open it during the process.

Two rules keep the result honest:

  • Exclude email, calendar, and general-purpose chat. Every firm has them, they inflate every count, and they are not the problem.
  • Do not accept an estimate. Ask people what they think the number is before the walkthrough, write it down, and compare it to what the walkthrough produces. The gap between the two is usually the most instructive part of the exercise.

Do the walkthrough with the person who did the work, not the person who designed the process. Those produce different answers, and only one of them is real.

What to record

FieldWhy it matters
Tools in the engagement pathThe headline count. Name each one; a list is more useful than a number.
Tools holding client identityName, contact details, engagement history. Three or more is the consolidation trigger.
Stated system of recordWhat the firm says is authoritative.
Actual system of recordWhat the walkthrough shows. The gap here predicts agent failure.
Reconciliation hours per weekTime spent only because information lives in more than one place.
Loaded hourly rateTurns hours into a number finance will engage with.
Service linesTool count tends to track service lines more closely than headcount.

The four numbers that matter

Everything else is context. These four determine what you should do next.

  1. Tool count in the engagement path. Not the number of subscriptions the firm pays for - the number someone has to open to deliver one piece of work.
  2. Number of systems holding client identity. If the same client exists as four partial records, you do not have client records; you have fragments a human has been assembling from memory.
  3. Reconciliation hours per week. Re-entering the same data, checking whether two systems agree, chasing which version is current, and hand-assembling views that should be queries.
  4. The estimate gap. Guessed count minus actual count. A large gap means process decisions are being made on a picture of the firm that is not accurate.

Convert the third into money using a loaded hourly rate, and state which rate you used. If you use the firm's own billing rate, say so - it is both defensible and a larger number than a salary-based estimate, because the hours are billable in principle and unbilled in practice.

How to read your result

Client identity in one system. You are in good shape. Spend your effort elsewhere.

Client identity in two systems with a clear authority rule. Manageable. Write the rule down so it survives staff turnover, and make sure anyone building automation knows it.

Client identity in three or more systems, or an authority answer that begins "it depends what you are asking about." This is the consolidation project, and everything else is secondary. It is also the specific reason AI pilots stall: an agent can only work from the record it can read, and a firm whose client record is spread across four systems cannot point an agent at the client record, because there is not one.

What to do with the result

Three things, ordered by how quickly they pay off.

  • Write the authority map first. For every field that matters, name the system of record. This costs an afternoon and removes most of the confusion that reconciliation work exists to resolve.
  • Attack the reconciliation hours, not the subscription bill. The subscriptions are the smallest line item. The labour is where the money is, and it is invisible because it is paid in salary rather than invoiced.
  • Do not start with AI. Deploying an agent on a fragmented stack amplifies the fragmentation rather than papering over it. The agent reads one of four systems and answers confidently from partial data, which erodes trust faster than doing nothing.

If you run this across several engagements, resist the urge to turn a small sample into percentages. Eight of eleven is honest and quotable. Seventy-three percent from a sample of eleven is neither.

Butterbase gives knowledge service firms one substrate where intake, CRM, client portals, support, and invoicing write to the same client record from the outset - so reconciliation work disappears because there is nothing left to reconcile. Related reading: What Your Firm Actually Spends to Run on Five Tools and Why AI Agent Pilots Stall at Service Firms.

Frequently asked questions

Pick one closed client engagement and trace it from first inbound contact to final invoice, recording every tool that held required data or that someone had to open. Exclude email, calendar, and general-purpose chat, since every firm has them and they inflate the count without explaining anything. Do the walkthrough with the person who did the work rather than the person who designed the process, and write down the guessed tool count first so you can measure the gap.

There is no reliable published figure, and any number quoted without a stated methodology should be treated with suspicion. What is consistent in our own conversations is that firms guess low: the count produced by an end-to-end walkthrough is routinely higher than the count the same people gave before the walkthrough. The useful number is your own, and it takes about two hours to produce.

The system of record is the source that is authoritative for a given field when two systems disagree. It matters for AI because an agent has no unconscious knowledge of which system to trust - staff know billing is right about status and the tracker is right about assignment, but that knowledge is nowhere written down. Without a documented authority map, the agent answers confidently from whichever fragment it read, which is the fastest way to lose user trust.

Any work that exists only because information lives in more than one place: re-entering the same client details across systems, checking whether two records agree, chasing down which version is current, and manually assembling a view that should have been a query. It is worth measuring separately from ordinary admin because it is the portion that disappears entirely under consolidation, rather than being reduced.

At minimum, resolve authority before deploying. Full consolidation is the durable answer but it is a longer project; a written authority map is the cheap version and removes most of the failure modes. Deploying an agent onto a fragmented stack without either amplifies the fragmentation, because the agent produces fluent answers from partial data and nobody can tell which answers were built on the wrong fragment.