Guide / AI Consulting

What Your Data Has to Look Like Before AI Is Worth Buying

AI assistants inherit your file mess and your permission mistakes. A six-part readiness audit you can run before you spend a dollar on licenses.

The short answer

Before deploying an AI assistant, audit six areas: permission sprawl, version chaos, file naming and structure, what should never be indexed, the systems of record holding your operational data, and retention and disposal. An assistant has no judgment about which of six copies of a document is current, and it reads anything the asking employee can reach.

The pitch for an AI assistant that reads your company files is genuinely good. Ask a question in plain language, get an answer sourced from your own documents, skip the twenty minutes of hunting through folders.

What the pitch leaves out is that the assistant has no judgment about which of your documents deserves to be believed. It will read the 2022 pricing sheet with the same confidence as the current one. It will surface the draft contract and the executed contract without knowing the difference. And it will read anything the asking employee has permission to read, which in most small business tenants is considerably more than anyone intended.

Here is the audit we run before a client deploys anything.

One: permission sprawl

Start with the tenant-wide question. What can an average employee actually reach.

The specific failures we find most often: a SharePoint site created for a project in 2021 with company-wide read access that now contains financial models. An HR folder where permission inheritance broke during a migration. A shared drive holding scans of employee documents that somebody set to "anyone with the link" to make a one-time transfer easier.

None of these caused visible harm before, because nobody browses folders they have no reason to open. An AI assistant removes that friction entirely. The exposure was always there. AI just makes it queryable.

Fix this first. It is the highest-risk item on the list and also the one with the clearest remedy.

Two: version chaos

Count how many copies of your most important documents exist. Your master services agreement, your price list, your employee handbook, your standard proposal.

If the answer is more than one live copy in more than one location, an AI assistant is going to make version mistakes and it is going to make them confidently. The remedy is not technical. Somebody has to designate a single source of truth for each critical document class and archive the rest somewhere the assistant does not index.

This is boring work. It is also the difference between an assistant that saves your team time and one that creates rework.

Three: file naming and structure

You do not need a formal taxonomy. You do need enough consistency that a machine can tell a client deliverable from an internal draft.

The minimum standard: dated documents carry the date in the filename in a consistent format. Client work lives under a folder named for the client. Drafts are marked as drafts or live in a drafts folder. Superseded material is moved out, not just renamed.

Four: what should never be indexed

Decide deliberately which locations an AI assistant may read. Some things belong outside that boundary regardless of permissions: legal hold material, active litigation files, HR investigations, merger and acquisition work, and anything covered by a confidentiality obligation that would not survive being surfaced in a summary.

Most platforms give you a way to exclude sites or drives from indexing. Use it, and write down what you excluded and why. That document is worth having when someone asks later.

Five: the systems that hold your real data

Files are the easy part. The information a small business actually runs on usually lives in a CRM, an accounting system, a practice management platform, or a line of business application. Whether an AI tool can reach those systems, and whether you want it to, is a separate decision from documents.

Inventory those systems, note which have an API and which do not, and note where the same customer exists under three different spellings. Data quality problems in your CRM do not get fixed by AI. They get amplified by it.

Six: retention and deletion

If you have no retention policy, you have an infinite one, and the AI assistant is reading fifteen years of everything. That includes the email thread from 2014 that resolved a dispute in a way you would not want summarized back to a client today.

A basic retention schedule, applied to email and files, does more for your risk posture than most security tools. It is also a prerequisite for defensible discovery if you ever get sued.

What this actually costs

For a 50-person company with a Microsoft or Google tenant, the audit itself takes about a week. The remediation ranges from two weeks to two months depending on how much history there is. Almost all of it is work you should have done anyway.

The businesses that skip this step often get a mediocre result, conclude that AI is overhyped, and cancel the licenses at renewal. The tool was reading a filing cabinet nobody had opened in six years.

Get the workbook

The Data Readiness Audit Workbook walks all six areas with a scored checklist, a permissions review sheet, a source-of-truth register for your critical document classes, an exclusion log, and a remediation tracker with owners and dates.

Score under 60 and you are not ready to deploy. Score over 80 and you will get real value out of day one.

Frequently asked questions

Why do AI deployments fail in small businesses?

Usually because the assistant was pointed at disorganized data. It surfaces the 2022 price sheet as confidently as the current one, and it reads whatever the asking employee has permission to read. The result feels underwhelming, the company concludes AI is overhyped, and the licenses get cancelled at renewal.

What is the biggest security risk of deploying an AI assistant?

Existing over-permissioning becoming queryable. Most small tenants have a site or shared drive readable company-wide that contains something sensitive, usually from a migration where permission inheritance broke. Nobody browses folders they do not need, so the exposure stayed theoretical. An assistant removes that friction entirely.

How long does data preparation take before an AI rollout?

For a 50-person company the audit itself takes about a week. Remediation ranges from two weeks to two months depending on how much history there is. Nearly all of it is work you should have done anyway.

What should we exclude from AI indexing?

Legal hold material, active litigation files, HR investigations, merger and acquisition work, and anything under a confidentiality obligation that would not survive being surfaced in a summary. Most platforms let you exclude sites or drives. Document what you excluded and why.

Related reading

Put this into practice

AEGITz Data Readiness Audit Workbook

Use the working resource connected to this guide. No sales gate and no dead-end file link.

Download Excel workbookView resource details

Related reading

Keep following the decision.

Need help applying it?

Bring the real operating problem.

Schedule a Discovery Conversation