Your AI agent is only as smart as the data you let it see. I keep repeating this sentence in client meetings, because almost every conversation about AI starts in the wrong place. Companies want to talk about which model to use, which vendor to pick, which agent framework is winning this quarter. Almost nobody wants to talk about where their data actually lives.
And yet that's where most projects quietly die. After dozens of engagements with Swiss SMEs, I can say it plainly: most companies don't have a model problem. They have a spreadsheet problem.
What the Spreadsheet Problem Actually Looks Like
Here's what I keep running into with clients: the first thing that blocks AI adoption isn't policy, compliance, or model quality. It's that the data lives in too many places. Not in one messy system (that would be fixable in a week), but smeared across a dozen tools, personal drives and inboxes, each holding a slightly different version of the truth.
The pattern is so consistent that I can usually predict it before I walk into the room. Three examples I see in almost every company:
The Personal Excel File
The campaign benchmarks the whole marketing team relies on live in someone's personal Excel file. It's meticulously maintained, by exactly one person, on their laptop, with their own naming conventions. When they're on holiday, the numbers are frozen. When they leave the company, the numbers leave with them. No AI agent will ever see this file, and neither will half the team.
The Half-Updated CRM
Customer information is split between a CRM that's half-updated and a SharePoint folder nobody maintains. Sales trusts neither, so they keep their real notes in email threads. Ask three people for the current status of a key account and you get three answers. An agent querying this CRM doesn't get the truth. It gets a confident snapshot of last quarter.
The Google Doc from 2023
Product specifications sit in a Google Doc last edited in 2023. Two product revisions have shipped since. The document is still the first result when anyone searches for specs, which means it's also what an AI assistant will retrieve and confidently quote to your customers. Outdated data isn't neutral; it's actively misleading at machine speed.
None of these are exotic failures. They're the normal state of a healthy, growing business that adopted tools one at a time over fifteen years. The problem is that this normal state is now the single biggest blocker between you and useful AI. You can connect the best model on the market. If it can't see clean, current data, it gives you garbage. Faster than ever, in fluent prose, with full confidence.
You can connect the best model on the market. If it can't see clean, current data, it gives you garbage.
Why This Matters Right Now
For years, the honest answer to many automation ideas was: the models aren't good enough yet. That excuse is gone. METR, a research organization that measures AI capabilities, published numbers showing that AI agent capability is doubling roughly every 4.3 months. Not every 4.3 years. Every 4.3 months. The tasks an agent can complete reliably keep getting longer and more complex at a pace no procurement cycle can match.
This changes where the bottleneck sits. Model capability is no longer the constraint for the vast majority of SME use cases: drafting offers, answering customer questions, reconciling reports, monitoring inboxes. The constraint is access: can the agent reach the data it needs, and can it trust what it finds there? A brilliant model reading a stale spreadsheet is still wrong.
Here's the uncomfortable implication: while you wait for your data to be 'ready someday', the gap between what agents could do for you and what they actually do for you widens every quarter. The companies that sorted out their data in 2025 are now deploying agents in weeks. The ones that didn't are still running pilots that impress nobody, wrongly concluding that AI doesn't work for their business.
The 30-Second Test
I've started telling clients: before we talk about agents, show me where your data lives. If the answer takes more than 30 seconds, that's the real project. Here's the full test: three questions, no slides, no workshop required:
Can you say where your data lives in 30 seconds?
For each core data type (customers, products, prices, projects, finances), can you name the one system that holds the current version? If the answer involves the words 'it depends', 'mostly' or someone's name, you've found your project.
Can your AI agent query it right now, without a CSV export?
'But we have a data team / data lake / governance framework.' Sure. But can an agent query it right now, programmatically, without someone exporting a CSV first? If a human has to manually extract data for the AI, you don't have AI automation. You have a slower human with extra steps.
Is it free of data you can't share with an agent?
Is the source clean of sensitive data (health information, salaries, data covered by confidentiality agreements) that must never reach an external model? If sensitive and shareable data are mixed in the same folder or table, every AI integration becomes a compliance discussion instead of an engineering task.
That's the bar. Not a data warehouse, not a three-year transformation roadmap. Just: known location, machine-accessible, safe to share. Most companies I meet fail at least two of the three questions, and most are genuinely surprised, because each individual team thinks their corner is fine.
The Fix: One Source of Truth per Data Type
The fix isn't a smarter AI tool, and it usually isn't a big-bang data platform either. It's picking a source of truth for each data type and actually keeping it current. Boring, unglamorous, and worth more than any model upgrade. Here's how I walk SMEs through it:
1. Inventory your data types
List the 5–10 data types your business actually runs on: customers, products, prices, projects, contracts, campaign results. For each one, write down every place it currently lives. This takes an afternoon, and the resulting list is usually embarrassing. Good: embarrassment is the budget argument you need.
2. Name one owner and one home per type
For each data type, pick exactly one system as the source of truth and exactly one person who owns its quality. Not a committee. A name. Customer data lives in the CRM, full stop. Product specs live in the product database, full stop. Everything else becomes a copy that's allowed to be wrong.
3. Set an update cadence
A source of truth that's current is infrastructure; one that's stale is a liability with a fancy name. Define how fresh each data type must be: prices weekly, pipeline daily, specs at every release. Then make updating it part of an existing routine, not an extra task that dies after three weeks of enthusiasm.
4. Separate sensitive data deliberately
Decide upfront what agents may never see: health data, salaries, personal data under the Swiss FADP and GDPR, customer secrets. Move it into clearly separated systems or fields with their own access controls. When sensitive data has a defined fence, everything outside the fence becomes instantly usable, and your compliance conversations get dramatically shorter.
5. Make it machine-readable and connectable
An agent can't click through your intranet. Each source of truth needs a programmatic door: an API, a database connection, or increasingly an MCP server, the emerging standard that lets AI agents query business systems directly and securely. PDFs of scanned tables and screenshots in slide decks do not count. If a machine can't read it, for AI purposes it doesn't exist.
Notice what's not on this list: a new AI tool, a data lake, a twelve-month program. For a typical SME this is weeks of focused work per data type, not years. Every step delivers value even before any agent is connected, because your humans suffer from the spreadsheet problem just as much as your AI does.
What Becomes Possible Once Your Data Is Ready
Here's the part that makes the boring work worthwhile. Once each data type has a current, machine-readable home, agent projects stop being research projects and start being plumbing. In our own client work, the difference is stark: with ready data, a working agent (one that drafts offers from the actual price list, answers support questions from the actual product specs, or briefs the sales team from the actual CRM) ships in days to weeks. Without it, the same project burns months on data archaeology before the first useful output.
There's a compounding effect, too. The first agent forces you to clean up one or two data sources. The second agent reuses them for free. By the third or fourth use case, you're connecting new automation to existing clean sources in an afternoon. Data readiness is the rare investment that gets cheaper to exploit every time you use it.
And remember the METR curve: capability doubling roughly every 4.3 months means whatever agents can do today is the floor, not the ceiling. Clean, connected, current data is how you make sure each new capability jump lands in your business in weeks instead of bypassing you entirely.
Before we talk about agents, show me where your data lives. If the answer takes more than 30 seconds, that's the real project.
How we work · Work · Insights