AI

You are not ready for AI until you own your own data

AI value does not come from the model. It comes from your own, clean data in one place that you control. Most companies do not have that.

Berkan Alci, founder of YK TechnologiesBerkan Alci7 min readExecutives, IT and operations

In brief

  • About 55% of your company data is dark: stored but unfindable and therefore unusable for AI (Splunk).
  • Around 90% of your unstructured data is never analysed; your AI reasons on a tenth of reality (IDC).
  • Poor data quality costs companies an average of 12.9 million dollars a year; a model amplifies that error instead of catching it (Gartner).
  • SaaS silos hold your data behind limited exports and separate APIs. On paper it is yours, in practice you rent access.
  • The order is fixed: first data in your own repository, then clean, then AI. Reverse it and you automate your own noise.

Every executive team gets the same question today: what are you going to do with AI. The demos look strong, the first pilot starts smoothly, and then it stalls. Almost always in the same place. Not at the model, but at the data underneath.

That is actually good news. It means your advantage is not sitting and waiting at a vendor. It is under your own roof, in data you already have but do not yet own in a way a model can use.

The problem is not the model, it is the data

The models have become a commodity. The large commercial variants and the open-source models you run yourself perform within the same margin, and they get better every month without you lifting a finger. That is not where you win.

You win at what the model gets to see. An AI that knows your quotes, your service history, your stock movements and your margins does something no general model can. An AI that cannot reach any of that gives you a smooth summary of the internet. Correct in tone, empty in content.

The reason everyone looks at the model is understandable. The model is the visible part, the part that demonstrates, the part the vendor talks about. The data underneath is plumbing, and nobody likes buying plumbing. But the demo runs on tidy sample data. Your company runs on the messy real version. That difference is exactly where pilots fall over.

So the distinction between the two is not in the choice of model. It is in the question of whether your data is reachable, whether it is correct, and whether it sits in one place. At most companies the answer is no three times over. That is where the AI story ends before it begins.

There is a temptation to sidestep those three questions: buy an off-the-shelf AI product that promises to unlock your data for you. That only moves the problem. You give a new vendor access to data you still do not own centrally, and you add a silo instead of removing one. The bill for the clean-up comes later, and larger.

Dark data: the half you do not see

Splunk calculates that about 55% of all company data is dark: stored, but not findable, not classifiable, not usable (Splunk). You pay for the storage, you carry the risk that comes with it, and you get no value out of it.

Dark data is not exotic. It is the ordinary work that falls beside your systems.

  • Emails with the real agreements that never make it into a system
  • PDFs and quotes in a shared folder that no one can find again
  • Machine logs and sensor data that are recorded and never read
  • Exports that someone once made and that have quietly gone stale since

Keeping it is not free either. You pay storage for information you never consult. You carry the risk if there is personal data among it that no one keeps track of anymore. In an audit or an incident, data you cannot find again is not an asset, it is an exposure.

That mountain is also growing faster than the rest. Unstructured data, emails, documents, images, logs, takes a bigger share every year. Do nothing and the gap between what you store and what you can use keeps widening.

IDC estimates that of all the unstructured data companies create, about 90% is never analysed (IDC). For AI that is not one figure among others, it is the heart of it. A model is only as good as the context it gets. If nine out of ten documents stay invisible, your AI reasons on a tenth of your reality, and presents that tenth with the certainty of a whole.

Why SaaS holds your data hostage

That data lies scattered and out of reach is rarely carelessness. It is architecture. Each team once chose its own tool. Sales in one system, invoicing in a second, tickets in a third, planning in a spreadsheet that one person manages. Each of those tools holds your data in its own format, behind its own API, on its own terms.

On paper that is your data. In practice you rent access. You may look at it as long as you pay. The export is limited to what the vendor provides, the history only goes back so far, and every connection between two systems is a separate project with its own invoice. The vendor has no interest at all in data that flows out smoothly. Lock-in is not a side effect, it is the business model.

Take a simple situation. A customer calls in, angry. To know what is going on, someone has to look in sales at what was promised, in the ticket system at what went wrong, in invoicing at what is outstanding and in planning at when the next delivery leaves. Four systems, four logins, four exports that do not line up. What a person scrapes together with effort, a model cannot do at all, because it never gets to see those four pieces together.

As long as your data sits in five silos that do not know each other, you cannot ask a single question that crosses the boundary of one system. And those are exactly the questions that pay off: where margin leaks away per customer segment, which deliveries predict a complaint, where dead stock is sitting idle.

Data quality: the hidden drag

Suppose you do get your data into one place. Then comes the layer underneath: is it correct. Gartner calculates that poor data quality costs organisations an average of 12.9 million dollars a year (Gartner). Duplicate customer records, addresses that no longer exist, fields each team fills in differently, amounts in the wrong currency.

With manual work a person catches those errors. Someone sees that an order is not right and rings back. A model does not do that. It takes your data literally and builds on it with full conviction.

Bad data does not produce a visible error. It produces a smooth, credible answer that is simply wrong.

That is the insidious part of AI on a shaky foundation. The error does not disappear, it is multiplied and dressed up to look credible. Whoever first puts their data in order and then switches on a model wins twice. Whoever reverses the order automates their own noise and pays to spread it faster.

Quality problems also stay hidden as long as systems stand apart from each other. Each system is more or less internally consistent. Only when you bring them together do the definitions clash: which date counts, which customer is the real one, which amount includes VAT. The same exercise that makes your data ready for AI exposes those cracks. Better that you see them than that a model smooths them over invisibly.

What ownership changes

Ownership means one concrete thing. Your data sits in a repository that is yours, in a format you can read, on infrastructure you control. Not rented, not scattered, not held hostage behind the terms of a third party.

Once that is in place, the question shifts. No longer which model you choose, but what you want to know. Sales, service, planning and invoicing sit in the same source, in the same language. A model takes them in at a single glance. The AI you then build works on your full reality, and you can trace right back to the source where every answer comes from.

The right order First your data in one place that you own. Then the data clean. Then a model. Every step you skip lets AI amplify the error instead of solving it.

That one source pays for itself long before the first AI application. Reporting that costs half a day of cut-and-paste today becomes a query. New people find their way in one system instead of seven. And the day you switch on a model, the hard half of the work is already done.

Ownership also makes you independent of which model wins. Today you pick one, next year a better or cheaper one. If your data is in one place and it is yours, you switch models without touching your foundation. If your data is locked inside a vendor's product, you switch nothing without starting over.

The path there is not a new AI project on top of the existing sprawl. It is clearing up that sprawl itself. Start with an audit of where your data sits today, what state it is in and who actually owns it. Then build towards one platform you own that replaces your scattered tools and brings everything together in one place you control. Only once that base is in place is the question about AI worth asking.

That is the order we hold to at YK Technologies. We build one open-source platform that you own, with the code in your own repository from day one, on EU cloud and with security at NIS2 level. No per-user licences, no lock-in. We work in waves, with a go/no-go at each step, so you are never stuck with a choice you can no longer undo.

AI comes after that. Not because it matters less, but because it only works with solid ground underneath. The model is the easy part, it is ready and gets better on its own. The data underneath is your company. Whoever owns it owns their AI. Whoever rents it builds on sand.

Want to apply this to your own situation?

Belgian, founder-led and built to hand over. One email is enough.