AI and maintenancePillar article

AI data leaks: the risk comes from your people and your contractors

Data leaks through AI do not come from the AI itself but from the people and third parties who use it: shadow AI, contractors and what the AI Act now requires.

8 min read

Chemical process unit and distillation structures
Chemical process unit and distillation structures

When sensitive data escapes through AI, the first instinct is to blame the tool. That is the wrong culprit. Artificial intelligence does not go looking for your data: people hand it over, and third parties process it on your behalf. The risk is not technical, it is human and contractual. And that is good news, because a human risk can be framed and controlled.

This article looks at where AI data leaks really come from, what the law now requires of you, and how to deal with it without falling back on a ban that does not hold.

The essentials

AI data leaks come from three sources, none of which is the AI itself: your staff pasting documents into consumer tools (shadow AI), your contractors and suppliers processing your data with their own tools, and the absence of any framework just as the AI Act comes into force. The answer is not to ban, but to map your actual usage, measure its risks and obligations, and take back control, where needed through AI deployed on-premise. Banning without an alternative makes the risk invisible; framing it makes it manageable.

The first risk is your own teams

Your teams already use AI, whether you decided so or not. ChatGPT, copilots built into office software, browser assistants: these tools have become working reflexes. And every time a member of staff pastes in an extract from a report, a table of measurements or a set of minutes, that data leaves the company for a third-party server.

This is not an abstract fear. According to a Cyberhaven analysis of 1.6 million workers, 11% of the data employees paste into ChatGPT is confidential, and an average company leaks sensitive information there hundreds of times a week. The same work highlights a point every manager should know: fewer than 1% of staff account for 80% of these leaks. In other words, the problem is not solved by raising everyone's awareness equally, but by identifying actual usage.

The best-known case is Samsung. In 2023, within a matter of weeks, engineers fed proprietary source code and the minutes of an internal meeting into ChatGPT, simply to save time. The company responded by banning generative AI for its staff. The detail that matters: none of those engineers was acting in bad faith. They were doing their job, with a handy tool, without knowing that what they typed left their perimeter for good.

This is what is known as shadow AI: AI that enters the company through the back door, without a decision, without controls, carrying off precisely the data you wanted to protect. In industry, what leaves is far from trivial: an inspection report describes the real condition of your equipment, a process parameter touches on trade secrets, a failure history reveals your weak points.

In industry, what leaks describes your plant

The Samsung example, or a manager pasting their strategy into a chatbot, is about code and office documents. In industry, the data that escapes is of a different nature, and often more revealing.

An inspection report pasted into a consumer tool means your wall thickness readings, your corrosion rates, your condition monitoring locations: the map of where your equipment is thinning and when a shutdown is approaching. A process parameter, in chemicals or food and beverage, touches on trade secrets. A failure history says where your production line gives way. A clean-in-place sequence is as much about know-how as it is about food safety. Aggregated, this information describes your site better than a competitor ever could.

That is the fundamental difference from a source-code leak: a company can rewrite software, but it cannot undo the fact that a third party now knows the real condition of its pressure equipment. On a major-hazard (Seveso) site, a pharmaceutical plant or a chemical unit, this data touches at once on industrial secrecy, safety and compliance. The leak mechanism is the same as at Samsung; the stakes are heavier, because they bear on equipment whose failure has physical consequences.

The second risk is your contractors and suppliers

The leak does not only come through your staff. It also comes through everyone who processes your data on your behalf. An inspection contractor using their own AI to write up reports, a software vendor adding an assistant feature, a subcontractor outsourcing a task: at every link in the chain, your data can pass through a tool you did not choose and whose workings you do not know.

Take an inspection campaign entrusted to a contractor. They take your thickness readings, write their report, and to work faster, lean on an online AI assistant to format their findings. Your measurements have just left your site, without your knowledge or authorisation. Multiply that by the number of contractors who come and go across a plant, each with their own tools and habits, and the exposure surface is no longer yours alone: it is that of your entire technical chain. A CMMS vendor bolting on an AI feature, an inspection body outsourcing data entry, a supplier hosting your history: so many doors you do not see, and yet for which you remain accountable.

The problem is that accountability cannot be outsourced. You remain responsible for what becomes of your data, even when a third party is handling it. A supplier that feeds a model with your reports, or sends them into a cloud beyond your control, creates an exposure you answer for. The question to ask before entrusting a task is therefore no longer only "is it done properly", but "what happens to my data while it is being done".

What the law now requires

The framework has tightened, and it targets exactly this chain. The European regulation on artificial intelligence, the AI Act, classifies uses by their level of risk and requires, for the most sensitive, documentation, human oversight and control of data. It comes into force in stages, which leaves little time to discover, on the day of an inspection, that you cannot answer.

To this is added the general data protection regulation. As soon as personal data enters the perimeter, an operator's name in a report for example, GDPR applies, and it draws a clear line between the data controller and its processors. A contractor processing your data remains bound by a contract that must state what it does with it. Finally, the artificial intelligence management standard offers a governance framework: how an organisation decides, documents and monitors its use of AI, just as it already does for quality or safety.

The reading rule remains the same as elsewhere: these texts are cited by their subject, never by an article number or a threshold you have not verified. What to take from them is simple: you are now required to know what your teams and your contractors do with AI, and to be able to demonstrate it.

Banning is not enough, and even makes things worse

Samsung's reaction, banning, is understandable, but it settles nothing on its own. Where AI is blocked on company workstations, teams do not stop working: they copy an extract from a report into a consumer tool from their personal phone, convinced they are saving time. The ban does not remove the risk, it makes it invisible and uncontrollable. Depriving your teams of a useful tool without offering them a controlled one is a sure way to guarantee they will find one behind your back.

How to deal with it, in practice

The right approach is not to deny the usage or to ban it, but to take it back in hand. It comes in three steps.

Map the actual usage. First of all, you have to know what is really going on: which tools are used, by whom, for what tasks, with what data, shadow AI included. This honest snapshot is the only solid basis. You cannot frame a usage you do not know about.

Measure the risks and obligations. For each usage, you assess what is at stake: which data is exposed, what reliability you can expect, which supplier you depend on, and which legal obligations apply. Not all uses are equal: asking for an email to be reworded does not carry the same weight as pasting in a technical file.

Decide, and provide a controlled alternative. Out of this analysis comes a prioritised action plan: what you allow, what you frame, what you replace. This is where technology meets governance. For uses that touch sensitive data, the most robust answer is a tool deployed on-premise, on your own infrastructure, where data is processed in place and never leaves. The leak vector then disappears by design, not by prohibition. This logic is developed in why responsible industrial AI is local and off-cloud.

In an industrial context, this local alternative is not a specialist's luxury. Maintenance and inspection data is precisely the kind that must not circulate, and processing it on site with a controlled tool amounts to removing the leak vector rather than chasing it after the event. The concrete deployment of such a tool is detailed in running local, off-cloud AI on an industrial site.

This trio, map, measure, decide, does not require AI expertise to be understood. It requires method, and the courage to look at what is already happening rather than turning a blind eye. A manager who knows where AI touches their data, who has measured what that exposes and who has decided what they allow no longer has a blind spot: they have a case they can defend, before a board as much as before an auditor.

This approach is not a certification, and it does not claim to be. It is a preparation: map, measure, decide, so as not to be caught out the day the question is asked, by an auditor, a board or an authority. The cost of such framing bears no comparison with that of a leak which, unlike the framing, cannot be undone.

The real issue is not AI, it is control

Take the three sources of leakage again: your staff, your contractors, the absence of a framework. None is the tool itself. All come down to the same point: knowing where your data is and who processes it. This is exactly what responsible AI designed in from the start covers, and what local deployment makes possible in practice. AI does not leak your data. It simply reveals whether you had kept control of it.

In industry, that control is not a matter of principle alone: the data at stake describes physical, safety-critical assets, and once a competitor or an uncontrolled provider holds the real condition of your installations, no software patch undoes it. The question is therefore not whether to use AI, but whether you can still say, at any moment, where your data lives and who answers for it.

Written by Adama CamaraAI Consultant · Industry · view profile

Published on July 29, 2026

Support

Custom AI systems for industry

Agents that put your data to work and extend your existing tools. Designed and run on site, off the network.

Visit Assets 4.0