Confidential data goes to AI? Experts point out where companies make the biggest mistake
“The biggest gap is between how quickly employees started using AI and how companies prepared for it.”
Artificial intelligence has become a permanent presence in Polish business and it has done so at an express pace. Faster than regulations were created, faster than companies had time to prepare procedures, and sometimes even faster than common sense suggested what could be pasted into the chat window. Today, AI helps write, analyze and speed up work, but along with convenience a new question has arisen: what actually happens to the data we entrust to algorithms?
I talk to Wojciech Chmiel and Jakub Dolata, creators of Anonimizer – a Polish technology startup, about the gap between the pace of AI implementation and security, trust in technology providers and why, before the document reaches the model, it is worth “cleaning” it of what the model does not need to know.
Beata Anna Święcicka, “Wprost”: AI entered companies faster than security procedures. Where is the biggest gap today: in technology, regulations or employee awareness?
Wojciech Chmiel, Anonymizer: The biggest gap today is between how quickly employees started using AI and how companies have prepared for this change. In many organizations, AI is already an everyday work tool, while security rules are still based mainly on instructions and prohibitions.
If an organization does not provide employees with tools that allow them to use AI in a safe manner, some people use publicly available, consumer solutions. Then external models may receive customer data, fragments of contracts, correspondence or other information that should not leave the organization.
Therefore, simply banning the use of such tools does not solve the problem. The company should give the employee a safe alternative. Our approach is to limit the scope of transmitted information and secure data that AI does not need to perform the task before using an external model.
Jakub Dolata, Anonymizer: From the technological side, it is also important not to require the employee to constantly make decisions about safety. The user should have a tool that helps him prepare the document before using AI. In Anonymizer, this stage takes place locally, on the user’s device. The source document does not leave his computer.
Companies make employees sign declarations regarding the safe use of AI, and a moment later one of these employees pastes the contract or customer data into ChatGPT to get their work done faster. What poses a greater threat to companies’ data security today: an external cyber attack or careless use of AI by employees?
Wojciech Chmiel: I wouldn’t try to rank these threats. Cyberattacks remain a very serious risk, but companies have known for years that they need to protect against them, and there are many experts who specialize in this area. The flow of data into AI tools is a newer issue for many organizations.
An employee can paste a fragment of a contract or customer data into the model and not feel that he has just created a risk. From his perspective, he is simply using a tool that allows him to get the job done faster.
Simple rules, education and convenient tools can significantly reduce this risk. The employee usually doesn’t want to compromise security, just get the job done quickly. Therefore, the company should design the process of using AI in such a way that data protection does not depend solely on whether the user remembers all the rules every time.
The AI Act is intended to increase security, but are we at risk of a situation in which companies will have more procedures than real control over what happens to data?
Wojciech Chmiel: We can discuss the risk of overregulation, because excess formal obligations do not always translate into greater safety. However, in the case of the AI Act, I see a lot of positives, especially an increase in the awareness of users and organizations regarding the responsible use of AI. The key will be whether the new obligations will translate into real processes and tools, and not just further documents.
Regulations are valuable when they help organize responsibility and lead to a real change in the way of working. If an organization adopts the principle that certain information should not be included in AI models, it must be able to enforce this principle in practice and know where the data actually flows.
For me, this is where the line between formal compliance and real control lies. The procedure sets a rule, technology allows it to be applied in everyday work.
Is data safe just because an AI provider declares that it does not use it to train the model? What should a company check before uploading a confidential document to external AI tools?
Jakub Dolata: NO. This is politics and policies change. This spring, Microsoft announced that data from GitHub will be used by default to train models, also on paid plans. The change only affected business contracts. The rest of the users had to go into the settings and turn it off themselves, even though they had accepted the terms a year earlier.
Training is only one of the risks. The provider can store conversation history, and sent documents also become part of this environment. OpenAI confirmed an error that caused some users to see titles and fragments of other people’s conversations. It had nothing to do with training. It was a failure, and such failures happen to everyone we entrust data and documents to, not just AI providers.
You can sign an agreement prohibiting the use of your data, but such terms are not always available in basic plans and do not remove all risk.
The real question is different: which information from this document is the model even needed? When analyzing a contract, the real name of the client or the name of the project is rarely needed. You can secure this in advance, at home.
Wojciech Chmiel: From the company’s point of view, this is an important change in thinking about security. Instead of basing protection solely on the terms and conditions of a specific AI vendor, an organization can minimize the scope of information that leaves its environment in the first place.
You have decided to process your data locally. In a cloud-based AI world, could “data never leaves the device” become the new security standard?
Jakub Dolata: I don’t think this will become a universal standard, nor should it. The best models will stay in the cloud because that’s where the computing power is. But the anonymization itself should be local. By doing it in the cloud, we make the entire original document available to the supplier, and we anonymize it in order not to share it. Division of labor becomes the standard. Preparing the document, i.e. finding data and replacing it, can be done by the user. The actual task is then performed by the external model, but with no content that is not needed for the task.
It makes sense to do this only now. A few years ago, such a product would have required a server because the office equipment was too weak. Today, on an office computer, we can automatically clean a document in seconds.
But locality has another feature that is rarely talked about. It can be checked. You can only trust the promise of the cloud provider, because we will not see what happens to the file after sending. Here, just disconnect your computer from the Internet and check if the program still works.
Wojciech Chmiel: From a business perspective, this is exactly the point: companies want to benefit from increasingly better AI models while maintaining control over confidential information. Solutions that combine these two needs will be an increasingly important part of a company’s AI infrastructure.
Given the current dynamics of AI development, what might data security look like in 4-5 years? Will companies protect data even more against AI, or will security be built into the very tools we will use in the near future?
Wojciech Chmiel: I expect that security will become increasingly embedded in the way we use AI. With the growing scale of these tools, it will be difficult to base data protection on hundreds of individual employee decisions. Companies will need solutions that apply their security principles directly to the moment they work with information.
Jakub Dolata: Companies should protect data from AI, but not only from it. Even if suppliers guarantee that their systems have built-in security, it doesn’t change much, because we already get such guarantees today. Whenever we share documents with some cloud service, it stores them. Its employees make mistakes and systems have failures. The development of technology will not change this.
In a few years, documents will increasingly be sent not by a person, but by an agent acting on his behalf and without asking for consent. At this scale, being mindful of what we share and with whom will be more important than ever. Remember that if the data is not there somewhere, it cannot leak from there.
Thank you for the interview.
