ChatGPT is already in use in many companies. The real question isn’t whether, but where your data goes the moment an LLM processes it.
In many companies, ChatGPT has long been in use, officially or unofficially. That settles the question “Should we use LLMs?”. What remains is the more important one: What happens to your data along the way?
“GDPR-compliant” is written on almost every vendor PDF today. What that means technically varies enormously. And nobody can seriously give you a blanket guarantee of compliance anyway: it depends on your individual case and needs legal review.
What can be said: which paths exist, what they can do and what they cost. That is exactly what this article is about.
The new question is data flow
Where data is stored is now part of every software company’s standard repertoire. Encryption, EU data centers, access control: largely solved problems.
AI adds a different question: Where does your data go the moment an LLM processes it? Every prompt containing customer data, contracts or health information takes that path. And the path differs fundamentally from provider to provider.
On top of that comes a layer that often gets overlooked: it makes a difference whether your employees use ChatGPT as a tool or whether an LLM is built into your product and your processes. At the latest when you build it in, you need a deliberate architecture decision.
For using LLMs with company data, three paths have become established.
Path 1: Frontier models directly
The direct line to the top models: Claude, GPT or Gemini via the provider’s API.
In favor: the most functionality, the fastest path and the cheapest way to use frontier models. Providers like Anthropic don’t train on your data, and transmission is encrypted.
Against: the data leaves the EU, the servers are mostly in the US. “Not used for training” doesn’t mean “stays in the country”. In practice, a prompt from Vienna is usually processed in the US.
For many applications that’s not a problem, and then the simplest path is often the best. With health data, HR data or trade secrets it’s a decision you should make deliberately, not by default.
Path 2: Frontier models via an EU cloud
The middle path: via AWS Bedrock, models like Claude run in an EU region. An additional security layer sits in between, and here too, your data isn’t used for training.
Against: it’s more expensive and sometimes offers fewer features than the direct API. And AWS remains a US corporation. For many cases with higher protection needs, it’s still a workable compromise.
Important for context: this is about the models available via Bedrock. Which ones those are changes constantly and should be checked at project time.
Path 3: Open-source models, hosted locally in Austria
The sovereign path: proven open-source models, hosted on Austrian servers. Your data never leaves the country, at no point during processing.
Economics speak for it too: above a certain volume and with the right task, local operation is cheaper, and it makes you independent of fluctuating token prices.
Against: as a rule of thumb, open-source models are about half a year behind the top models. For most tasks that’s enough. For the most demanding ones, not always.
So what does “GDPR-compliant” mean?
The honest answer remains: GDPR compliance depends on the individual case and needs legal review. You should question any blanket guarantee.
What technology can do: make compliance achievable. Strictly separate sensitive data. Log via prompt tracing what goes to a model and what comes back. And make sure nothing quietly leaks to an external model.
One thing stays the same regardless of the technology: you remain responsible for data protection. Technology doesn’t take that responsibility off your shoulders. But it makes it manageable, through traceability instead of flying blind.
In the end, the answer is usually a mix: sensitive data local, the rest via EU cloud or direct API. Cut along your data, not along a sales pitch.
The mix has a side effect that is often underestimated: resilience. Large model providers have outages, and token prices fluctuate. With a tested local system in reserve and automatic LLM routing, your product stays available even when the provider is down.
We know these questions from practice. With our own start-up Novid20, we supported 35 million lab tests strategically and technically, with sensitive health data, live and under the scrutiny of authorities and the media.
Storage location has become a commodity question. Data flow hasn’t, by a long way. If you can answer it for your data, you don’t have to fear AI. Not even with the most sensitive data.
N° 06 Contact
Tell me what you’re planning.
No form, no ticket system. Email me directly and you get me, not an inbox someone else manages.