Skip to content
Datos y confianza

What data from your company should not go into an AI model

The question is not «is it safe?» but «what happens if this turns up tomorrow inside somebody else’s answer?».

Published 9 min read Navhera

The question always arrives in the same form: “is it safe to paste this into ChatGPT?”. And it makes no difference which tool it is —ChatGPT, Claude, Gemini, Copilot or whichever one comes built into the program you already use—: the criterion is the same, and the terms vary more between plans of the same tool than between different tools. And it is always badly framed, because “safe” depends on a contract almost nobody read and on a setting almost nobody checked.

The useful question is a different one: what happens if the thing I am about to paste turns up tomorrow on someone else's screen, or in a lawsuit? That one can be answered without being a lawyer or an engineer.

What actually happens when you paste something

When you type into an AI tool, that text leaves your company and travels to a provider's server. What happens next depends on three things: whether the plan you pay for allows your data to be used to train models, how long the provider keeps it and who inside the provider can read it. All three are in the contract, and all three differ between the free plan and the enterprise plan of the same product.

That is the point that causes the most confusion. None of these tools is a single thing with a single policy. In almost all of them, the free account, the paid individual account and the enterprise account have different terms, and the difference that matters is not one of brand but of plan: the same provider can treat your content in two opposite ways depending on what you contracted. The same holds for the other providers. Saying “we use such-and-such tool” says nothing about the risk until you say under which plan.

The setting almost nobody checks

In consumer accounts —free and paid individual— the main tools usually come with a switch in the settings that decides whether your content can be used to improve the model, and it usually comes turned on. Find it, read it and decide. It is the only item on this list that is fixed in a minute.

With two warnings. The first: turning that switch off is a setting, not a contractual commitment; it can change with a terms update and it gives you nothing to claim. The second: it says nothing about the other two questions —how long it is kept and who can read it— which still live in the plan's contract.

A three-zone traffic light

This is the criterion we use, and it can be pinned to a wall.

Green zone — can go in without ceremony

  • Text you would publish on your website without thinking twice.
  • Drafts, ideas, document structures, writing corrections.
  • Public technical information: how a tool works, what a term means.
  • Made-up or genuinely anonymized data used to run a test.

Practical rule: if you would paste it into a company social media post, it is green zone.

Amber zone — only with a contract and with judgment

  • Internal, non-public information: cost prices, margins, processes, policies.
  • Working documents with the names of employees or of providers.
  • The company's own code.
  • Customer data already aggregated or dissociated, in a form that identifies nobody.

Here the free tool will not do. You need an enterprise plan with a written commitment not to use the data for training, and you need someone in the company to have read that commitment.

If you are not clear on what the contract of the plan you already pay for says, it is a short and concrete review: what is kept, for how long, who can read it and whether or not it feeds training. The technical division does it with you.

Red zone — does not leave the company

  • Third parties' personal data without a legal basis to handle it this way: national ID numbers, phone numbers, email addresses, home addresses, purchase histories tied to a name.
  • Sensitive data. Law 8968 (Ley de Protección de la Persona frente al tratamiento de sus datos personales, Costa Rica's personal data protection act) defines it as information belonging to a person's private sphere, and expressly mentions racial origin, political opinions, religious or spiritual convictions, socioeconomic status, biomedical or genetic information, and sexual life and orientation. That “socioeconomic status” surprises everyone and is the one most often stepped on by accident: income, debts, ability to pay.
  • Credentials: passwords, access keys, tokens, certificates. A secret pasted into a chat is a secret burned.
  • Material under someone else's confidentiality agreement. If you signed an agreement with a customer, that agreement probably does not contemplate their information passing through a third party.
  • Information that decides something about a person: performance reviews, credit decisions, disciplinary files. Note that credit decisions also fall into the previous category.

With sensitive data, Costa Rican law does not limit itself to asking for more care: it prohibits processing it, and permits it only in narrow, listed cases —safeguarding a vital interest of someone who cannot consent, the legitimate activities of a foundation or association with respect to its own members, data the person made public voluntarily or that is needed to exercise a right in judicial proceedings, and purposes of diagnosis or medical care in the hands of professionals bound by professional secrecy—. A general-purpose AI tool is none of those cases. The red zone, on this point, is genuinely red.

The case that comes up most. Someone pastes in a customer list “just so it sorts it for me”. That list is a file of third parties' personal data that none of them authorized sharing with a foreign provider. Sorting it in a spreadsheet takes five minutes longer and does not create the problem.

What Costa Rican law says, in short

Law 8968 starts from the premise that personal data belongs to the person, not to whoever collects it. Two direct consequences for this topic follow from that:

  1. Consent is for a specific purpose. That a customer gave you their phone number to coordinate a delivery is not permission to pass it on to a third party for another purpose.
  2. Transferring data to third parties needs a basis of its own. The law only allows data to be transferred when the data subject has authorized it expressly and validly. An AI provider that processes your customers' data is a third party, even if you experience it as “a tool”.

This is not legal advice, and a company that handles sensitive data or large volumes should check it with a lawyer. But the operating criterion holds: if the data identifies a person and that person does not know it is going to pass through there, it does not pass.

How to decide in thirty seconds

Four questions, in this order. The first one that gives a “yes” stops the process.

  1. Does this identify a specific person who is not me?
  2. Is it a credential, or does it open access to something?
  3. Is it covered by a confidentiality agreement with a third party?
  4. Would it bother me to see it with my name on it in the press?

The minimum policy every company should have in writing

Half a page is enough, and it avoids the difficult conversation later on:

  • Which tools are approved and under which plan. Not “AI may be used”, but which ones and in what form.
  • The traffic light, with examples from your own business, not generic ones.
  • What to do when someone gets it wrong. Having a way to report it without punishment is what makes people report it. Without one, nobody finds out.
  • Who decides when a case does not fall clearly into any zone.

That it exists in writing matters more than how long it is. A half-page policy that people know works better than a twenty-page manual nobody opened. How we handle the data this site collects is in the privacy policy.

How we do it here

When we work with a company's data we apply the same traffic light this article describes: what goes to a model is decided before starting, it is written down, and if something falls in the red zone another way of solving it is found — anonymizing, working with a sample, or leaving that step in a person's hands.

How we can help with this

This article ends by asking three things of you: classify your data, review the contract of the tools you already use and write half a page of policy. All three are jobs with a beginning and an end, and they are part of what we do:

  • Classifying your data into the traffic light, with the examples of your business and not the generic ones in this article. That is what turns a table into something your team can use on Monday.
  • Reviewing what the terms of the tools you already pay for say and translating them into what they mean for your case.
  • Writing the policy and leaving it in language your team understands, with the route for reporting when someone gets it wrong.
  • Putting the data in order before automating anything, which is the step that appears when the information is spread across files that do not match.

If you want to start here, the technical division will review it with you in a thirty-minute meeting, no cost, no obligation to hire us.

Frequently asked questions

Can I paste my customers' data into ChatGPT, Claude or Gemini?

As a rule, no. Data that identifies a specific person is handled under the purpose for which that person provided it, and passing it to an external AI provider usually falls outside that purpose. If the business case demands it, you need an enterprise plan with a written commitment not to train on that data, and it is worth reviewing it with a lawyer.

Is the paid version of an AI tool safe for company data?

Paying is not the same as protecting. What changes the risk is the plan contracted: enterprise plans usually include a commitment not to use the content to train models, along with retention controls, and paid individual plans do not always. The answer is in the contract of the specific plan, not in the price.

What is sensitive data under Costa Rican law?

Law 8968 defines it as information relating to a person's private sphere, and cites racial origin, political opinions, religious or spiritual convictions, socioeconomic status, biomedical or genetic information, and sexual life and orientation. The law prohibits processing it except in listed cases, so it should not go into general-purpose AI tools. It is worth noting that socioeconomic status —income, debts, ability to pay— is on that list, which is not the case in other countries.

Does turning off training in the tool's settings help?

It helps and it is worth doing, but it does not solve the problem. It is an account setting, not a contractual commitment: it can change with a terms update and it leaves you nothing to claim. It also says nothing about how long the provider keeps your content or about who inside the provider can read it, which are the other two questions and live in the plan's contract.

Does anonymizing the data solve the problem?

Only if the anonymization is real. Removing the name and leaving the national ID number, the phone number or a combination of data that allows the person to be re-identified anonymizes nothing. To test a tool it is usually cleaner to make up example data than to try to depersonalize the real data.

Sources and notes

  1. Law No. 8968, Ley de Protección de la Persona frente al tratamiento de sus datos personales (Costa Rica, 2011), and its implementing regulation, Decreto Ejecutivo 37554-JP. Text on the MICITT site. Article 3, definition of sensitive data; article 9, prohibition of its processing and exceptions; article 14, transfer of data.
  2. Processing and retention terms vary by provider and by plan; the valid source is the current contract of the plan contracted, not the product's general documentation. That is why this article does not reproduce the terms of any specific tool: they would go out of date before they were of any use.
  3. This article describes an operating criterion and does not constitute legal advice.

Written by the Navhera team and reviewed before publishing. If you spot an error, write to us and we will correct it with a note.