Data boundaries, a plain explanation for site owners
You are responsible for what your assistant does with visitor data, including the parts your vendor handles.
When you put an assistant on your website, you take on a position you may not have thought about explicitly: you are the party your visitors are trusting, regardless of who built the software. They did not choose your vendors. They chose to type something into your site.
The boundary you control
Three decisions are genuinely yours, and they are the ones with the most effect on exposure:
**What you index.** If you upload internal documents to make the assistant more capable, you have moved private content into a pipeline that includes third parties. Sometimes that is fine. It should be deliberate.
**What you ask for.** Every field you collect is data you now hold. A lead form asking for a phone number creates an obligation that no form would not have.
**How long you keep conversations.** Longer retention gives you better analytics and a larger exposure. This is a real tradeoff with no universally correct answer.
The boundary your vendor controls
Model provider terms, retention at the provider, processing location, and subprocessor relationships. You cannot change these but you can ask about them, and you should before you publish anything describing them.
The specific questions worth asking any vendor are in the article on vendor evaluation. The short version: ask for it in writing, and treat a vague answer as an answer.
What to tell visitors
A short, plain disclosure at the point of interaction is better than a long policy nobody opens. Something that covers: that they are talking to an AI assistant, that the conversation is stored, roughly what for, and how to reach a human instead.
Three sentences near the chat input does more real work than three pages of policy, and the two are not alternatives. You need the policy as well. The disclosure is what people actually read.
Things not to say unless you can evidence them
- That conversations are not used for model training. This depends on provider terms you may not have verified.
- That data never leaves a particular region. Verify the processing location first.
- That you are compliant with a named framework. Compliance is a determination, not a description.
- That data is deleted immediately. Deletion propagation through backups and provider systems takes time.
- That the assistant cannot make mistakes. It can.
Each of these is commonly written on websites and each is checkable. The cost of being caught overstating data handling is disproportionate to whatever the claim was worth.
The honest position for a pre launch product
If you are launching something and the details are not settled, say so. "We are finalising our provider agreements and will publish retention terms before general availability" is a sentence that costs nothing and protects everything.
Visitors do not expect a pre launch product to have a mature compliance posture. They do expect it not to lie about having one.
About the author
Vishal, CTO, Creoglyph. Writes about the systems view: retrieval, routing, reliability, data boundaries and deployment.