Essays on AI softwareWednesday, October 7, 2026Independent · No sponsors
TalkAI Tools
EssayBy The Talk AI Tools editorsOctober 7, 2026

Who reads what you type

Who reads what you type

There is a sentence that every small company now looks for before switching on an AI assistant, and every vendor now supplies it. Microsoft's version, in the Copilot privacy documentation updated September 30, 2026, reads: prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs. Google's, in the Workspace privacy hub, reads: Workspace does not use customer data for training models without customer's prior permission or instruction. Anthropic's Team plan page lists no model training on content by default. Having found the sentence, most buyers stop reading. This essay is about the pages beneath it, which is where the product actually lives.

What this note covers, in order.
What this note covers, in order.

The sentence, and the two words that qualify it

Read the three versions again with a lawyer's patience. Microsoft's is unconditional on its face. Google's carries the phrase without customer's prior permission or instruction, which is a door, and doors are there to be walked through, usually by a setting an administrator does not remember enabling. Anthropic's carries by default, which is the same door with a different handle. None of this is sinister; it is what a vendor writes when it wants to offer an opt-in later without re-issuing a promise. But it does mean that the sentence a buyer takes as a wall is, in two of three cases, a wall with a gate, and the gate is controlled from an admin console that in a ten-person company is managed by whoever set up email three years ago.

The more useful observation is that training was never the main risk for a small company. A model trained on a fragment of your proposal does not leak your proposal; it adjusts a few hundred billion weights by an amount nobody could recover. The risks that actually land on a small business are duller and nearer: who inside the company can see what the assistant was asked, how long the record of asking survives, and which third parties handle the text on the way. The training sentence answers none of those. The rest of the documentation does, and it is considerably less quoted.

Retention is where the data actually goes

Microsoft's documentation is explicit that interactions are kept. The user's prompt and Copilot's response, with citations to the grounding sources, are stored as what the page calls the content of interactions, forming a Copilot activity history. The page says this is encrypted at rest, is not used for training, and can be searched by administrators with Content search or Microsoft Purview, with retention policies set by the organisation. Users can delete their own history through the account portal. In other words, every question an employee asks the assistant is, by design, a discoverable record that an administrator can read, and the length of its life is a policy decision the company may not know it has made.

Google's hub describes the same shape with different defaults: retention for Gemini in Workspace is set by administrators and ranges from 90 days to indefinite, the consumer Gemini app keeps data for up to 36 months, and Gemini Notebook prompts and responses are not retained after the session ends. Notion's pricing page offers zero data retention as an Enterprise feature, which tells you both that it is possible and that it is sold as an upgrade. Anthropic's October 6, 2026 cyber program announcement notes that retention is required for monitoring in that program, with exceptions for customers already holding zero-retention access. The pattern is consistent across vendors: retention is the default, zero retention is the premium, and the number of days is a dial somebody has to find.

The assistant inherits your permissions, including the wrong ones

Both large vendors make a promise that sounds like protection and functions as a mirror. Microsoft's page says Copilot only surfaces organisational data to which individual users have at least view permissions, and adds, in the same paragraph, that it is important to be using the permission models available in services such as SharePoint. Google's hub says that if the user does not have access to a document or email, Gemini will not retrieve that content. Both are true. Both also mean that an assistant will cheerfully summarise the salary spreadsheet that was shared with Everyone in 2022 because somebody needed it for a day.

This is the clause that should change a small company's behaviour more than any other, and it rarely does, because it does not read like a warning. Before an assistant, a badly shared folder was a theoretical exposure: the file existed, but nobody searched for it. After an assistant, search is the product. The right response is not to distrust the vendor; it is to run the sharing audit that was postponed because nothing was looking. Microsoft's documentation even points at the relevant tools, Purview labels and restricted permissions, which the assistant is said to honour. The tools were there before the assistant. The assistant is what makes using them urgent.

The third model in the room

A buyer who has read the privacy page of one vendor has, increasingly, read the terms for one of three. Microsoft's documentation states that it offers third-party models within Copilot, naming Anthropic and OpenAI as subprocessors, and that an administrator decides whether those models are used, with additional terms possibly applying. It also notes that models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary. Notion's September 15, 2026 release notes describe a model picker offering Opus 5, GPT-5.6 Sol and Kimi K3, and a September 9 release lets workspace owners restrict which models agents may use, on Business and Enterprise plans. The workspace you bought is the lobby; the models are tenants, and the lease terms differ by door.

For a small company this creates a question the privacy page cannot settle on its own: which of the models is actually answering, and under whose data terms? The honest answer, from the documentation, is that the administrator chooses, the default may change with a release, and the contractual position can differ between the model you were shown in a demo and the one selected on a Tuesday. Microsoft's page says that model updates do not change security, privacy or compliance settings, which is reassuring about the settings and silent about the subprocessor list. The place to look is the admin console, not the marketing page, and the thing to write down is the model list on the day you signed.

How a small company should read these pages

None of these steps requires trusting or distrusting a vendor. They require reading pages the vendor has already published, which is a lower bar than it sounds, because the pages are long and the sentence everyone wants is near the top. The vendors are not hiding the rest. They are relying, reasonably, on nobody scrolling.

Bottom line

The training promise is real and it is the least important thing on the page. What decides whether an AI assistant is safe inside a small company is how long it remembers, who can read the memory, what it can reach through permissions you forgot you granted, and which models are answering under which contract. All four are documented, all four are configurable, and all four default to the setting that is easiest for the vendor to support rather than the one a cautious buyer would choose. The privacy page is the product specification. Read it as one, as of the day you read it, and then read it again when the release notes mention a new model.