8a. PII protection (Settings → PII)#

Two independent, opt-in toggles — either can be on alone:

  • Mask PII in answers: masks emails, phone numbers, card/account numbers, and ID-shaped values (SSN/SIN) in what answer surfaces show — the dashboard's Ask tab, /v1/answer, /v1/query, and MCP search_knowledge. The model still sees the original text to generate the answer, and the underlying documents/chunks are stored and re-served untouched — this is a presentation-layer mask, not data scrubbing.
  • Scrub PII at ingest: masks the same patterns in a document's text (and, for agent memories, the statement) before it is ever stored — identifiable data never persists in the knowledge base. Applied to every ingest path (uploads, connector crawls, scanned-image transcription, deep page rendering, and agent memory deposits) and to a translated document's preserved original-language text alike. Forward-only: turning this on applies only to content ingested from that point forward — anything already indexed is unchanged unless you re-ingest it (re-upload, re-crawl, or turning translation/deep-render loose on it again).

Both are off by default, per workspace. v1 is deterministic pattern matching only (the same five categories: email, phone, card, IBAN, SSN/SIN) — this is not HIPAA Safe Harbor de-identification and should not be relied on as such; a named-entity-recognition mode is a possible future addition, not something either toggle claims today.

LLM de-identification (Safe Harbor categories) — a third, stronger opt-in: an LLM pass at ingest masks names, locations, dates, ages, and record numbers (the categories patterns can't catch). Adds per-document LLM cost, and fails closed: if the pass errors, the document is NOT indexed (retry or disable) — for privacy, silent failure is worse than loud failure. This is LLM-assisted de-identification designed around the Safe Harbor categories, not a HIPAA compliance guarantee.