Privacy
How we handle your transcripts
Communication is sensitive. This page explains exactly what happens to your data — what we redact before any AI sees it, what we deliberately keep and why, and what Myc accesses when you connect a meeting tool.
Your meeting transcripts may contain sensitive personal and business information. Before any transcript is sent to the AI for analysis, Making Yourself Clear automatically removes identifying information using a defense-in-depth approach. The redaction is designed to protect third parties who aren’t in your meeting (customers, competitors, colleagues mentioned in passing) and sensitive identifiers like SSNs, employee IDs, addresses, and credentials. This page explains exactly what happens — and what we deliberately leave intact, and why.
What happens when you upload a transcript
- 1
You upload a transcript
The file is parsed to text and stored in your account. It is not sent anywhere external yet.
- 2
PII redaction (before any AI call)
A local redaction layer (Microsoft Presidio plus custom recognizers) scans the transcript before any AI call. Every person’s name — the meeting participants (speakers) and anyone mentioned but not present alike — is reduced to bare initials (e.g. "Sarah Chen" becomes "SC", used consistently throughout), so the AI can still follow who said what without seeing real names. Other sensitive items — emails, phone numbers, addresses, government identifiers, employee IDs, API keys, dates of birth — are replaced with placeholder tokens like <EMAIL_1>, <EMPLOYEE_ID_1>. The original text never leaves our backend.
- 3
Redacted text is sent to the AI provider
Only the redacted transcript is sent to the AI provider (OpenAI by default). We send every request with store=false, which disables OpenAI-side storage of the conversation, and OpenAI does not train on API data. Standard provider retention (up to 30 days, for abuse-monitoring) still applies.
- 4
AI response references speakers by name
The AI only ever saw initials — every person’s name, yours included, is reduced to bare initials before the request leaves our backend. A reversible mapping of initials to names is kept encrypted on our side and applied when your analysis is displayed, so the coaching reads with real names even though none were sent. You can confirm exactly what was sent on any analysis via “View what was sent to the AI”, which always shows the redacted version. Other redacted details — emails, phone numbers, ID numbers — stay masked everywhere, including for you.
What we detect and redact
The redaction layer is configured with multiple aggressiveness levels (Conservative, Standard, Permissive). Standard is the default. The categories below are detected across all levels.
Identity
- Names of people (everyone) — Every person’s name is reduced to bare initials — both the meeting participants (speakers) and anyone mentioned but not present (customers, competitors, absent colleagues, family members). The same person maps to the same initials consistently throughout the transcript.
- Dates of birth — Explicit DOB references ("DOB 03/14/1968", "born March 14, 1968").
- Employee / staff IDs — Common formats like EMP-####, EMPID-####, STAFF-####.
Contact information
- Email addresses — Any string matching standard email patterns.
- Phone numbers — US and international formats.
- Physical addresses — US, Canada, and UK street address patterns.
Government & financial identifiers
- Social security numbers — US SSN patterns.
- Passport numbers — US passport identifiers.
- Driver license numbers — US driver license identifiers.
- Medical license numbers — Healthcare professional license identifiers.
- Credit card numbers — Full and partial references ("Amex ending in 1234", expiration dates).
Credentials
- API keys and tokens — OpenAI keys (sk-/pk-), Stripe keys, webhook secrets, generic labelled secrets.
Network & location
- IP addresses — IPv4 and IPv6 addresses.
- URLs — Web addresses that may identify internal systems.
- Locations — Cities, regions, countries (Conservative aggressiveness only).
- Nationalities / religions / political groups — Detected as NRP entities.
What we don’t redact (and why)
Some categories of information are intentionally preserved in the redacted transcript. These are deliberate design choices, not oversights — each has a specific reason.
Who-said-what (turn attribution)
We keep the conversation’s turn-by-turn structure and a consistent initials label for each speaker (e.g. "SC"), so the AI can follow who said what and give coherent coaching — "SC, you interrupted DP twice in the first five minutes" — without ever seeing a real name. If your security team wants speaker turns anonymised further, contact us.
Organization names (off by default)
Detection of organization names is configurable and off by default. If your admin turns it on in admin settings, organisation names mentioned in the transcript will also be redacted.
Meeting dates, times, and structural markers
Timestamps, turn numbering, and other structural metadata are kept so the AI can reason about timing and flow. These don’t identify individuals on their own.
Before and after
A representative example of a short 1:1. Every name — the speakers (Sarah Chen, Daniel Park) and people mentioned but not present — is reduced to consistent initials; other sensitive identifiers become opaque tokens.
What you upload
Sarah Chen: Daniel, I’m worried about hitting the Q3 number. Mark from sales said his contact Alex is pushing back hard, and our customer called me on +1-415-555-0142 to complain. My employee ID for the escalation is EMP-4421. Daniel Park: Tell me more about what Mark said.
What the AI sees
SC: <DP>, I’m worried about hitting the Q3 number. <M> from sales said his contact <A> is pushing back hard, and our customer called me on <PHONE_NUMBER_1> to complain. My employee ID for the escalation is <EMPLOYEE_ID_1>. DP: Tell me more about what <M> said.
What happens at the AI provider
- Requests are sent with store=false on OpenAI, which disables OpenAI-side storage of the conversation; OpenAI does not train on API data. On the default tier, OpenAI’s standard 30-day abuse-monitoring retention applies. Zero data retention is available through OpenAI’s enterprise agreement for higher-volume deployments; bring-your-own-endpoint or self-host route around vendor retention entirely.
- Customer organizations can configure their own LLM endpoint per-org (Azure OpenAI, enterprise OpenAI, OpenAI-compatible proxy, or Anthropic-direct) so transcripts route through their own contract instead of ours. Set up by your organization admin in Settings.
- We never send raw transcripts to any third-party service for storage, analysis, or any other purpose.
How your data is stored on our side
- Customer data lives in a managed Postgres database (Railway managed Postgres on Google Cloud Platform, US West region). Customer content is encrypted at rest at the application layer with AES-256-GCM under a per-organization key; on our hosted service the root key is held in AWS KMS (adding decrypt audit + key revocation). Standard cloud-managed AES-256 volume encryption applies to the database and all backups beneath that.
- Each organization's data is logically isolated via row-level tenancy keys (organization_id) and enforced at every API endpoint. Coachees outside an organization see only their own data.
- Raw transcript text and the redacted version (the exact text sent to the AI) are stored separately. You can view the redacted version on each analysis result.
- A reversible mapping of placeholder tokens to original values is kept securely — encrypted at rest, configurable per-organization — so admins can audit exactly what was redacted on any given run.
- You can delete a transcript and its analyses at any time from the run detail page.
Connecting your meeting tools (Zoom, Google Meet)
Instead of uploading each transcript by hand, you can connect a meeting account once and Myc will import new transcripts from meetings you host automatically. Connecting is optional, always your choice, and you can disconnect at any time. Here is exactly what Myc accesses, and why.
Google Meet
When you connect a Google account, Myc requests:
- Your Meet transcripts — to read the text transcript of Meet meetings you host, so there is something to analyze. Text only — never audio or video, and never meetings you don’t host.
- Your upcoming calendar events (read-only) — only if you turn on automatic transcription. Myc looks at your upcoming meetings solely to find the Meet meetings you organize, so it can switch Google’s transcript generation on for them. This is used in the moment and not stored, and Myc never writes to your calendar.
- Your meeting settings — with your permission, Myc turns on the “generate transcript” setting for your own meetings, and nothing else — it does not touch recording or notes.
- Your basic Google profile (name, email) — to identify your account and match your own voice in the transcript.
Zoom
When you connect a Zoom account, Myc reads:
- Your cloud-recording transcripts — the audio-transcript text from cloud recordings of meetings you host — text only, never audio.
- Your basic Zoom profile — to identify your account.
Imported transcripts are treated exactly like transcripts you upload yourself: PII-redacted before any AI call, encrypted at rest, and never used to train any model (see the sections above). The access tokens for your connected accounts are encrypted at rest. Google and Zoom user data is used only to provide the coaching features described on this page — we never sell it, never use it for advertising, and never share it with third parties except the AI provider that produces your analysis, and then only the redacted text.
You can disconnect a meeting account any time from Connected Apps in Myc, which stops all future imports. You can also revoke Myc’s access directly from your Google Account (myaccount.google.com/permissions) or your Zoom account settings.
Myc’s use of information received from Google APIs adheres to the Google API Services User Data Policy, including the Limited Use requirements.
Questions or concerns?
For a deeper technical security overview — architecture, encryption, access controls, sub-processors, and compliance posture — see our security overview. If your security or legal team needs additional documentation (data flow diagram, DPA, white paper, SOC 2 status), contact security@makingyourselfclear.com.