
At some point, probably recently, you said something near a smart speaker or a phone and received an advertisement for it shortly after. Maybe you dismissed it as coincidence, or as the algorithm being eerily good at predicting your interests. But the less comfortable explanation is also the more accurate one: the devices and services you interact with every day are collecting a great deal of data about you, and most people have never looked at the settings that govern what happens to it.
The question of whether your AI assistant is “spying” on you depends on how you define surveillance. None of the major AI platforms are secretly listening to your conversations without any activation trigger. The more precise concern is subtler and considerably more widespread: every interaction with an AI assistant — every query, every voice command, every completed conversation — is being collected, stored, used to train models, and in many cases shared with affiliated companies, advertising partners, or third-party analytics providers. The extent of that collection, and what you can do to limit it, varies by platform and requires active configuration. The default settings across virtually every major AI assistant are calibrated in favour of data collection, not privacy.
Research published in PMC’s peer-reviewed literature identifies AI assistants and IoT voice devices as operating on always-listening architectures that create a fundamental privacy vulnerability through accidental activation — with voice systems estimated to trigger approximately once per hour through misheard commands, capturing unintended recordings without user awareness. That figure has practical significance beyond the privacy concern: if your voice assistant is capturing unintended audio, what it captures, where that audio goes, and how long it’s retained are all questions that the following audit addresses.
What the Federal Government Has Said About AI Data Collection
Before getting into the platform-specific settings, understanding the regulatory context is important for calibrating how seriously to take the audit — because the answer is: more seriously than most people do.
The Federal Trade Commission issued orders to seven companies providing consumer-facing AI-powered chatbots in 2025, seeking detailed information about data collection practices, model training, retention policies, and safeguards designed to protect minors. In September 2025, the FTC used its Section 6 investigatory authority to issue formal information-gathering orders to companies offering generative AI companion products and services, specifically directing them to account for their advertising, safety, and data handling practices. The signal from the agency is clear: AI data collection practices are under active federal scrutiny, and the practices that regulators are examining are the same ones most consumers have never reviewed in their own settings.
The National Institute of Standards and Technology’s Generative AI Profile — NIST-AI-600-1, published July 26, 2024 and still the substantive federal AI risk management baseline in 2026 — identifies privacy and data management as core risk categories for AI systems, and in December 2025 NIST released a preliminary draft CSF Profile for AI specifically designed to help organisations manage cybersecurity risks uniquely associated with AI systems. NIST frameworks are not legally binding on consumers, but they establish the standard of care that well-governed AI systems should meet — and comparing that standard against how your AI assistant actually handles your data reveals the gap most consumer products are still operating in.
The Children’s Online Privacy Protection Act — COPPA — applies directly to AI assistants used in households with children under 13, requiring parental consent for data collection and strict handling obligations. The FTC’s expanded COPPA scrutiny in 2025 and 2026 specifically targeted AI chatbot and companion products’ practices around minors, reflecting the agency’s sensitivity to the particular vulnerability of children’s data in AI training contexts. If children use any of the AI assistant products in your home, COPPA compliance is a concrete parental concern, not just a regulatory abstraction.
The Platform-by-Platform Privacy Picture in 2026
The five major AI assistant platforms — ChatGPT, Google Gemini, Microsoft Copilot, Amazon Alexa, and Apple Siri — have meaningfully different privacy architectures, and the differences matter practically for what the audit finds and what settings you can change.
Amazon Alexa occupies the most vulnerable position in the 2026 privacy landscape. Alexa sends every voice command to the cloud — and critically, Amazon removed Alexa’s last local-only voice processing option in 2026, meaning there is no configuration available that keeps voice queries on-device. Every question you ask Alexa, every command, and every ambient audio captured during accidental activations travels to Amazon’s servers. Amazon uses this data for product improvement, and it has disclosed data-sharing practices with advertising partners. The audit step for Alexa users is specific and consequential: navigate to the Alexa app → More → Settings → Alexa Privacy → Review Voice History, where you can delete individual recordings and configure automatic deletion. Disable cloud storage of voice recordings by turning off “Use Voice Recordings” and “Use Messages to Improve Transcriptions” in the Alexa Privacy settings. These settings don’t prevent Amazon from receiving the audio — they limit what’s retained after processing.
Apple Siri has historically been the most privacy-protective major voice assistant through its on-device processing architecture. The 2026 picture is more complicated. Siri’s Apple Intelligence overhaul introduced a Gemini-powered backend — Apple’s partnership with Google brings Gemini’s reasoning into Siri as an optional backend, triggered when a query exceeds what Apple Intelligence can handle locally. Users can also invoke ChatGPT for open-ended queries. The result is that Siri handles personal context with Apple’s on-device privacy architecture, but escalates to Gemini or ChatGPT when queries require it. This means a query that begins on your device may end in Google’s or OpenAI’s data infrastructure, depending on its complexity. The audit step: Settings → Siri & Search → disable “Share with App Developers”; then Settings → Siri & Dictation History to review and delete stored queries. For the Gemini and ChatGPT handoff features, review the individual settings for those services as separate platforms.
Google Gemini occupies a structural tension that the platform’s own documentation reflects honestly. Gemini’s privacy framework includes strict access rules, region-aware storage, and enterprise-grade encryption — but in the standalone app and API consumer tiers, training use and long retention windows quietly persist by default. The audit step for Google’s AI products runs through myaccount.google.com → Data & Privacy → My Activity → Web & App Activity. Specifically, disable “Gemini Apps Activity” to stop conversations from being stored and used for model improvement. Google’s data infrastructure is interconnected across its product suite, meaning Gemini activity that isn’t explicitly opted out of can inform targeting and personalisation across Search, YouTube, and advertising products.
OpenAI/ChatGPT has made its data management settings more accessible than most competitors — you can delete chats with a 30-day delay, export your data, and opt out of model training from the Settings menu. The key action is: Settings → Data Controls → “Improve the model for everyone” → disable. Enterprise users are excluded from training by default. The “30-day delay” on deletions is the detail worth noting: content you delete from your view continues to exist in OpenAI’s systems for 30 days before actual deletion. Disabling training participation limits the use of your conversations for model improvement but does not eliminate data retention for safety and legal compliance purposes.
Microsoft Copilot integrates with the broader Microsoft 365 ecosystem, meaning that for users with Microsoft accounts, Copilot interactions can intersect with data stored across Outlook, OneDrive, Teams, and other Microsoft products. The audit step: privacy.microsoft.com → Copilot conversation history → clear; then account settings → AI improvement → off. All major voice assistants encrypt data transmission using TLS 1.3 or similar protocols, and voice recordings are encrypted both in transit and at rest on cloud servers — but encryption during transmission does not limit what the platform does with that data once it arrives at their servers. Encryption is a security property, not a privacy property.
The Structural Issue Your Settings Can’t Fix
Working through all the settings above produces meaningful improvements to your AI privacy posture. It does not resolve the deeper structural issue that regulators and researchers are increasingly focused on.
When you disable AI training participation on any of these platforms, you’re limiting one use of your data. You’re not necessarily limiting collection, storage, retention for safety compliance, sharing with affiliated entities, or the data that flows through your queries into whatever third-party services the AI calls to fulfil your requests. The FTC has examined AI companion product data handling practices specifically because the data collected through these products — including emotional context, relationship patterns, and personal disclosures that users share in conversation — represents a category of personal information whose commercial use raises concerns that standard privacy policy disclosures don’t adequately address. The concern is not hypothetical. It’s the subject of active federal investigatory activity.
The broader data economy into which AI assistant data flows is explored in our analysis of the data rights economy and who controls the information your devices generate. The smart home security dimension — including how connected devices create data collection endpoints throughout the home environment — is mapped in our piece on how IoT devices create security and surveillance vulnerabilities. And the attention and engagement dimension of AI assistant design — the ways these platforms are built to maximise interaction frequency in ways that increase data collection — connects to the analysis in our piece on how tech platforms compete for your focus.
What the settings audit achieves is practical and meaningful: it reduces the volume of data being retained and used for purposes you haven’t actively chosen, it limits model training participation, and it gives you a defensible understanding of what you’ve consented to. Privacy is not a one-time setting — it requires routine review because manufacturers frequently modify default settings, often at software update moments when users are clicking “agree” on update notifications without reviewing what’s changed. The quarterly audit cadence recommended by security researchers is the right operating tempo: once per year misses policy changes; more frequently is unnecessary for most users.
Frequently Asked Questions
No AI assistant from a major provider is secretly recording everything you say continuously. All major voice assistants — Alexa, Siri, Google Assistant — are designed to activate only after detecting their wake word. The documented concern is accidental activation, which research estimates occurs approximately once per hour through misheard phrases that the device interprets as a wake word. During accidental activations, audio is captured and transmitted to the platform’s servers. Additionally, generative AI assistants like ChatGPT, Gemini, and Copilot retain all text conversations you actively submit — the surveillance concern there is not passive background listening but the retention and use of voluntary interactions you may not have considered carefully. The most honest framing is that these devices are not spying in a traditional sense, but they are collecting considerably more of your data than most users realise or have actively chosen to share.
In the ChatGPT web app or mobile app, navigate to Settings → Data Controls → locate the toggle for “Improve the model for everyone” → disable it. Once this setting is off, your future conversations will not be used for model training. Note that ChatGPT’s conversation deletion has a 30-day delay — when you delete a chat, it disappears from your view but remains in OpenAI’s systems for 30 days before actual deletion. OpenAI retains conversations for safety and legal compliance purposes regardless of training opt-out status. Enterprise and API users are excluded from training data use by default without needing to change settings. The FTC has investigated AI chatbot data handling practices and issued formal information-gathering orders to AI companion product companies specifically regarding these data handling questions, signalling that training data consent is an active area of regulatory scrutiny.
Siri maintains the strongest privacy posture in 2026 through Apple’s on-device processing architecture for personal context queries, though its privacy position has become more complex following the 2026 Apple Intelligence overhaul. Siri now escalates queries beyond its on-device capability to a Gemini backend (through Apple’s partnership with Google) and optionally to ChatGPT — meaning complex queries may be processed by Google’s or OpenAI’s data infrastructure rather than entirely on-device. Alexa is the least private major voice assistant in 2026: Amazon removed its last local-only processing option, meaning every voice command travels to Amazon’s cloud servers. Google Gemini offers strong enterprise-grade encryption and access controls, but consumer-tier accounts retain training use rights and long retention windows by default unless explicitly disabled. For users who prioritise keeping data off the cloud, Apple hardware with Apple Intelligence configured for on-device-only processing provides the most control — at the cost of reduced capability for complex queries.
The NIST AI Risk Management Framework — specifically the Generative AI Profile NIST-AI-600-1, published July 26, 2024 and still the substantive federal AI risk-management baseline in 2026 — is a voluntary guidance framework published by the National Institute of Standards and Technology to help organisations identify, assess, and manage risks associated with AI systems, including privacy and data management risks. It is not legally binding on private companies or directly enforceable by consumers. In December 2025, NIST released a preliminary draft Cybersecurity Framework Profile for AI systems, with final guidance expected through 2026. The NIST frameworks establish the standard of care that well-governed AI systems should meet — comparing a platform’s practices against NIST guidance is a way to evaluate whether the company is taking AI risk seriously. Consumer protection from AI data practices comes primarily from the FTC’s authority over deceptive and unfair practices, state-level privacy laws, and COPPA for children’s data, rather than from NIST directly.
Security researchers and privacy specialists recommend auditing AI and voice assistant privacy settings quarterly — every three months — rather than once at setup and never again. The reason is structural: manufacturers frequently modify default settings at software update moments, and new data-sharing features or policy changes are often rolled out with defaults favouring collection rather than privacy. Operating system updates, app version updates, and account policy changes can all reset or alter previously configured privacy settings. The quarterly audit cycle — checking training participation settings, reviewing stored conversation history, clearing voice recording archives, and reviewing active permissions for each AI platform you use — takes less than twenty minutes total across all platforms and ensures that the settings you configured six months ago still reflect the privacy posture you intended. Set a calendar reminder for the same week each quarter.
The Bottom Line
The honest answer to “is my AI assistant spying on me?” is: not in the way a human spy would, but it is collecting, retaining, and using your data in ways that most people have never reviewed and many would find surprising if they did. The platforms are not malicious. They are commercial entities whose business models depend on data, and whose default configurations reflect that dependency.
The FTC’s active scrutiny of AI chatbot data handling practices, its formal information-gathering orders to seven AI companies in 2025, and its expanded COPPA enforcement in AI contexts all signal that the regulatory environment recognises what individual users are still catching up to: AI assistants collect personal data at a scale and intimacy level that warrants the same deliberate attention people give to financial and medical privacy. The settings audit in this guide takes about twenty minutes. Running it quarterly takes about twenty minutes a year, divided across four sessions. That’s a modest investment for meaningful improvement in what you’ve actually consented to share with the AI systems you use every day.
This article is for informational purposes only. Privacy settings, data retention practices, and platform policies change frequently — always verify current settings in your specific platform’s privacy documentation before assuming prior configurations remain in effect.




