Far-Field Voice Anonymization for Private Remote ASR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice control devices like smart speakers and voice assistants face privacy concerns as they offload voice data to remote computers, allowing these systems to gather personally identifiable information such as identity, gender, age, and emotional state, which is not intended for processing.

Innovation Solution

Implement a system with a first and second module to anonymize voice data before transmission to remote ASR modules, using audio transformations to mask sensitive information, ensuring only relevant data is processed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice data is transmitted to remote computers for processing, then speech recognition capability is improved, but user privacy is compromised due to extraction of personally identifiable information

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprivacy violation
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent applies audio transformations (pitch shifting, time stretching, spectral modifications) to the voice data before transmission to remote ASR services. This preliminary anonymization processing removes personally identifiable information while preserving the speech content, ensuring that the remote processor cannot extract identity, gender, age, or emotional state data while still performing accurate speech recognition.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediate processing layer between the local voice capture and remote ASR transmission. This intermediary module applies transformations that act as a mediator, converting the voice data into an anonymized form that maintains speech recognition functionality while blocking the extraction of sensitive personal attributes. The transformed data serves as an intermediary representation that preserves linguistic information but removes biometric information.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If audio transformations are applied to anonymize voice data, then privacy protection is improved, but speech recognition accuracy may deteriorate

Engineering Contradiction:
Improveprivacy protectionVSAvoidspeech recognition accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent carefully selects and applies audio transformations that modify specific parameters of the voice signal (pitch, temporal characteristics, spectral properties) while preserving other parameters critical for speech recognition. By changing only the parameters related to speaker identification and leaving the linguistic content parameters intact, the system achieves anonymization without sacrificing speech recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12585821B2Voice privacy for far-field voice control devices that use remote voice services
Publication Date: 2026.03.24 AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
  • US12585821B2 patent drawing
  • US12585821B2 patent drawing
  • US12585821B2 patent drawing

AI summary

A system includes a first module and a second module. The first module may be configured to perform operations including generating voice data based on an input audio, anonymizing the voice data by applying a first audio transformation, and transmitting the anonymized voice data to a first remote ASR module for generating speech recognition data. The second module may be configured to perform operations including separating the input audio into a first data and a second data, anonymizing the first data by applying a second audio transformation to the first data, generating an anonymized audio data by combining the anonymized first data and the second data, and transmitting the anonymized audio data to a second remote ASR module for generating speech recognition data.