Real-Time Voice Anonymization with Local Embedding Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice anonymization technologies are not performed in real time, are less secure due to cloud-based implementations, and lack deepfake detection capabilities within typical audio stacks on computing devices.

Innovation Solution

The implementation of a system that anonymizes a speaker's voice in real time on the speaker's computing device, using a voice anonymization module that generates a transformed voice and a voice embedding for secure transmission and subsequent voice recovery on the listener's device, while also incorporating deepfake detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If cloud-based voice anonymization is used, then voice processing capability is improved, but security deteriorates due to transmission over wireless connections

Engineering Contradiction:
Improvevoice processing capabilityVSAvoidsecurity
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

Instead of anonymizing voice in the cloud and then recovering it, the patent inverts the approach by extracting voice embeddings locally on the user's device before transmission. This ensures the actual voice data never leaves the device, while still enabling cloud-based processing capabilities. The embedding extraction is performed using a neural network model running locally, transforming the voice into a compact representation that preserves speaking style but cannot be used for deepfakes without the original voice file.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent extracts only the essential voice characteristics (embeddings) from the full voice signal, separating the identifying features from the speaking style. This extraction is performed locally using a neural network that captures speaker identity and style in a compressed form. The extracted embeddings are then transmitted to the cloud for processing, while the original voice data remains secure on the user's device.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If real-time voice anonymization is implemented, then privacy protection is improved, but system complexity increases

Engineering Contradiction:
Improveprivacy protectionVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces complex real-time voice processing systems with a more efficient approach using pre-trained neural network models for embedding extraction. Instead of implementing full voice anonymization pipelines in real-time, the system uses lightweight embedding models that can rapidly transform voice signals into secure representations. This substitution of heavy mechanical processing with optimized neural network inference reduces system complexity while maintaining real-time performance.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent performs voice embedding extraction as a preliminary action before voice transmission or processing. By preparing the voice embeddings in advance on the user's device, the system avoids the need for complex real-time anonymization processing during active communication. The pre-extracted embeddings can be quickly transmitted and used for various purposes without requiring intensive processing during the actual voice interaction.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If voice embeddings are transmitted securely, then voice recovery capability is improved, but transmission security requirements increase

Engineering Contradiction:
Improvevoice recovery capabilityVSAvoidtransmission security
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent changes the parameter representation of voice data from raw audio signals to compressed embeddings. This transformation reduces the amount of data that needs to be transmitted securely while preserving the essential voice characteristics needed for recovery. The embeddings are a compact mathematical representation that captures speaker identity and style, requiring fewer security measures for transmission compared to full-resolution audio data.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a simplified copy of the voice data in the form of embeddings rather than transmitting the original voice files. These embedding copies contain sufficient information for voice recovery and analysis but are much smaller and less valuable to attackers. The copying process transforms the voice into a form that is easier to transmit securely while maintaining the ability to recover and reconstruct the original voice characteristics.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250124171A1Secure real time voice anonymization and recovery
Publication Date: 2025.04.17 INTEL CORP
  • US20250124171A1 patent drawing
  • US20250124171A1 patent drawing
  • US20250124171A1 patent drawing

AI summary

Voice anonymization systems and methods are provided. Voice anonymization is done on the speaker's computing device and can prevent voice theft. The voice anonymization systems and methods are lightweight and run efficiently in real time on a computing device, allowing for speaker anonymity without diminishing system performance during a teleconference or VoIP meeting. The anonymization system outputs a transformed speaker voice. The anonymization system can also generate a voice embedding that can be used to reconstruct the original speaker voice. The voice embedding can be encrypted and transmitted to another device. Sometimes, the voice embedding is not transmitted and the listener receives the anonymized voice. Systems and methods are provided for the detection of voice transformations in received audio. Thus, a listener can be informed whether the speaker voice output from the listener's computing device is the original speaker's voice or a transformed version of the original speaker voice.