Segmented Pitch Randomization for Edge Audio Anonymization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio anonymization techniques fail to effectively protect speaker identities while maintaining audio usability, as they either obfuscate rather than anonymize data or require significant computational resources, and existing DSP-based methods offer lower privacy and utility compared to ML-based algorithms.
Innovation Solution
A lightweight DSP-based audio anonymization algorithm employing pitch shifting and segmentation with probabilistic pitch randomization to enhance privacy and utility, using a base pitch generation with probability to prevent determinism and segment audio signals to increase recognition difficulty.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If speech-to-text conversion is used to protect privacy, then computing power consumption is reduced, but conversion accuracy deteriorates and additional information is lost
Solution Approach 1:
The patent employs a lightweight DSP-based anonymization algorithm that can be executed on edge devices with limited resources, replacing the need for heavy ML-based speech-to-text conversion. This disposable-like approach processes audio directly without requiring complex model inference, reducing computing power consumption while maintaining privacy protection.
Solution Approach 2:
The patent substitutes ML-based speech-to-text conversion with a DSP-based audio processing system. Instead of converting audio to text through complex neural networks, the system directly processes audio signals using digital signal processing techniques (pitch detection, pitch shifting) to achieve anonymization, replacing the mechanical conversion process with a more efficient signal processing approach.
2Reliability
If ML-based algorithms are used for audio anonymization, then privacy and utility are improved, but algorithm complexity, size and latency increase
Solution Approach 1:
The patent replaces complex ML-based anonymization algorithms with a lightweight DSP-based approach that can run on resource-constrained edge devices. The system uses simple pitch detection and pitch shifting operations instead of complex machine learning models, significantly reducing algorithm complexity and device requirements while maintaining privacy protection effectiveness.
Solution Approach 2:
The patent substitutes ML-based audio anonymization with a DSP-based system that uses traditional signal processing techniques. Instead of relying on complex neural networks and large model sizes, the system employs pitch detection algorithms and pitch shifting operations that are computationally efficient and have low latency, replacing the heavy ML infrastructure with lightweight signal processing.
3Reliability
If pitch shifting is applied to anonymize audio, then speaker identity is protected, but audio quality may deteriorate
Solution Approach 1:
The patent applies pitch shifting by modifying the frequency parameter of the audio signal. By changing the pitch (fundamental frequency) of the speaker's voice, the system protects speaker identity while maintaining audio quality. The pitch shifting is performed in a controlled manner to preserve the naturalness and intelligibility of the speech.
Solution Approach 2:
The patent performs pitch detection and pitch shifting as preliminary processing steps before any further audio processing or transmission. By pre-anonymizing the audio through pitch modification, the system ensures speaker identity protection is established early in the processing chain, allowing subsequent processing to focus on maintaining audio quality without compromising privacy.
Data Source
AI summary
A method and an electronic device for generating an anonymized audio output are provided. The method, executable by the electronic device, comprises acquiring an audio recording of a speaker; stochastically determining a base pitch value based on at least a first probabilistic function; segmenting the original audio input into a plurality of audio segments, each of the plurality of audio segments being associated with a respective pitch. For each audio segment, the method further comprises generating a pitch adjustment value using a combination of the base pitch value of the segment and a value determined using a second probabilistic function; generating an adjusted audio segment by adjusting the pitch of the audio segment using the pitch adjustment value, the adjusted audio segment having an adjusted pitch that is different from the original pitch; generating the anonymized audio output by combining the adjusted audio segments.


