Audio Privacy Processing for Speech Detection and Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies lack effective methods to process audio input recordings and address privacy concerns, particularly in public spaces where external sound recordings are necessary, such as for autonomous driving or noise monitoring.

Innovation Solution

An apparatus and method that process audio input recordings by detecting speech using machine-learning algorithms and applying various processing rules to modify or filter out speech, ensuring privacy while allowing non-speech components to remain usable.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If speech detection and filtering is implemented to address privacy concerns, then privacy protection is improved, but audio signal processing complexity increases

Engineering Contradiction:
Improveprivacy protectionVSAvoidaudio signal processing complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The system performs speech activity detection and processing rule application in real-time during audio recording, rather than as post-processing. The processor continuously monitors incoming audio signals, detects speech portions, and applies appropriate processing rules (masking, deletion, or preservation) immediately, ensuring privacy protection is built into the recording process itself rather than added afterward.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies different processing rules to different portions of the audio signal based on local characteristics. Speech portions are identified and treated differently from non-speech portions, with each speech segment evaluated individually against processing rules to determine whether to mask, delete, or preserve it, rather than applying a uniform processing approach to the entire audio signal.

Inventive Principle:
Principle #3Local quality

2Object-affected harmful factors

If speech portions are masked or removed to protect privacy, then privacy protection is improved, but information loss increases

Engineering Contradiction:
Improveprivacy protectionVSAvoidaudio information loss
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The system dynamically adjusts the level of processing applied to speech portions based on contextual evaluation. Rather than consistently masking or removing all speech, the processor evaluates each speech portion against configurable processing rules and selectively applies masking, deletion, or preservation based on the specific context, allowing the system to adapt its privacy protection level to each situation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system modifies audio signal parameters selectively based on speech detection results. When speech is detected, the processor changes parameters such as amplitude (for masking) or presence (for deletion) only for the detected speech portions, while leaving non-speech portions unchanged, thereby minimizing overall information loss while maintaining privacy protection.

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If low resolution recording is used to ensure privacy, then privacy protection is improved, but recording quality deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoidrecording quality
Core Design Contradiction:
Object-affected harmful factorsVSManufacturing precision

Solution Approach 1:

The audio signal is segmented into speech portions and non-speech portions through speech activity detection. This segmentation allows the system to apply privacy protection measures selectively only to speech portions while preserving the full quality of non-speech portions, rather than degrading the entire recording to low resolution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts speech portions from the audio signal using speech activity detection and applies processing rules specifically to these extracted portions. By separating speech from non-speech content, the system can protect privacy in speech portions while maintaining high recording quality for all other audio content.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP4323901B1Apparatus and method for processing an audio input recording to obtain a processed audio recording to address privacy issues
Publication Date: 2025.01.15 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP4323901B1 patent drawingFigure 1
  • EP4323901B1 patent drawingFigure 2
  • EP4323901B1 patent drawingFigure 3

AI summary

An apparatus for processing an audio input recording to obtain a processed audio recording according to an embodiment is provided. The apparatus comprises an input interface (110) for receiving a plurality of audio input portions of the audio input recording. Moreover, the apparatus comprises a processor (120) for processing a plurality of audio input portions of the audio input recording to obtain a processed audio recording. The processor (120) is configured to determine, whether or not an audio input portion of the plurality of audio input portions comprises speech. If the processor (120) has detected that the audio input portion comprises speech, the processor (120) is configured to generate the processed audio recording by modifying the audio input portion to obtain a modified audio portion, and by generating the processed audio recording such that the processed audio recording comprises the modified audio portion instead of the audio input portion. Or, if the processor (120) has detected that the audio input portion comprises speech, the processor (120) is configured to generate the processed audio recording, such that the processed audio recording does not comprise the audio input portion.