Neural Speech Masking Signals for Adaptive Privacy Protection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech masking technologies generate fixed masking signals for different speakers and speech contents, limiting their applicability and user experience in various scenarios.

Innovation Solution

A neural network-based approach is used to dynamically generate masking signals tailored to specific speech contents and masking effects, utilizing a training process that considers factors like speech intelligibility, recognition accuracy, and comfort level, enabling adaptable masking in diverse scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a fixed masking signal generation method is used for different speakers and speech contents, then the system complexity is reduced and the implementation is simplified, but the adaptability to different scenarios and speech contents deteriorates, resulting in poor masking effects

Engineering Contradiction:
Improveadaptability to different scenariosVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies dynamics by replacing the fixed masking signal generation method with a dynamic neural network model that adapts to different speakers and speech contents in real-time. The model processes input speech features and generates customized masking signals based on the specific characteristics of each speech instance, enabling the system to adapt to various scenarios while maintaining manageable complexity through automated processing

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the masking signal generation process by using a neural network model that adjusts masking signal characteristics based on input speech features. Instead of using fixed parameters, the system dynamically modifies masking signal properties (such as frequency content, temporal characteristics, and signal strength) according to the detected speech content, thereby improving adaptability without requiring overly complex manual configuration

Inventive Principle:
Principle #35Parameter changes

2Reliability

If a fixed masking signal generation method is used, then the ease of operation is improved, but the masking effectiveness for different speech contents deteriorates

Engineering Contradiction:
Improvemasking effectivenessVSAvoidease of operation
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent applies self-service by enabling the system to automatically analyze speech contents and generate appropriate masking signals without requiring manual intervention or complex user configuration. The neural network model self-adjusts to different speech patterns and automatically optimizes masking effectiveness for each speech instance, maintaining ease of operation while significantly improving masking reliability across different scenarios

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback mechanisms where the system continuously monitors and evaluates the masking effect based on speech content characteristics. The neural network model uses feedback information about the input speech to iteratively refine and adjust the generated masking signals, ensuring optimal masking effectiveness while keeping the user interface simple and easy to operate

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250336384A1Speech masking method and system, electronic device, and non-transitory computer readable storage medium
Publication Date: 2025.10.30 AAC ACOUSTIC TECH (SHANGHAI) CO LTD
  • US20250336384A1 patent drawing
  • US20250336384A1 patent drawing
  • US20250336384A1 patent drawing

AI summary

The disclosure relates to the field of communication security and discloses a speech masking method and system, an electronic device, and a non-transitory computer readable storage medium. In the disclosure, after the target speech is obtained, the target speech is not masked according to the traditional fixed masking method, but the target masking effect is determined in advance, and the target masking effect can be determined according to different requirements. After that, the neural network model is trained according to different target masking effects, and the neural network model trained according to the different target masking effects can dynamically provide different masking signals for the target speech. In this way, different masking signals can be generated for different target speeches according to different needs, more scenarios can be applied, and good masking effects can be obtained, thereby improving user experience.