A speech privacy protection method and device based on high-frequency acoustic signal mixing

By constructing a personalized corpus to generate high-frequency interference noise and combining it with automatic interference control, the high cost and vulnerability of existing voice assistant privacy protection technologies are solved, achieving effective voice privacy protection and user-friendliness.

CN119864051BActive Publication Date: 2025-10-24HUAZHONG UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411902920.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-10-24
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing voice assistant privacy protection technologies rely on dedicated ultrasonic equipment, which is costly and vulnerable to attack. Continuous interference affects the user experience, and existing noise reduction solutions are difficult to effectively prevent unauthorized recording.

Method used

A voice privacy protection method based on high-frequency sound signal mixing is adopted. A personalized corpus is constructed by collecting user voice data, generating and modulating high-frequency interference noise, and combined with an automatic interference control strategy, the nonlinear characteristics of commercial microphones are utilized to detect voice activity in real time and inject interference signals.

Benefits of technology

It effectively prevents unauthorized recording, reduces the likelihood of eavesdroppers capturing clear speech, reduces energy consumption, improves user experience, eliminates reliance on dedicated equipment, and enhances noise diversity and resistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119864051B_ABST
    Figure CN119864051B_ABST
Patent Text Reader

Abstract

The application discloses a speech privacy protection method and device based on high-frequency sound signal mixing, first, user speech data is collected to build a personalized user corpus; then, the user customizes the scene requirement, selects the text theme of speech synthesis, and expands the user corpus; then, the user selects the noise theme related to the interference from the selected text theme according to the specific life situation, randomly extracts the speech fragment from the user corpus corresponding to the noise theme to generate the interference noise; the interference noise is modulated to the preset frequency band interval; finally, when it is detected that the user is talking, the controller sends the modulated interference signal to the jammer, and the jammer injects the interference signal into the nearby microphone. The application has practicability and safety, can effectively protect the user privacy, and can resist attacks of various de-noising technologies such as speech enhancement and blind source separation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of information security, and relates to a voice privacy protection method and device, in particular to a voice privacy protection method and device based on high-frequency sound signal mixing. BACKGROUND

[0002] Voice assistants are widely used in smart speakers in today's society. These assistants enable users to control smart speakers through voice commands without manual operation. With its convenience, voice assistants have gradually become an indispensable part of many people's daily lives. However, in order to ensure real-time response, voice assistants usually capture and record all incoming audio signals and stream them to the cloud for processing and response. Once a user's voice data is recorded by a voice assistant, the control and flow of this data are difficult to grasp and trace, which raises serious privacy concerns for users. Numerous research reports indicate that voice assistants have certain defects and face the risk of being misused for illegal surveillance.

[0003] In existing voice privacy protection technologies, researchers usually take advantage of the inherent nonlinear characteristics of commercial microphones to achieve privacy protection by playing carefully designed ultrasonic interference signals near the smart speaker. When the interference signal is received by the microphone of the smart speaker, demodulation will occur and audible interference noise will be generated, effectively confusing unauthorized recordings and achieving protection of user privacy. For example, Chinese patent document No. CN202310653928, published on July 4, 2023, describes a voice privacy protection method based on ultrasonic interference. The invention discloses a method, system and device for designing anti-eavesdropping ultrasonic interference samples, which includes: preprocessing a normal human voice data set; obtaining the audio transfer function of the digital domain audio modulated by ultrasonic waves and demodulated by an eavesdropping microphone; obtaining the playback length of the ultrasonic interference sample; establishing a target function to optimize and train the interference sample, and obtaining the ultrasonic interference sample release time, ultrasonic interference sample optimized amplitude and normal human voice optimized amplitude after training; calculating the mixed audio based on the audio transfer function, ultrasonic interference sample release time, ultrasonic interference sample optimized amplitude and normal human voice optimized amplitude, and inputting the mixed audio into the voice recognition system after amplitude limiting processing; inputting the training set data into the voice recognition system after processing, calculating the average loss function, and performing gradient descent optimization; evaluating the data in the prediction set to complete the design of the ultrasonic interference sample.

[0004] Although the ultrasonic method takes into account effectiveness and user-friendliness, its practical application is subject to the requirement of a specific ultrasonic transmitter or array. Existing ultrasonic interference schemes use ultrasonic signals for noise injection, often relying on dedicated ultrasonic transmitters or arrays, which are expensive. Moreover, existing noise schemes have certain limitations. Gaussian white noise can be easily filtered out by attackers using existing noise reduction algorithms, cosine wave noise is difficult to mask the entire speech frequency band, and phoneme-level noise has relatively insufficient actual interference effect. In addition, continuous playback of interference signals not only leads to unnecessary energy consumption, but also affects the voice interaction between users and smart speakers, severely reducing user experience. SUMMARY

[0005] To solve the above problems, the present application provides a speech privacy protection method and device based on high-frequency sound signal mixing based on the insensitivity of the human ear to near-ultrasonic signals and the inherent non-linear defects of commercial microphones.

[0006] The technical scheme adopted by the method of the present application is: a speech privacy protection method based on high-frequency sound signal mixing, comprising the following steps:

[0007] Step 1: Collect user speech data and build a personalized user corpus;

[0008] Step 2: User defines scenario requirements, selects text topics for speech synthesis, and expands the user corpus;

[0009] Step 3: The user selects noise topics related to this interference from the text topics selected in step 2 according to the specific life situation it is in, and randomly selects speech segments from the user corpus corresponding to the noise topics to generate interference noise;

[0010] Step 4: Modulate the interference noise to a preset frequency band interval;

[0011] Step 5: When it is detected that the user is talking, the controller sends the modulated interference signal to the jammer, and the jammer injects the interference signal into the nearby microphone.

[0012] As a preferred embodiment, the specific implementation process of step 2 is to use speech synthesis technology to expand the user's personal corpus. First, audio samples from different speakers are collected, and then emotion category labels are added to the collected data.

[0013] As a preferred embodiment, the text encoder and duration predictor of the VITS model are optimized so that they can accept speaker ID and emotion category embeddings;

[0014] The optimization of the text encoder is to add two embedding layers for speaker ID and emotion category at the input layer of the text encoder, which respectively convert the discrete speaker ID and emotion category into continuous vector representation, and the word embedding of the text is spliced with the embedding vectors of the speaker ID and emotion category.

[0015] The optimization of the duration predictor is to increase the number of layers and the number of nodes in each layer of the duration predictor according to the new input dimension, increase the original number of layers from 2 to 4, and increase the number of nodes in each layer from 256 to 512, to ensure that the model can effectively process these additional embedding information and enhance the expression ability of the model.

[0016] As a preferred, the specific implementation of step 3 includes the following steps:

[0017] Step 3.1: accelerate the randomly selected speech segment using a random acceleration factor a;

[0018] Step 3.2: remove the speech gap whose energy is continuously lower than a certain threshold from the accelerated speech segment;

[0019] Step 3.3: time-domain superposition of the processed speech segment to generate mixed speech interference noise.

[0020] As a preferred, in step 3.1, the randomly selected speech segment is converted to the frequency domain using Fourier transform; in the frequency domain, each frequency component has its own phase evolution with time, by adjusting the phase information and the step of phase evolution, the length of the signal is shortened; finally, inverse Fourier transform is performed to convert the audio signal to the time domain, thereby obtaining the accelerated audio that maintains the pitch.

[0021] As a preferred, in step 3.2, the removal of the speech gap whose energy is continuously lower than a certain threshold is achieved by analyzing the sample data in the accelerated audio stream frame by frame, calculating the short-time energy of the audio signal, if a certain segment of audio has a duration longer than threshold A and contains all audio frames whose short-time energy is lower than threshold B, then this segment of audio is marked as silent; remove all audio parts marked as silent from the audio stream.

[0022] As a preferred, in step 4, the upper sideband modulation technology is used to modulate the interference noise to the high frequency band of 17kHz to 21kHz.

[0023] As preferred, in step 5, the voice activity detection technology is used for real-time voice activity detection; the controller collects the environmental audio signals in real time, divides the audio signals into frame signals of a preset length, analyzes each frame independently, extracts the energy distribution, spectral tilt and peak value characteristics of each frame to distinguish the characteristics of voice and non-voice; if at least N consecutive frame signals are detected to have voice characteristics, it is determined that there is user voice activity at this time; wherein N is a preset value. As preferred, in step 5, the automatic interference control strategy is used for interference control; the specific implementation includes the following steps:

[0024] Step 5.1: Real-time voice activity detection is performed; when voice activity is detected, the following step 5.2 is executed; otherwise, the following step 5.3 is executed;

[0025] Step 5.2: Start interference & wait for voice command, through the phoneme-level wake-up word detection technology, when the controller identifies the first phoneme in the pre-defined wake-up word phoneme sequence, the controller instructs the interferer to switch to the stop interference state while continuing to monitor the subsequent phonemes until the entire phoneme sequence is identified; if the wake-up word phoneme sequence is not continuously detected, the interferer returns to the interference state;

[0026] Wherein, the trained BiLSTM-CRF model is used for single phoneme recognition, in the training, the features of the current frame and the features of the previous and subsequent 10 frames are combined together as the features of the frame for training, after training, the wake-up word data set and a part of the original voice data set are combined to fine-tune the model;

[0027] Step 5.3: Determine whether interference is being performed; if yes, stop the interference; if no, do nothing; determine whether the voice assistant uses follow-up mode or non-follow-up mode; when the voice assistant uses non-follow-up mode, use voice activity detection and a preset silence period strategy for interference recovery; when the voice assistant uses follow-up mode, once the voice assistant responds to the user's instruction, the controller resumes monitoring the voice activity and resumes the interference when voice activity is detected again.

[0028] The technical scheme adopted by the device of the application is: a voice privacy protection device based on high-frequency sound signal mixing, comprising:

[0029] One or more processors;

[0030] A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the voice privacy protection method based on high-frequency sound signal mixing.

[0031] Compared with the prior art, the beneficial effects of the application include:

[0032] (1) The present application uses speech synthesis technology to expand the user's personal corpus, and the user does not need to access the Internet during operation, thus eliminating the risk of user voice privacy leakage. In addition, the user can choose the text theme of speech synthesis according to his specific situational needs, such as economy, culture, and daily life, to better disturb the eavesdropper.

[0033] (2) The present application removes gaps from speech segments used to compose noise, not only effectively reducing long periods of silence and reducing the possibility of eavesdroppers capturing clear speech, but also increasing the effective components of interference noise, thereby improving the overall interference effect.

[0034] (3) The present application superimposes user speech segments in the time domain to generate mixed speech interference noise. Increasing the number of simultaneously superimposed speech segments can produce more complex noise signals, but in the case of maintaining the same signal-to-noise ratio, it also dilutes the energy allocated to each segment, thereby reducing the overall interference effect. At the same time, using too few superimposed signals can obtain stronger interference effects, but also make the noise more easily understood, vulnerable to attacker speech separation models, and thus potentially expose the original speech.

[0035] (4) The present application accelerates, gap-removes, and then time-domain superimposes randomly extracted speech segments, significantly improving the interference effectiveness of mixed speech noise.

[0036] (5) The present application generates interference noise based on the user corpus, which can more effectively confuse unauthorized recordings and has stronger resistance to various denoising models such as speech enhancement and blind source separation. And the present application expands the user corpus through speech synthesis means, increasing the diversity of noise.

[0037] (6) The present application eliminates the dependence on special ultrasonic devices: the present application uses near-ultrasonic waves of 17 kHz to 21 kHz to inject interference noise, allowing any commercial speaker with wireless communication capabilities to be used as an interferometer without the need for a special ultrasonic transmitter or array.

[0038] (7) The present application controls the activation, suspension, and resumption of interference through the design of an automatic interference control strategy, which effectively prevents unauthorized recordings while preserving normal voice interaction between the user and the smart speaker, achieving automatic interference control. BRIEF DESCRIPTION OF DRAWINGS

[0039] The technical solutions of the present application are further illustrated below using examples and specific embodiments. In addition, some drawings are also used in the process of explaining the technical solutions. For those skilled in the art, other drawings and the intent of the present application can also be obtained from these drawings without creative effort.

[0040] Figure 1 Method flowchart for embodiments of the present application;

[0041] Figure 2 Automatic interference control strategy flowchart for embodiments of the present application. DETAILED DESCRIPTION

[0042] In order to facilitate those skilled in the art to understand and implement the present application, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0043] The present application is mainly based on the privacy security of mobile devices, considering the nonlinear effect of commercial microphones, and proposes a voice privacy protection method based on high-frequency sound signal mixing. The present application makes full use of the nonlinear effect of commercial microphones, introduces a mixed noise interference scheme with strong interference and high robustness, and combines with an automatic interference control strategy, while retaining normal voice interaction between the user and the smart device, effectively preventing unauthorized recording.

[0044] See Figure 1 The voice privacy protection method based on high-frequency sound signal mixing provided by the embodiment includes the following steps:

[0045] Step 1: Collect user voice data and build a personalized user corpus;

[0046] In one embodiment, during the registration process, the user needs to record a small part of the voice sample according to the prompt to build a personalized user corpus. In addition, the user can choose to provide additional voice recording under time permitting to further enrich the personal corpus.

[0047] Step 2: User customizes scenario requirements, selects text topics for voice synthesis, and expands the user corpus;

[0048] In one embodiment, the present embodiment generates interference noise based on the user's corpus, and the diversity of the generated noise is limited by the amount of speech data available to the user. The less speech data in the user's corpus, the more likely the noise will repeat, which can potentially allow an adversary to separate the original speech from the interfered recording. Therefore, to reduce this risk, the present embodiment employs speech synthesis techniques to expand the user's personal corpus. After obtaining IRB approval, the present embodiment collected audio samples from different speakers, totaling approximately 40,000 sentences with an average sentence length of 8 seconds. In the collected data, the present embodiment added emotional classification labels such as rhythm, pause, and intonation. These labels help generate more natural and expressive synthesized speech. The present embodiment modified the text encoder, duration predictor, and flow layer of the VITS model to accept speaker ID and emotion category embeddings. The training process followed the standard procedure provided by VITS. In this way, the trained model can read input text according to different base speakers and generate speech in different emotions. The user only needs to provide a small portion of speech samples, and the present embodiment can extract the unique timbre features of the speaker and apply them to the synthesized speech when generating speech.

[0049] By optimizing the text encoder and duration predictor of the VITS model to accept speaker ID and emotion category embeddings, the present embodiment adds two embedding layers for speaker ID and emotion category to the input layer of the text encoder. These embedding layers convert discrete speaker IDs and emotion categories into continuous vector representations. The word embeddings of the text will be concatenated with the embedding vectors of the speaker ID and emotion category. In this way, the input to the text encoder not only contains text information, but also contains information about the speaker's identity and emotional state. According to the new input dimension, the present embodiment increases the number of layers and nodes in each layer of the duration predictor, increasing the original number of layers from 2 to 4 and the number of nodes in each layer from 256 to 512 to ensure that the model can effectively process these additional embedding information and enhance the model's expression ability. During training, a diverse dataset containing multiple speakers and emotions is used to ensure that the model can learn the impact of different speakers and emotions on generated audio features.

[0050] It is worth noting that the user does not need to access the Internet during operation, thus eliminating the risk of user speech privacy leakage. In addition, the user can choose the text topic of the speech synthesis according to their specific situational needs, such as economy, culture, and daily life, to better disrupt the eavesdropper.

[0051] Step 3: The user selects a noise topic related to the current interference from the text topics selected in Step 2 according to the specific life situation he is in, and randomly extracts a speech segment from the user corpus corresponding to the noise topic to generate interference noise;

[0052] In an embodiment, the implementation process is as follows:

[0053] Step 3.1: This embodiment accelerates the randomly extracted speech segment. This embodiment uses Fourier transform to convert the audio signal to the frequency domain. In the frequency domain, each frequency component has its own phase evolution over time. The phase change between adjacent frames reflects the speed of the signal frequency components. In order to accelerate the audio, this embodiment adjusts the phase information so that the transition between frames is faster. After adjusting the phase information reasonably, this embodiment adjusts the step size of the phase evolution to shorten the time length of the signal without affecting the pitch. Finally, this embodiment performs inverse Fourier transform on the modified signal to convert the audio signal to the time domain, thereby obtaining accelerated and pitch-maintained audio. By accelerating, more meaningless semantic information can be generated in the same time period, thereby enhancing the noisy feeling of the interference noise and improving the interference effect. Through experiments, this embodiment determines that the optimal range of the random acceleration factor a is [1.4, 1.8], and applies it to each speech segment.

[0054] Step 3.2: This embodiment generates interference noise based on the user-registered corpus. However, humans often naturally pause when speaking, which may reveal some original speech. To this end, this embodiment performs gap removal processing on the speech segments used to compose noise, removing speech gaps with energy continuously below a certain threshold. This embodiment analyzes the sample data in the accelerated audio stream frame by frame. For each audio frame, this embodiment calculates the short-time energy of the audio signal. If a segment of audio lasts longer than 0.1s and contains all audio frames with short-time energy below '-40dB', the segment of audio is marked as silent. This embodiment removes all audio parts marked as silent from the audio stream. This processing not only effectively prevents long silent intervals, reducing the likelihood of the eavesdropper capturing clear speech, but also increases the effective components of the interference noise, thereby improving the overall interference effect.

[0055] Step 3.3: This embodiment performs time-domain superposition on the processed user speech segments described above to generate mixed speech interference noise. Increasing the number of simultaneously superimposed speech segments can produce more complex noise signals, but it will also dilute the energy allocated to each segment, reducing the overall interference effect while maintaining the same signal-to-noise ratio. At the same time, using too few superimposed signals can obtain stronger interference effect, but it will also make the noise more easily understood, vulnerable to the attacker's speech separation model, and thus potentially expose the original speech. To balance the interference effectiveness and resistance to speech separation models, this embodiment sets the number of superimposed speech segments to 2.

[0056] Step 4: Modulate the interference noise to the preset frequency band interval;

[0057] In theory, most commercial loudspeakers have a maximum output frequency of 24 kHz. However, in practical applications, these loudspeakers significantly weaken their frequency response above 21 kHz. At the same time, the human ear has almost no perceptual ability for sounds above 17 kHz. To maximize the effective speech components of the demodulated noise, it is necessary to ensure that the part of the noise between 50 Hz and 4 kHz before modulation is modulated to the frequency range that the microphone is relatively sensitive to.

[0058] Therefore, in one embodiment, this embodiment uses the upper sideband modulation technique to modulate the interference noise to the near-ultrasonic frequency band of 17 kHz to 21 kHz. This embodiment uses the user's mobile device as a controller and a commercial audio speaker as an interferer, and deploys the scheme on the controller side to control and manage the interferer. The user places the interferer near the smart device, and the controller transmits the modulated interference signal to the interferer through Bluetooth. The interferer injects interference signals into the surrounding microphones, preventing unauthorized eavesdropping and protecting user privacy.

[0059] Step 5: When it is detected that the user is talking, the controller sends the modulated near-ultrasonic interference signal to the interferer, and the interferer injects the interference signal into the nearby recording microphones.

[0060] Continuous interference has a high power requirement for both the controller and the interferer. In addition, injecting interference signals when the user is not speaking can inadvertently trigger the response of the smart device, thus reducing the user experience.

[0061] In an embodiment, the present embodiment detects the user's speech activity in real time. The controller collects the ambient audio signal in real time, and the present embodiment divides the audio signal into frame signals of 20 ms in length, and analyzes each frame independently. The present embodiment extracts the energy distribution, spectral tilt and peak features of each frame to distinguish the features of speech and non-speech. When at least 10 consecutive frame signals are detected to have speech features, it indicates that there is user speech activity at this time. When the user is detected to be speaking, the controller sends the modulated near-ultrasonic interference signal to the jammer through Bluetooth. Then, the jammer injects the interference signal into the nearby recording microphone.

[0062] See Figure 2 In an embodiment, an automatic interference control strategy is used for interference control; the specific implementation includes the following steps:

[0063] Step 5.1: Real-time speech activity detection; when speech activity is detected, the following step 5.2 is executed; otherwise, the following step 5.3 is executed.

[0064] Step 5.2: Start interference & wait for speech command, while to preserve the normal interaction between the user and the smart device, the present embodiment continues to provide a strategy for interference suspension. The controller is usually located outside the effective interference radius of the jammer and close to the user, in a state not affected by interference. Therefore, the controller constantly captures the speech signal to monitor the user's intention, and when it detects the user's intention to issue a wake-up command, it instructs the jammer to temporarily stop interference. The present embodiment implements a phoneme-level wake-up word detection technology for user speech. When the controller identifies the first phoneme in the predefined wake-up word phoneme sequence (for example, " / h / " in "Hey Alexa"), the controller instructs the jammer to switch to the stop interference state, while continuing to monitor the subsequent phonemes until the entire phoneme sequence is identified. If the phoneme sequence of the keyword is not continuously detected, the jammer returns to the interference state. In order to identify a single phoneme, the present embodiment uses a trained BiLSTM-CRF model for single phoneme recognition, and in training, the present embodiment does not directly use the mfcc features of each frame for training, but combines the features of the previous and subsequent 10 frames with the features of the current frame as the features of the frame for training. And after training, the present embodiment combines the wake-up word dataset and a part of the original speech dataset to fine-tune the model. This not only makes the model have a preference for specific phonemes of the wake-up word, but also prevents the model from classifying all phonemes as wake-up word phonemes, ensuring the effectiveness of interference on normal speech.

[0065] Step 5.3: At this time, no user voice activity is detected, so there is no need to perform the interference. It is determined whether the interference is being performed at this time; if yes, stop interference is performed; if no, no operation is performed. When the voice assistant returns to the inactive state, the controller will continue to monitor the voice activity to determine whether to resume the interference. The voice assistant has two interactive states, a follow-up mode and a non-follow-up mode. In the follow-up mode, the voice assistant remains in the active state after the initial interaction and waits for further commands without the need for a wake-up word command again. If no further interaction occurs within a predefined time period, the voice assistant will automatically switch to the inactive state. For this mode, the present embodiment simulates the workflow of the voice assistant to determine when it ends the activity. Specifically, the present embodiment uses voice activity detection and a preset silence period, and if there is no voice activity during the silence period, it is considered that the voice assistant is in the inactive state. In the non-follow-up mode, the voice assistant responds to a single voice command and then returns to the inactive state immediately after the command is executed. In this mode, as soon as the voice assistant starts responding, the controller resumes monitoring the voice activity and resumes the interference when voice activity is detected again. In particular, the present embodiment distinguishes the accurate time when the voice assistant starts responding by comparing the Mel-frequency cepstral coefficient (MFCC) features of the captured signal frames with the voice features of the user and the voice assistant.

[0066] The present embodiment also provides a voice privacy protection device based on high-frequency sound signal mixing, comprising:

[0067] One or more processors;

[0068] A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the voice privacy protection method based on high-frequency sound signal mixing.

[0069] The present application is further described below through specific experiments.

[0070] The present application designs and implements experiments in the digital domain and the physical domain respectively. In the digital domain experiment, the present application tests audio data of about 1400 hours covering 11 different signal-to-noise ratios (-5 to 5); in the physical domain experiment, the present application collects more than 100 hours of audio data for testing. The present application uses three commercial automatic speech recognition models (Tencent, Google, and Xunfei) and two open-source ASR models (Wenet and Whisper) for evaluation. The results show that, whether in the digital domain or the physical domain environment, when the signal-to-noise ratio is 0, the speech recognition error rate caused by the mixed speech noise of the present application is more than 60%, which is 22% higher than that of the Gaussian white noise; and when the signal-to-noise ratio decreases to -4, the error rate further increases to more than 90%, which is 18% higher than that of the Gaussian white noise. In addition, the speech enhancement and blind source separation techniques have limited effect on reducing the speech recognition error rate caused by the mixed speech noise, and the error rate only decreases by less than 3% and 2% respectively.

[0071] Through a large number of experimental verification in the digital domain and the physical domain, the present application significantly improves the security and user-friendliness of speech privacy protection. Compared with the prior art, the present application has better speech privacy protection effect under the same signal-to-noise ratio, does not require special ultrasonic equipment, and preserves normal speech interaction between the user and the smart speaker while effectively interfering.

[0072] The present application adopts a novel and robust mixed speech noise scheme, and through an automatic interference control strategy, realizes effective interference while maintaining normal speech interaction between the user and the smart speaker. Since the near-ultrasonic signal used in the present application is free of dependence on special equipment and the mixed speech noise effectively confuses the acoustic characteristics of the original speech, the present application has both practicality and security, which not only effectively protects user privacy, but also can resist attacks of speech enhancement, blind source separation and other denoising technologies.

[0073] It should be understood that the above-described embodiments are part of the embodiments of the present application, rather than all the embodiments. In addition, the technical features in each embodiment or single embodiment provided by the present application can be combined with each other arbitrarily to form a feasible technical solution, and such combination is not restricted by the order of steps and / or structure composition mode, but must be based on the realization by the person skilled in the art, when the combination of technical solutions appears contradictory or unachievable, it should be considered that such combination of technical solutions does not exist, nor within the protection scope required by the present application.

[0074] It should be understood that the above description is merely a detailed explanation of the preferred embodiments and is not intended to limit the patent protection scope of the present application. Any modification or alternation made by those skilled in the art without departing from the scope of the present application shall fall within the patent protection scope of the present application. The patent protection scope of the present application shall be subject to the appended claims.

Claims

1. A speech privacy protection method based on high-frequency acoustic signal mixing, characterized by, The method comprises the following steps: Step 1: Collect user voice data and build a personalized user corpus; Step 2: User customizes the scene requirement, selects the text theme of voice synthesis, and expands the user corpus; Step 3: The user selects the noise theme related to this time from the text theme selected in step 2 according to the specific life situation it is in, randomly extracts a voice segment from the user corpus corresponding to the noise theme to generate interference noise; The specific implementation of step 3 comprises the following steps: Step 3.1: Use a random acceleration factor α to accelerate the randomly extracted voice segment; Step 3.2: Remove the gaps of the accelerated voice segment, remove the voice gaps whose energy is continuously lower than a certain threshold; Step 3.3: Superimpose the processed voice segment in the time domain to generate mixed voice interference noise; Step 4: Modulate the interference noise to a preset frequency band interval; Step 5: When it is detected that the user is talking, the controller sends the modulated interference signal to the jammer, and the jammer injects the interference signal into the nearby microphone.

2. The speech privacy protection method based on high-frequency acoustic signal mixing according to claim 1, characterized in that, The specific implementation process of step 2 is: using voice synthesis technology to expand the user's personal corpus, first collecting audio samples from different speakers, then adding emotion category labels in the collected data.

3. The speech privacy protection method based on high-frequency acoustic signal mixing according to claim 2, characterized in that: Optimize the text encoder and duration predictor of the VITS model to enable it to accept speaker ID and emotion category embedding; The optimization of the text encoder is to add two embedding layers for speaker ID and emotion category in the input layer of the text encoder, which converts discrete speaker ID and emotion category into continuous vector representation, and the word embedding of the text is spliced with the embedding vector of the speaker ID and emotion category; The optimization of the duration predictor is to increase the number of layers and nodes of each layer of the duration predictor according to the new input dimension, increase the original number of layers from 2 to 4, and increase the number of nodes of each layer from 256 to 512, to ensure that the model can effectively process these additional embedding information and enhance the expression ability of the model.

4. The speech privacy protection method based on high-frequency acoustic signal mixing according to claim 1, characterized in that: In step 3.1, Fourier transform is used to convert the randomly selected voice segment to the frequency domain; in the frequency domain, each frequency component has its own phase evolution with time; by adjusting the phase information and the step length of phase evolution, the length of the signal is shortened; finally, inverse Fourier transform is performed to convert the audio signal to the time domain, thereby obtaining an accelerated audio that maintains the pitch.

5. The speech privacy protection method based on high-frequency acoustic signal mixing according to claim 1, characterized in that: In step 3.2, the removal of the voice gap whose energy is continuously lower than a certain threshold is achieved by analyzing the sample data in the accelerated audio stream frame by frame, calculating the short-time energy of the audio signal, and if a certain segment of audio has a duration longer than threshold A and contains all audio frames whose short-time energy is lower than threshold B, the segment of audio is marked as silent; all audio parts marked as silent are removed from the audio stream.

6. The speech privacy protection method based on high-frequency acoustic signal mixing according to claim 1, characterized in that: In step 4, the upper sideband modulation technology is used to modulate the interference noise to the high frequency band of 17kHz to 21kHz.

7. The speech privacy protection method based on high-frequency acoustic signal mixing according to claim 1, characterized in that: In step 5, the voice activity detection technology is used for real-time voice activity detection; the controller collects the environmental audio signal in real time, divides the audio signal into frame signals of a preset length, analyzes each frame independently, extracts the energy distribution, spectral tilt and peak value characteristics of each frame to distinguish the characteristics of voice and non-voice; If at least N consecutive frame signals are detected to have voice characteristics, it is determined that there is user voice activity at this time; wherein N is a preset value.

8. The speech privacy protection method based on mixing high frequency acoustic signals according to any one of claims 1-7, characterized in that: In step 5, an automatic interference control strategy is used for interference control; the specific implementation includes the following steps: Step 5.1: Real-time voice activity detection; when voice activity is detected, the following step 5.2 is executed; otherwise, the following step 5.3 is executed; Step 5.2: Start interference & wait for voice command, through the phoneme-level wake-up word detection technology, when the controller recognizes the first phoneme in the predefined wake-up word phoneme sequence, the controller instructs the interferometer to switch to the stop interference state while continuing to monitor the subsequent phonemes, until the entire phoneme sequence is recognized; if the phoneme sequence of the wake-up word is not continuously detected, the interferometer returns to the interference state; Wherein, the trained BiLSTM-CRF model is used for single phoneme recognition, in the training, the features of the current frame and the previous and next 10 frames are combined together as a frame of features for training, after training, the wake-up word data set and a part of the original voice data set are combined to fine-tune the model; Step 5.3: Determine whether the interference is ongoing; if yes, stop the interference; if no, do nothing; determine whether the voice assistant uses follow-up mode or non-follow-up mode; when the voice assistant uses non-follow-up mode, use voice activity detection and a preset silence period strategy for interference recovery; when the voice assistant uses follow-up mode, once the voice assistant responds to the user's instruction, the controller resumes monitoring voice activity and resumes interference when voice activity is detected again.

9. A speech privacy protection device based on mixing of high frequency acoustic signals, characterized in that Comprise: One or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the high-frequency sound signal mixing-based voice privacy protection method as claimed in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Anti-eavesdropping ultrasonic interference sample design method, system and device

    CN116388884A

  • Privacy protection method and device based on white-box voice confrontation sample

    CN115208507A

  • Voice interference noise design method based on human voice structure

    CN115841821A