A method for recording interference based on ultrasonic waves for Chinese speech anti-eavesdropping
By generating coupled noise highly correlated with the user's voice and modulating it with ultrasound, the problem of existing ultrasonic recording interference systems being bypassed by eavesdroppers' noise reduction methods is solved, improving system security and user experience, and protecting voice privacy.
Patent Information
- Application Number
- CN202411294767.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-14
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-09-14
AI Technical Summary
Existing ultrasonic recording jamming systems have security vulnerabilities when facing noise reduction techniques used by eavesdroppers, as these techniques can be bypassed by algorithms. Furthermore, user experience is affected by hardware imperfections, and audible noise may be generated.
An ultrasonic recording jamming method for Chinese speech is designed. By generating coupled noise that is highly correlated with the user's speech, and using phoneme processing, coupled noise generation, ultrasonic carrier modulation and correction algorithms, ultrasonic waves are generated and emitted to interfere with eavesdropping and avoid the generation of audible noise.
It enhances system security and user experience, increases the difficulty of separating eavesdropping noise, reduces the risk of audible sound leakage, and protects voice privacy.
Smart Images

Figure CN119360885B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of eavesdropping defense, in particular to an ultrasonic-based recording interference method for Chinese speech eavesdropping prevention. BACKGROUND
[0002] With the growing demand for voice privacy protection, various anti-eavesdropping technologies have emerged. Among them, the ultrasonic recording interference system has received widespread attention from researchers and consumer markets due to its unique advantages, such as: human ears cannot perceive, widely applicable devices, large coverage, and low deployment cost. The ultrasonic interference system uses the nonlinear characteristics of microphones to generate low-frequency noise in electronic recording devices using ultrasonic signals that are inaudible to the human ear, thereby effectively resisting eavesdropping and protecting voice privacy. Due to the nonlinear characteristics of microphones in electronic recording devices, high-frequency ultrasonic signals will be distorted when they exceed the frequency response range, and then be distorted into low-frequency noise in the audible frequency band. By carefully designing and modulating ultrasonic signals, electronic recording devices can be efficiently interfered with to prevent eavesdropping and harm. However, in actual scenarios, eavesdroppers will use various technical means to reduce noise interference in recordings, thereby recovering the original voice information and achieving the purpose of stealing voice privacy.
[0003] To address the existing security risks in the field of eavesdropping interference, the present application proposes an ultrasonic recording interference method designed specifically for Chinese speech. Considering that eavesdroppers may use various noise reduction methods, the present application designs a coupling noise generation algorithm based on the characteristics of Chinese speech. Without real-time collection of user voice content, the ultrasonic interference noise generated by this algorithm is highly correlated with the voice signal to be protected and is difficult to separate. In addition, to improve the practicality and user experience of the system, the present application designs an ultrasonic carrier modulation strategy, which effectively avoids the generation and leakage of audible frequency band noise caused by hardware imperfections such as ringing effect, ensuring that the user's hearing is not affected during use, and taking into account the efficient noise injection of different types of eavesdropping devices. SUMMARY
[0004] To address the problems raised in existing methods, the present application provides an ultrasonic-based recording interference method for Chinese speech eavesdropping prevention, which solves the security risks that can be bypassed by algorithms. The present application is realized through the following scheme:
[0005] The present application discloses an ultrasonic-based recording interference method for Chinese speech eavesdropping prevention, comprising the following steps:
[0006] Obtain a small segment of voice of the user to be protected;
[0007] Generate a large amount of corpus data for the user using voice generation technology based on the small segment of voice of the user;
[0008] The corpus data of the user is processed by a phoneme processing algorithm to generate a phoneme-level corpus of the user;
[0009] The coupling noise generation algorithm randomly extracts any number of phonemes from the phoneme-level corpus each time noise is generated to generate coupling noise;
[0010] The coupling noise is modulated by an ultrasonic carrier wave, and the corrected sound wave is emitted by an ultrasonic loudspeaker through an ultrasonic correction algorithm.
[0011] As a further improvement, the phoneme processing algorithm according to the present application is specifically:
[0012] 1) The intensity threshold of the corpus data is calculated using the maximum inter-class variance method, and the corpus data of the user is cut into phonemes;
[0013] 2) The cut phonemes are clustered according to monophthongs, diphthongs and consonants, and different tones of the same phoneme are regarded as different categories during clustering;
[0014] 3) The average signal of each phoneme of the user is obtained by aligning and averaging the clustered phoneme signals using a dynamic time warping (Dynamic Time Warping) algorithm, thereby obtaining the phoneme library.
[0015] As a further improvement, the coupling noise generation algorithm according to the present application is specifically:
[0016] 1) Randomly extract phonemes from the phoneme-level corpus of the user, i.e. randomly extract any number of phonemes every short interval to generate a user's semantic-free phoneme sequence s(t);
[0017] 2) Generate a noise signal n1(t) with similar characteristics to the user's voice based on the user's phoneme sequence s(t):
[0018]
[0019] where M 1×2 (t) is a 1x2 non-singular random matrix; n r1 (t) is a random noise; L[·] is a non-integrable nonlinear aliasing function:
[0020]
[0021] 3) Based on n1(t), a time-frequency noise signal n2(t) is generated by introducing a random noise in the time domain and the frequency domain through convolution operation respectively:
[0022]
[0023] Wherein, * is convolution operation;n ri (t)(i=1,2,3) is mutually independent and distributed different random noise;
[0024] 4) Based on n2(t), the frequency spectrum bandwidth is widened by using frequency hopping spread spectrum technology, and finally the coupling noise n3(t) is formed:
[0025]
[0026] Wherein, Is the frequency hopping signal;ω n Is the nth frequency point in the frequency table; Is the phase modulation function, wherein d is the error coded signal code, t is time, and Δω is the frequency difference component before and after frequency hopping.
[0027] As a further improvement, the ultrasonic noise correction algorithm described in the application is specifically:
[0028] 1) Smooth filtering is performed on the ultrasonic signal using a frequency domain filter, and the expression of the frequency domain filter is:
[0029]
[0030] Wherein, ω is the angular frequency;
[0031] 2) The smooth filtered ultrasonic signal is decomposed into a plurality of signals with a relatively narrow bandwidth, and the frequency of each signal is less than 50Hz, and each decomposed signal is emitted using a different ultrasonic emitter.
[0032] The beneficial effects of the application are as follows:
[0033] The application discloses a recording interference method based on ultrasonic waves.
[0034] The application designs a user-unaware voice eavesdropping interference technology for the problem of Chinese voice eavesdropping, fully considers the ability of eavesdroppers and the acoustic privacy leakage risk caused by the eavesdroppers under real conditions, and enhances the effectiveness and security of the system.
[0035] The application uses a voice generation technology to expand user corpus, cuts the corpus into phonemes without semantics according to the characteristics of Chinese voice, reduces the voice registration time cost of the user, protects the privacy of the voice of the user, and enhances the protection ability for Chinese voice.
[0036] The application generates noise coupled with the voice of the user according to the voice of the user, increases the difficulty for eavesdroppers to separate the noise, and more completely protects the voice of the user from being eavesdropped.
[0037] The present application modifies the ultrasonic signal in view of the audible sound leakage problem of the ultrasonic emission device, weakens the leakage degree of audible sound, enhances the user experience, and reduces the risk of the ultrasonic emission device being discovered by an eavesdropper. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 The flowchart of the system of the present application. DETAILED DESCRIPTION
[0039] The technical solutions of the present application will be further described below in combination with the drawings of the specification through specific embodiments:
[0040] The present application discloses a recording interference method based on ultrasonic waves for Chinese speech anti-eavesdropping, Figure 1 The flowchart of the system of the present application, aiming at the problem of Chinese speech anti-eavesdropping, proposes a user-unaware voice eavesdropping interference technology, and generates unique coupling noise according to the user voice characteristics, thereby enhancing the effectiveness and security of the system.
[0041] The present application comprises the following steps:
[0042] Step 1: Obtain a small piece of voice of the user to be protected;
[0043] Step 2: Use a speech generation technology to generate a large amount of corpus data of the user according to the small piece of voice of the user;
[0044] Step 3: Use a phoneme processing algorithm to generate a phoneme-level corpus library of the user from the corpus data of the user;
[0045] Step 4: Use a coupling noise generation algorithm to randomly extract any number of phonemes from the phoneme-level corpus library each time noise is generated to generate coupling noise;
[0046] Step 5: Modulate the coupling noise through an ultrasonic carrier wave, and use an ultrasonic correction algorithm to emit the corrected sound wave through an ultrasonic loudspeaker.
[0047] Specifically, the details of the steps of the present application are as follows.
[0048] Step 1: Obtain a small piece of voice of the user to be protected:
[0049] Use a recording device to collect the user's voice when speaking for a few seconds. This is used for subsequent user corpus generation.
[0050] Step 2: Use a speech generation technology to generate a large amount of corpus data of the user according to the small piece of voice of the user:
[0051] Use the SV2TTS speech generation network to generate more user corpus from the collected short-duration user voice data.
[0052] Step 3: The user's corpus data is processed by a phoneme processing algorithm to generate a user's phoneme-level corpus:
[0053] To further improve the correlation between the coupling noise based on the generated speech signal and the content to be protected, the present application uses the initial consonant system and tone features of Chinese to generate more targeted user corpus. The present application segments the generated corpus content according to phonemes, classifies them according to initial consonants and tones, and obtains the average corpus of each phoneme of the user to be protected. Specifically, the following steps are taken:
[0054] 1) The maximum inter-class variance method (also known as the OTSU method) is used to calculate the speech signal intensity threshold, and the phonemes are cut based on this threshold.
[0055] 2) Each phoneme is clustered according to monophthong, diphthong, and consonant. Considering that the tone in Chinese speech has a great influence on the speech spectrum features, different tones of the same phoneme are also considered as different categories. The final number of categories = (number of monophthong categories + number of diphthong categories + number of consonant categories) x number of tones. The number of categories can be adjusted according to the language type, for example: there are double consonant phonemes in some dialects, and there are "yin up" and "yang up" tones in Wu dialect.
[0056] 3) The corpus signals in the same tone phoneme are aligned based on the Dynamic Time Warping (DTW) algorithm and then averaged to obtain the average signal of each phoneme of the user.
[0057] Step 4: Through the coupling noise generation algorithm, any number of phonemes are randomly selected from the phoneme-level corpus each time the noise is generated to generate the coupling noise:
[0058] To further improve the security of the coupling noise and avoid the use of noise cancellation algorithms by eavesdroppers to evade the interference of the anti-eavesdropping device, the present application randomly selects phoneme combinations from the user's phoneme-level corpus to generate coupling noise, enhancing the anti-noise ability of the noise. Specifically, the following steps are taken:
[0059] 1) Randomly extract phonemes from the user's phoneme-level corpus, i.e., randomly extract any number of phonemes every short interval to generate a user's semantic-free phoneme sequence s(t).
[0060] 2) Perform time-domain nonlinear aliasing on the phoneme sequence s(t) and random noise to generate a noise signal n1(t) with similar characteristics to the user's voice:
[0061]
[0062] where M 1×2(t) is a 1x2 non-singular random matrix; n r1 (t) is a random noise; L[·] is a non-integrable nonlinear aliasing function:
[0063]
[0064] Due to the many-to-many mapping relationship of nonlinear aliasing, its inverse transformation is not unique. This feature makes the generated coupling noise have high complexity and robustness, increasing the difficulty of separating noise from speech signals.
[0065] 3) Based on the noise signal n1(t) similar to the user's voice characteristics, its interference ability is enhanced in time domain and frequency domain respectively. Specifically, a random noise is introduced in each convolution operation in time domain and frequency domain. The obtained coupling noise is:
[0066]
[0067] Where, * is convolution operation; n ri (t) (i = 1, 2, 3) are random noises that are independent of each other and have different distributions. Through convolution operation in time domain and frequency domain, the obtained noise is tightly coupled with the speech signal, enhancing the ability of the noise to resist the de-noising means of the eavesdropper.
[0068] 4) Based on n2(t), in order to resist the de-noising means such as ultrasonic sniffing using time-frequency domain features, the frequency hopping spread spectrum technology is used to widen the frequency bandwidth of the coupling noise. The obtained coupling noise is:
[0069]
[0070] Where, is a frequency hopping signal; ω n is the nth frequency point in the frequency hopping table; is a phase modulation function, where d is the error encoded signal code, t is time, and Δω is the frequency difference component before and after frequency hopping. Through this method, the bandwidth of the coupling noise n3(t) is greatly expanded, and in this case, the eavesdropper cannot obtain all the characteristics of the interference noise, so it is impossible to recover the speech signal.
[0071] Step 5: The coupling noise is modulated by an ultrasonic carrier wave, and the modified sound wave is emitted through an ultrasonic loudspeaker through an ultrasonic correction algorithm:
[0072] Due to the imperfect characteristics of the ultrasonic wave emitting device, audible sound noise will be generated during the emission process, affecting the daily use of the user, and more likely to arouse the vigilance of the eavesdropper, so as to take measures to avoid eavesdropping. In order to solve this problem, the present application designs an ultrasonic correction algorithm from two aspects of reducing the ringing effect and cutting the wideband spectrum.
[0073] 1) Reduce the ringing effect: Ringing effect can cause audible noise at the transmitting end of the ultrasonic signal. When the frequency of the acoustic signal changes sharply, a pulse signal will be generated in the hardware device. This pulse wave has high energy in the audible frequency band, and the oscillation on the sound waveform is similar to the ringing of a bell, so it is called ringing effect. In the process of ultrasonic speaker sound production, a band-pass filter is used to cut off the signal frequency to ensure the generation of ultrasonic signals of a specific frequency without deviation. However, due to the imperfection of the hardware of the band-pass filter, a sudden change in frequency may occur in the process, triggering the ringing effect. In principle, the sudden change in the acoustic signal can be regarded as being truncated by a rectangular window function, and the frequency spectrum distribution characteristics of the rectangular window function are as follows: when the input frequency ω is less than the cutoff frequency ω0, X(ω) = 1, that is, the signal in this frequency band is retained; otherwise, when the input frequency ω is greater than or equal to the cutoff frequency ω0, X(ω) = 0, that is, the signal in this frequency band is filtered out. The inverse Fourier transform of the band-pass filter in the time domain is a Sine function:
[0074]
[0075] The response of the band-pass filter in the time domain is a main pulse with high amplitude, followed by countless periodic pulses that gradually decay. In this phenomenon, the main pulse with the highest amplitude is called a ringing pulse wave signal. Due to the imperfections of the actual hardware, this significant ringing pulse wave signal may leak into the low-frequency audible frequency band.
[0076] In order to reduce the influence of the ringing pulse, the present application uses a frequency domain filter to smooth the ultrasonic wave signal. The present application applies a Sine function in the frequency domain to the band-pass filter to replace the rectangular window function. The expression of the Sine function is:
[0077]
[0078] According to the reversibility of Fourier transform, the expression of the Sine function band-pass filter in the time domain is similar to the rectangular window function. This filter design can effectively smooth the sudden change of the signal frequency, thereby eliminating the ringing pulse phenomenon. Although this processing will produce some out-of-band frequency components, the energy of these out-of-band frequency components is low and still within the ultrasonic frequency range. Therefore, this modulation method can significantly reduce the ultrasonic energy in the band-pass frequency range while avoiding the generation of audible noise.
[0079] 2) Cut the wideband spectrum: The internal amplifier of the ultrasonic speaker can produce nonlinear effects, causing the emitted ultrasonic signal to be demodulated into low-frequency sound. Assuming that the center frequency of the ultrasonic signal to be transmitted is f c, with bandwidth 2B. Due to the non-linear characteristic of the ultrasonic speaker, it will generate a low frequency sound signal with center frequency 0 and bandwidth B. This non-linear effect will make the ultrasonic signal become audible noise when it propagates in the air.
[0080] To solve this problem, the present application decomposes the ultrasonic signal into multiple signals with narrower bandwidth, and makes the frequency lower than 50Hz, i.e. converts it into infrasound which is not perceived by human ears, so as to avoid the appearance of audible noise. Taking the ultrasonic signal with bandwidth 2B as an example, it is equally divided in frequency domain, and is decomposed into k signals with bandwidth . Then the non-linear effect will generate an audible sound signal with center frequency 0 and bandwidth . By using a larger k, the low frequency signal generated by self-demodulation can not be perceived by human ears, i.e. . The k divided ultrasonic signals are emitted by k different ultrasonic emitters. The k signals are respectively demodulated and spliced into the original noise signal on the eavesdropper's device.
[0081] The above is only to illustrate the technical idea of the present application, and cannot limit the protection scope of the present application. Any modification made according to the technical idea of the present application on the basis of the technical scheme falls within the protection scope of the claims of the present application.
Claims
1. A method for recording interference based on ultrasonic waves for Chinese speech anti-eavesdropping, characterized in that, The method comprises the following steps: Obtaining a small piece of voice of a user to be protected; Generating a large amount of corpus data of the user according to the small piece of voice of the user by using a voice generation technology; Generating a phoneme-level corpus of the user by processing the corpus data of the user by a phoneme processing algorithm; Generating coupled noise by a coupled noise generation algorithm, randomly extracting any number of phonemes from the phoneme-level corpus each time noise is generated; Modulating the coupled noise by an ultrasonic carrier wave, and emitting the modified sound wave through an ultrasonic loudspeaker by an ultrasonic noise correction algorithm; The phoneme processing algorithm specifically comprises: 1) using the maximum inter-class variance method to calculate the intensity threshold of the corpus data, and cutting the corpus data of the user into phonemes; 2) clustering the cut phonemes according to monophthongs, diphthongs and consonants, and regarding different tones of the same phoneme as different categories during clustering; 3) aligning and averaging each category of phoneme signal after clustering by using a dynamic time warping algorithm, thereby obtaining the average signal of each phoneme of the user as the phoneme-level corpus; The coupled noise generation algorithm specifically comprises: 1) randomly extracting phonemes from the phoneme-level corpus of the user, i.e. randomly extracting any number of phonemes every short interval to generate a user's semantic-free phoneme sequence s(t); 2) generating a noise signal n1(t) with similar characteristics to the user's voice according to the phoneme sequence s(t): where M 1×2 (t) is a 1 x 2 nonsingular random matrix; n r1 (t) is a random noise; L[·] is a nonintegrable nonlinear aliasing function: 3) based on n1(t), introducing a random noise in the time domain and the frequency domain respectively by convolution operation to generate a time-frequency noise signal n2(t): where * is a convolution operation; n ri (t), i = 1, 2, 3 are mutually independent and differently distributed random noises; 4) based on n2(t), widening its spectral bandwidth by using frequency hopping spread spectrum technology to finally form coupled noise n3(t): wherein is a frequency hopping signal; ω n is the nth frequency bin in the frequency hopping frequency table; is a phase modulation function, where d is the error encoded signal code, t is time, and Δω is the frequency difference component before and after frequency hopping. The ultrasonic noise correction algorithm specifically comprises: 1) using a frequency domain filter to perform smoothing filtering on the ultrasonic signal, and the expression of the frequency domain filter is: where ω is the angular frequency; 2) decomposing the smoothed ultrasonic signal into multiple signals with relatively narrow bandwidth, and making the frequency of each signal lower than 50Hz, and each decomposed signal is emitted by using a different ultrasonic emitter.
Citation Information
Patent Citations
Generating method for shielding signals used for protecting Chinese speech privacy
CN104637485A
Anti-eavesdropping method and system based on ultrasonic injection technology
CN114337850A