Method, device and equipment for anti-voice deep forgery based on quantum formant disturbance
By adding quantum noise to the fundamental frequency and formant frequency of the speech signal, interfering with the speech feature learning of the deep forgery model, the problem of detection algorithms being bypassed in the prior art is solved, and efficient speech signal protection is achieved.
Patent Information
- Application Number
- CN202510152120.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-11
AI Technical Summary
Existing deep forgery detection algorithms may be bypassed by new forgery techniques, and traditional methods such as adding noise will reduce audio quality and affect user experience.
By reading the original speech signal, the fundamental frequency and formant frequency are extracted, and quantum noise is generated based on the parameterized quantum circuit, and added to the fundamental frequency and formant frequency, the disturbed signal is obtained and preprocessed to reconstruct the speech signal.
Effectively interfere with the learning and generation of speech features by the deep forgery model, increasing the difficulty of attack of the model, while ensuring that the voice after adding noise is minimal auditory difference from the original speech.
Smart Images

Figure CN119993113A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information security technology, and in particular to a method, device and equipment for countering voice deep forgery based on quantum resonance peak perturbation. Background Art
[0002] With the rapid development of artificial intelligence and deep learning technologies, voice deep fake (DeepFake) technology has been able to generate highly realistic voice and video content. These technologies are widely used in entertainment, film and television production and other fields, but they also bring serious security and ethical issues. For example, deep fake technology can be used to generate fake voices, impersonate others, commit fraud or spread false information. Current deep fake detection algorithms may be bypassed by new and more advanced fake technologies. Traditional methods of adding noise to audio may degrade audio quality and affect user experience. Generating adversarial samples requires understanding the details of the attack model and may not work well for different models. Summary of the invention
[0003] In view of this, the present application proposes a method, device and apparatus for anti-speech deepfake based on quantum formant perturbation, which is suitable for interfering with the learning and generation of speech features by the deepfake model.
[0004] According to one aspect of the present application, a method for anti-speech deepfake based on quantum formant perturbation is provided, comprising: reading an original speech signal, extracting a fundamental frequency and a formant frequency of the original speech signal;
[0005] Generating quantum noise based on a parameterized quantum circuit; adding the quantum noise to the extracted fundamental frequency and resonance peak frequency to obtain a disturbed signal;
[0006] The disturbed signal is preprocessed to obtain a reconstructed speech signal.
[0007] In a possible implementation, when extracting the fundamental frequency and the formant frequency of the original speech signal, the fundamental frequency of each time frame in the original speech signal and the first three formant frequencies of the current time frame are extracted.
[0008] In a possible implementation manner, quantum noise is generated based on a parameterized quantum circuit, and a bit value measured by the parameterized quantum circuit is converted into a real noise value.
[0009] In a possible implementation, the disturbed signal includes: a disturbed base frequency and a disturbed resonance peak frequency; generating an excitation signal based on the disturbed base frequency; and constructing a resonance peak filter based on the disturbed resonance peak frequency.
[0010] In a possible implementation, the parameterized quantum circuit is:
[0011]
[0012] in, represents the evolution operator of the entire quantum circuit; L represents the number of layers of the circuit; represents the parameterized quantum gate of the lth layer, with parameters
[0013] In a possible implementation, an excitation signal is generated based on the disturbed fundamental frequency; a formant filter is constructed based on the disturbed formant frequency, and further includes: passing the excitation signal through the formant filter to obtain the reconstructed speech signal.
[0014] In a possible implementation, optimizing the parameterized quantum circuit includes:
[0015] Calculating a perceptual difference between the original speech signal and the reconstructed speech signal;
[0016] The resonance peak frequency before adding noise and the change in the resonance peak frequency after the disturbance are calculated as the interference degree; the target loss function is obtained by combining the perceptual difference and the interference degree in a weighted manner; and the loss function is minimized.
[0017] According to another aspect of the present application, there is provided a device for anti-speech deepfake based on quantum formant perturbation, comprising: a feature extraction module, a noise perturbation module, and a signal reconstruction module;
[0018] The feature extraction module is configured to read the original speech signal and extract the fundamental frequency and formant frequency of the original speech signal;
[0019] The noise perturbation module is configured to generate quantum noise based on a parameterized quantum circuit; add the quantum noise to the extracted fundamental frequency and resonance peak frequency to obtain a perturbed signal;
[0020] The signal reconstruction module is configured to preprocess the disturbed signal to obtain a reconstructed speech signal.
[0021] According to another aspect of the present application, there is provided a device for anti-speech deep fake based on quantum resonance peak perturbation, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to execute the above method.
[0022] According to another aspect of the present application, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0023] Beneficial effects of this application: The optimized quantum noise generated by the quantum neural network can effectively interfere with the learning and generation of speech features by the deep fake model. Adding noise to the resonance peak frequency of the speech signal can accurately interfere with the key features of the speech, which is better than adding noise directly, and increases the difficulty of attacking the deep fake model. Through the design of the loss function and the adjustment of the optimization algorithm, the differences in human ear perception and the interference effect are considered at the same time to guide the optimization process. Let the noise interfere with the model without affecting the sound quality. Use linear predictive coding to reconstruct the speech signal to ensure that the speech after adding noise is minimally different from the original speech in hearing.
[0024] Other features and aspects of the present application will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the present application and, together with the description, serve to explain the principles of the present application.
[0026] Figure 1 A flow chart showing a method for anti-speech deep fake based on quantum formant perturbation according to an embodiment of the present application;
[0027] Figure 2 A schematic diagram of a quantum circuit according to an embodiment of the present application is shown;
[0028] Figure 3 The original audio waveform diagram of the embodiment of the present application is shown;
[0029] Figure 4 An audio waveform diagram after adding noise according to an embodiment of the present application is shown;
[0030] Figure 5 The deepfake model of the embodiment of the present application generates an audio fundamental frequency and a resonance peak frequency diagram;
[0031] Figure 6 A diagram showing the fundamental frequency and formant frequency of the deepfake model forged audio before adding noise in an embodiment of the present application;
[0032] Figure 7 A diagram showing the fundamental frequency and formant frequency of the forged audio of the deepfake model after adding noise in an embodiment of the present application;
[0033] Figure 8 The original audio spectrum diagram of the embodiment of the present application is shown;
[0034] Fig. 9 An audio spectrum diagram after adding noise according to an embodiment of the present application is shown;
[0035] Fig.10 A comparison diagram of the audio waveforms forged by the deepfake model before and after adding noise in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0036] Various exemplary embodiments, features and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0037] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0038] In addition, in order to better illustrate the present application, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present application can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present application.
[0039] This application is suitable for interfering with the learning and generation of speech features by deep fake models, and ensuring that the speech after adding noise has the least auditory difference from the original speech.
[0040] Example 1
[0041] Figure 1 A flow chart of a method for anti-speech deep fake based on quantum formant perturbation according to an embodiment of the present application is shown. Figure 1 As shown, the method includes:
[0042] Step S100, read the original speech signal, extract the fundamental frequency and formant frequency of the original speech signal. The fundamental frequency (F0) is generated by the vibration of the vocal cords and determines the pitch of the sound. The formants are generated by the resonance of the vocal tract and are local peaks in the speech spectrum, which mainly determine the timbre and sound quality of the speech.
[0043] Step S200, generating quantum noise based on parameterized quantum circuits. Parameterized quantum circuits refer to quantum neural networks (QNNs), which are composed of several quantum bits and parameterized quantum gates. Quantum neural networks are machine learning models that use parameterized quantum circuits. By adjusting circuit parameters, QNNs can generate specific probability distributions. The angles of quantum gates (such as Hadamard gates and RX revolving gates) in quantum circuits are adjustable parameters, where the Hadamard gates put quantum bits in a superposition state, making the measurement results random. The RX revolving gate is a revolving gate around the X-axis, and its rotation angle is an adjustable parameter. These parameters can be trained through optimization algorithms. Using the randomness and complexity of quantum computing, unpredictable quantum noise is generated.
[0044] Step S300, adding quantum noise to the extracted fundamental frequency and resonance peak frequency to obtain a disturbed signal. The disturbed signal refers to the fundamental frequency and resonance peak frequency after adding noise respectively.
[0045] Step S400, preprocessing the disturbed signal to obtain a reconstructed speech signal. Preprocessing refers to generating an excitation signal and a formant filter based on the disturbed signal to combine the speech signal.
[0046] In a possible implementation, the audio file is loaded, and the audio data and sampling rate are obtained, and an acoustic analysis tool (such as Praat or Parselmouth library) is used to extract the fundamental frequency F0(t) of each time frame in the original speech signal and the first three formant frequencies F1(t), F2(t), and F3(t) of the current time frame. Where t represents the time frame index, F0(t) is the fundamental frequency at time t, and Fi(t) is the i-th formant frequency at time t, i=1,2,3.
[0047] Adding quantum noise to the resonance peak frequencies of speech signals can directly interfere with the deep fake model's learning and replication of speech features.
[0048] Specifically, the audio signal is divided into frames of 10 milliseconds in length, and each frame is traversed: the fundamental frequency is extracted using the To Pitch method; and the formant frequency is extracted using the To Formant (burg) method.
[0049] Store the extracted fundamental and formant frequencies as an array for subsequent processing.
[0050] Generate quantum noise using a quantum neural network (QNN), which is a parameterized quantum circuit. Use ParameterVector to create a parameter vector theta for parameterizing the rotation gate. Apply a Hadamard gate and an RX rotation gate to each qubit to form a parameterized quantum circuit. Create a sampler using Qiskit's Sampler primitive and create a SamplerQNN object based on the circuit and sampler. Print and plot the quantum circuit for visualization.
[0051] The parameterized quantum circuit is:
[0052]
[0053] in, represents the evolution operator of the entire quantum circuit; L represents the number of layers of the circuit; represents the parameterized quantum gate of the lth layer, with parameters
[0054] By sampling the output probability distribution of the quantum circuit, we can obtain random noise that conforms to quantum characteristics. We measure the output quantum state of the quantum circuit, obtain random measurement results, and convert the measured bit string b into a real noise value:
[0055]
[0056] Where N(t) is the noise value at time t; int(b) means converting the bit string b into an integer; n is the number of quantum bits; and scale is the noise amplification factor.
[0057] like Figure 2 The figure shows a quantum circuit diagram of an embodiment of the present application, including 5 quantum bits and corresponding gates, wherein the Hadamard gate is used to put the quantum bits into a superposition state so that the measurement results are random. The RX revolving gate is a revolving gate around the X axis, and its rotation angle is an adjustable parameter.
[0058] The quantum noise is added to the extracted base frequency and the first three resonance peak frequencies to obtain the disturbed base frequency and the disturbed resonance peak frequency. The formula for adding noise is as follows:
[0059] The fundamental frequency after perturbation:
[0060] The resonance peak frequency after perturbation: i = 1, 2, 3;
[0061] in, is the fundamental frequency after the disturbance at time t; Fi noisy (t) is the frequency of the i-th resonance peak after the disturbance at time t; represents the value of the quantum noise sequence generated by QNN at time t. α, β are parameters that control the noise intensity of the fundamental frequency and the resonance peak frequency.
[0062] Use QNN's forward method to calculate the probability distribution of the output. Sample each probability distribution, get the measurement result, and map it to a noise value. Add the generated noise sequence to the corresponding resonance peak frequency, and use np.clip to ensure that the frequency value is non-negative. Update the weight coefficient by training the quantum neural network algorithm, and the weight coefficient determines the proportion of the noise sequence to be added.
[0063] Based on the fundamental frequency after the disturbance Generate an excitation signal; based on the resonance peak frequency after the disturbance Construct a resonance peak filter; pass the excitation signal through the resonance peak filter to obtain a reconstructed speech signal. Unlike adding noise directly to the audio signal, adding quantum noise to the resonance peak frequency of the speech signal accurately interferes with the key features of the speech, increasing the difficulty of attacking the deep fake model.
[0064] Furthermore, the linear predictive coding (LPC) method is used to filter each frame of the speech signal to simulate the filtering effect of the vocal tract. White noise is generated as an excitation source, and the synthesized speech signal is obtained through the filter. The synthesized audio signal is normalized to prevent signal overload.
[0065] Furthermore, the parameters of the quantum circuit are optimized through training, specifically, the perceived difference between the original speech signal and the reconstructed speech signal is calculated;
[0066] By comparing the Mel-frequency cepstral coefficients (MFCC) of the original speech signal and the reconstructed speech signal, the difference between the two is calculated:
[0067]
[0068] T is the total number of time frames; is the parameter vector of the quantum circuit; MFCC original (t) is the MFCC feature of the original audio at time t; is the MFCC feature of the processed audio at time t. The MFCC features of the original audio and the synthesized audio are extracted, and the mean absolute difference of the MFCC features is calculated as a measure of the perceptual difference.
[0069] The change in the resonance peak frequency before adding noise and the resonance peak frequency after the disturbance is calculated as the degree of interference:
[0070]
[0071] Where Fi(t) is the original i-th resonance peak frequency; is the frequency of the ith resonance peak after perturbation, which depends on the parameter
[0072] Combining the perceived difference and the interference degree in a weighted manner to obtain a target loss function;
[0073]
[0074] λ∈[0,1] is a weight parameter that balances the influence between P and D. Use classical optimization algorithms (such as gradient descent, Adam) to optimize quantum circuit parameters To minimize the loss function
[0075] In each optimization iteration, the gradient of the loss function is updated
[0076]
[0077] Where η is the learning rate; is the loss function with respect to the parameter gradient.
[0078] By defining a loss function, we measure the perceived difference between the audio after adding noise and the original audio and the degree of interference to the deep fake model, and find the optimal quantum circuit parameters through the optimization algorithm. The noise obtained under this optimal parameter is added to the resonance peak frequency, which can balance the auditory difference and interference effect. The COBYLA method (evolutionary algorithm based on constrained optimization) is used for optimization without giving a gradient. By setting the maximum number of iterations, the time of the optimization process is limited. By optimizing the parameters of the quantum circuit, the generated quantum noise can maximize the interference effect while minimizing the impact on the human ear, thus achieving a balance between interference and sound quality.
[0079] Save the processed speech file, draw the waveform of the original speech signal and the perturbed speech signal, as well as the change graph of the fundamental frequency and formant frequency, to help analyze and understand the processing effect. Specifically, use matplotlib to draw the time domain waveform of the original speech signal and the perturbed speech signal; draw the change curve of the fundamental frequency (F0) and the first three formant frequencies (F1, F2, F3) over time; display the original audio and the audio after adding noise respectively for easy comparison.
[0080] like Figure 3The figure shows the original audio waveform of a specific embodiment of the present application. The fundamental frequency and formant frequency are stable, and the speech features are clear. The horizontal axis represents time (s), and the vertical axis represents amplitude. After adding noise to the audio, the fundamental frequency and formant frequency are disturbed, the fundamental frequency is offset or reduced in certain time periods, and the formant frequency fluctuates significantly, such as Figure 4 shown.
[0081] Use the DeepFake model to generate high-quality fake speech with fundamental frequency and formant frequency similar to the original speech, such as Figure 5 As shown, Figure 6-7 The fundamental frequency and resonance peak frequency before and after adding noise forged using the DeepFake model. The blue part represents the fundamental frequency (F0), the red part is the resonance peak 1 frequency (F1) engraved at the same time, the green part is the resonance peak 2 frequency (F2) engraved at the same time, and the yellow part is the resonance peak 3 frequency (F3) engraved at the same time. It can be seen that the forged speech with quantum resonance peak frequency generated by deepfake has obvious noise, the fundamental frequency is almost offset, the human voice is distorted, and the resonance peak frequency is unstable.
[0082] like Figure 8 The original audio spectrum is shown in Fig. 9 The following is the audio spectrum after adding noise. The horizontal axis represents time (Time), the vertical axis represents frequency (Hz), and the colors in the figure from light to dark represent different decibels (dB) of the audio. There is no obvious difference between the original audio spectrum and the audio spectrum after adding noise. Ordinary listeners can hardly detect the obvious difference when listening.
[0083] By adding quantum formant noise to the original audio, and using deepfake to forge the original voice and the voice containing quantum formant noise, the quality of the generated forged voice containing quantum formant noise is significantly reduced, the noise is prominent, and the auditory perception is significantly different. Fig.10 The following figure shows a comparison of the audio waveforms forged by the deepfake model before and after adding noise.
[0084] The method for anti-deep fake speech proposed in this application adds optimized quantum noise generated by quantum neural network to the resonance peak frequency of speech signal. It can effectively interfere with the learning and generation of speech features by deep fake model. Through the design of loss function and adjustment of optimization algorithm, it is ensured that the difference between the speech after adding noise and the original speech is minimal in hearing, and the change of processed speech is not obvious to the listener. In this way, the interference with deep fake model is maximized while minimizing the difference in human ear perception.
[0085] According to another aspect of the present application, there is also provided a device for anti-speech deep fake based on quantum formant perturbation, comprising: a feature extraction module, a noise perturbation module, and a signal reconstruction module;
[0086] The feature extraction module is configured to read the original speech signal and extract the fundamental frequency and the formant frequency of the original speech signal;
[0087] A noise perturbation module is configured to generate quantum noise based on a parameterized quantum circuit; add the quantum noise to the extracted fundamental frequency and resonance peak frequency to obtain a perturbed signal;
[0088] The signal reconstruction module is configured to preprocess the disturbed signal to obtain a reconstructed speech signal.
[0089] Furthermore, according to another aspect of the present application, there is also provided a device for anti-speech deep fake based on quantum formant perturbation, comprising a processor and a memory for storing processor executable instructions. Wherein, the processor is configured to implement any of the aforementioned methods for anti-speech deep fake based on quantum formant perturbation when executing the executable instructions. Here, it should be noted that the number of processors can be one or more.
[0090] As a computer-readable storage medium, the memory can be used to store software programs, computer executable programs and various modules, such as: the program or module corresponding to the method of anti-speech deep fake based on quantum formant perturbation in the embodiment of the present application. The processor executes various functional applications and data processing of the device based on quantum formant perturbation anti-speech deep fake by running the software program or module stored in the memory.
[0091] According to another aspect of the present application, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, any of the aforementioned methods for anti-speech deep fake based on quantum resonance peak perturbation is implemented.
[0092] Application prospects of this application: Personal privacy protection: Provide security protection for personal voice data to prevent malicious cloning or tampering.
[0093] Secure communication: Prevent voice from being intercepted and forged in communications that require high confidentiality.
[0094] Digital copyright: protects the copyright of audio works and prevents unauthorized copying and distribution.
[0095] In summary, the quantum resonance peak perturbation anti-deep fake technology proposed in this application combines the advantages of quantum computing and speech signal processing, providing an innovative and effective solution to the security and ethical issues brought about by deep fakes, and has important theoretical significance and application value.
[0096] The embodiments of the present application have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for anti-speech deep fake based on quantum formant perturbation, characterized in that: include: Reading an original speech signal, and extracting a fundamental frequency and a formant frequency of the original speech signal; Generating quantum noise based on parameterized quantum circuits; Adding the quantum noise to the extracted fundamental frequency and resonance peak frequency to obtain a disturbed signal; The disturbed signal is preprocessed to obtain a reconstructed speech signal.
2. The method according to claim 1, characterized in that When extracting the fundamental frequency and the formant frequency of the original speech signal, the fundamental frequency of each time frame in the original speech signal and the first three formant frequencies of the current time frame are extracted.
3. The method according to claim 1, characterized in that Quantum noise is generated based on a parameterized quantum circuit, and bit values measured by the parameterized quantum circuit are converted into real noise values.
4. The method according to claim 1, characterized in that: The disturbed signal includes: a disturbed base frequency and a disturbed resonance peak frequency; generating an excitation signal based on the disturbed fundamental frequency; A formant filter is constructed based on the perturbed formant frequency.
5. The method according to claim 3, characterized in that: The parameterized quantum circuit is: in, represents the evolution operator of the entire quantum circuit; L represents the number of layers of the circuit; represents the parameterized quantum gate of the lth layer, with parameters 6. The method according to claim 4, characterized in that The method further comprises: generating an excitation signal based on the disturbed fundamental frequency; constructing a formant filter based on the disturbed formant frequency, and obtaining the reconstructed speech signal by passing the excitation signal through the formant filter.
7. The method according to claim 6, characterized in that Optimizing the parameterized quantum circuit, including: Calculating a perceptual difference between the original speech signal and the reconstructed speech signal; Calculating the resonance peak frequency before adding noise and the change amount of the resonance peak frequency after the disturbance as the interference degree; Combining the perceived difference and the interference degree in a weighted manner to obtain a target loss function; Minimize the loss function.
8. A device for anti-speech deep fake based on quantum formant perturbation, characterized in that: include: Feature extraction module, noise disturbance module, signal reconstruction module; The feature extraction module is configured to read the original speech signal and extract the fundamental frequency and formant frequency of the original speech signal; The noise perturbation module is configured to generate quantum noise based on a parameterized quantum circuit; add the quantum noise to the extracted fundamental frequency and resonance peak frequency to obtain a perturbed signal; The signal reconstruction module is configured to preprocess the disturbed signal to obtain a reconstructed speech signal.
9. A device for anti-speech deep fake based on quantum formant perturbation, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 7 when executing the executable instructions.
10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voice signal quantum encryption communication system based on random number
CN108768542A
Synthetic voice processing method and device, storage medium and electronic equipment
CN112289298A
Voice signal detection system and method
US20070100609A1