Method, device and equipment for anti-voice deepfake based on quantum resonance peak disturbance

By using the quantum resonant perturbation method, quantum noise is generated using parameterized quantum circuits to interfere with the speech feature learning of deepfake models. This solves the problem of the difficulty in effectively interfering with deepfakes in existing technologies and increases the difficulty of attacks without affecting the sound quality.

CN119993113BActive Publication Date: 2026-04-10RELATED (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RELATED (BEIJING) TECHNOLOGY CO LTD
Filing Date
2025-02-11
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing deepfake detection algorithms may be bypassed by new forgery techniques, and traditional methods may affect audio quality or be ineffective, making it difficult to effectively interfere with the deepfake model's learning and generation of speech features.

Method used

By using a quantum formant perturbation-based method, quantum noise is generated using parameterized quantum circuits and added to the fundamental frequency and formant frequency of the speech signal. Combined with optimization algorithms and loss functions, a reconstructed speech signal is generated to interfere with deepfake models.

Benefits of technology

It effectively interferes with the speech feature learning of deepfake models, increasing the difficulty of attacks, while keeping the sound quality unaffected, making the difference difficult for the human ear to detect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993113B_ABST
    Figure CN119993113B_ABST
Patent Text Reader

Abstract

The application relates to a method, device and equipment for anti-deepfake speech based on quantum resonance peak disturbance, wherein the method comprises the following steps: reading an original speech signal, extracting a fundamental frequency and a formant frequency of the original speech signal; creating a parameterized quantum circuit to generate quantum noise; adding the quantum noise to the extracted fundamental frequency and formant frequency to obtain a disturbed signal; and preprocessing the disturbed signal to obtain a reconstructed speech signal. The optimized quantum noise generated by the quantum neural network can effectively interfere with the learning and generation of speech features by a deepfake model. Through the design of a loss function and the adjustment of an optimization algorithm, the speech after adding the noise is ensured to have the smallest difference in hearing from the original speech.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information security, and in particular to a method, device and equipment for anti-voice deepfake based on quantum formant disturbance. BACKGROUND

[0002] With the rapid development of artificial intelligence and deep learning technology, voice deepfake technology has been able to generate highly realistic voice and video content. Current deepfake detection algorithms can be bypassed by new and more advanced fake technologies. Traditional methods of adding noise to audio can reduce audio quality and affect user experience. Generating adversarial samples requires understanding the details of the attack model, and may not work well for different models. SUMMARY

[0003] Therefore, the present application provides a method, device and equipment for anti-voice deepfake based on quantum formant disturbance. It is suitable for interfering with the learning and generation of voice features by a deepfake model.

[0004] According to an aspect of the present application, a method for anti-voice deepfake based on quantum formant disturbance is provided, comprising: reading an original voice signal, extracting the fundamental frequency and formant frequency of the original voice signal;

[0005] Generating quantum noise based on a parameterized quantum circuit; adding the quantum noise to the extracted fundamental frequency and formant frequency to obtain a disturbed signal;

[0006] Preprocessing the disturbed signal to obtain a reconstructed voice signal.

[0007] In one possible implementation, when extracting the fundamental frequency and formant frequency of the original voice signal, the fundamental frequency of each time frame and the first three formant frequencies of the current time frame in the original voice signal are extracted.

[0008] In one possible implementation, quantum noise is generated based on a parameterized quantum circuit, and the measurement bit value of the parameterized quantum circuit is converted into a real noise value.

[0009] In one possible implementation, the disturbed signal includes a disturbed fundamental frequency and a disturbed formant frequency; an excitation signal is generated based on the disturbed fundamental frequency; and a formant filter is constructed based on the disturbed formant frequency.

[0010] In one possible implementation, the parameterized quantum circuit is:

[0011]

[0012] wherein, represents an evolution operator of the entire quantum circuit; L represents the number of layers of the circuit; represents a parameterized quantum gate of the lth layer, and the parameters are .

[0013] In a possible implementation, an excitation signal is generated based on the perturbed fundamental frequency; a formant filter is constructed based on the perturbed formant frequency, and the method further includes: passing the excitation signal through the formant filter to obtain the reconstructed speech signal.

[0014] In a possible implementation, the parameterized quantum circuit is optimized, including:

[0015] A perceptual difference between the original speech signal and the reconstructed speech signal is calculated;

[0016] A change amount of the formant frequency before adding noise and the perturbed formant frequency is calculated as an interference degree; a target loss function is obtained by combining the perceptual difference and the interference degree in a weighted manner; and the loss function is minimized.

[0017] According to another aspect of the present application, an apparatus for anti-speech deepfake based on quantum formant perturbation is provided, including: a feature extraction module, a noise perturbation module, and a signal reconstruction module.

[0018] The feature extraction module is configured to read an original speech signal, and extract a fundamental frequency and a formant frequency of the original speech signal.

[0019] The noise perturbation module is configured to generate quantum noise based on a parameterized quantum circuit, and add the quantum noise to the extracted fundamental frequency and formant frequency to obtain a perturbed signal.

[0020] The signal reconstruction module is configured to pre-process the perturbed signal to obtain a reconstructed speech signal.

[0021] According to another aspect of the present application, an apparatus for anti-speech deepfake based on quantum formant perturbation is provided, including: a processor; a memory for storing processor-executable instructions; and wherein the processor is configured to execute the above method.

[0022] According to another aspect of the present application, a non-volatile computer readable storage medium having computer program instructions stored thereon is provided, wherein the computer program instructions are executed by a processor to implement the above method.

[0023] The beneficial effects of the present application: the optimized quantum noise generated by the quantum neural network can effectively interfere with the learning and generation of voice features by the deepfake model. Adding noise on the formant frequency of the voice signal can precisely interfere with the key features of the voice, which is better than directly adding noise, and increases the difficulty of attack of the deepfake model. Through the design of the loss function and the adjustment of the optimization algorithm, the human ear perception difference and the interference effect are considered to guide the optimization process. The noise can both interfere with the model and not affect the sound quality. The linear predictive coding is used to reconstruct the voice signal to ensure that the voice after adding noise has the smallest difference in hearing with the original voice.

[0024] Other features and aspects of the present application will become apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the present application and serve to explain the principles of the present application.

[0026] Figure 1 A flow chart of a method for anti-voice deepfake based on quantum formant disturbance according to an embodiment of the present application is shown;

[0027] Figure 2 A quantum circuit schematic diagram according to an embodiment of the present application is shown;

[0028] Figure 3 An original audio waveform diagram according to an embodiment of the present application is shown;

[0029] Figure 4 An audio waveform diagram after adding noise according to an embodiment of the present application is shown;

[0030] Figure 5 A deepfake model generated audio fundamental frequency and formant frequency diagram according to an embodiment of the present application is shown;

[0031] Figure 6 A fundamental frequency and formant frequency diagram of deepfake model forged audio before adding noise according to an embodiment of the present application is shown;

[0032] Figure 7 A fundamental frequency and formant frequency diagram of deepfake model forged audio after adding noise according to an embodiment of the present application is shown;

[0033] Figure 8 An original audio spectrum diagram according to an embodiment of the present application is shown;

[0034] Figure 9 An audio spectrum diagram after adding noise according to an embodiment of the present application is shown;

[0035] Figure 10A comparison chart of audio waveforms of deepfake model forgery before and after adding noise according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0036] Various exemplary embodiments, features, and aspects of the present application will be explained in greater detail below with reference to the accompanying drawings. Like reference numerals may be used to refer to like elements throughout. While various aspects of embodiments are illustrated, the embodiments need not be used to scale unless specifically indicated.

[0037] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0038] In addition, for the purpose of convenience and brevity, detailed descriptions of well-known functions and structures incorporated in the application are omitted. It will be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described herein, embody the principles of the application and fall within the spirit and scope of the application.

[0039] The present application is applicable to interfere with the learning and generation of speech features by a deep fake model, and to ensure that the speech after adding noise has the smallest difference in hearing with the original speech.

[0040] Embodiment 1

[0041] Figure 1 A flowchart of a method of anti-speech deep fake based on quantum formant perturbation according to an embodiment of the present application is shown. As shown in Figure 1 The method comprises:

[0042] Step S100, read the original speech signal, and extract the fundamental frequency and formant frequency of the original speech signal. The fundamental frequency (F0) is generated by vocal cord vibration and determines the pitch of the sound. Formants are generated by the resonance of the vocal tract and are local peaks in the speech spectrum, which mainly determine the timbre and tone quality of the speech.

[0043] Step S200, generating quantum noise based on a parameterized quantum circuit. The parameterized quantum circuit refers to a quantum neural network (QNN) composed of a number of qubits and parameterized quantum gates. The quantum neural network is a machine learning model using a parameterized quantum circuit. By adjusting the circuit parameters, the QNN can generate a specific probability distribution. The angles of quantum gates (such as Hadamard gates, RX rotation gates) in the quantum circuit are adjustable parameters, wherein the Hadamard gate places the quantum bit in a superposition state, so that the measurement result has randomness. The RX rotation gate is a rotation gate around the X axis, and the rotation angle is an adjustable parameter. These parameters can be trained by an optimization algorithm. By utilizing the randomness and complexity of quantum computing, quantum noise that is difficult to predict is generated.

[0044] Step S300, adding quantum noise to the extracted fundamental frequency and formant frequency to obtain a perturbed signal. The perturbed signal refers to the fundamental frequency and formant frequency after adding noise respectively.

[0045] Step S400, pre-processing the perturbed signal to obtain a reconstructed speech signal. The preprocessing refers to generating an excitation signal and a formant filter based on the perturbed signal, which is used to merge the speech signal.

[0046] In one possible implementation, an audio file is loaded, and audio data and a sampling rate are obtained. An acoustic analysis tool (such as Praat or Parselmouth library) is used to extract the fundamental frequency F0(t) of each time frame in the original speech signal and the first three formant frequencies F1(t), F2(t), F3(t) of the current time frame. Wherein, t represents the time frame index, F0(t) is the fundamental frequency at time t, and Fi(t) is the i-th formant frequency at time t, i = 1, 2, 3.

[0047] Adding quantum noise to the formant frequencies of the speech signal can directly interfere with the learning and replication of speech features by the deepfake model.

[0048] Specifically, the audio signal is divided into frames with a length of 10 milliseconds, and each frame is traversed: the fundamental frequency is extracted using the To Pitch method; and the formant frequencies are extracted using the To Formant (burg) method.

[0049] The extracted fundamental frequency and formant frequencies are stored as an array for subsequent processing.

[0050] A quantum neural network (QNN) is used to generate quantum noise. The QNN is a parameterized quantum circuit that uses a ParameterVector to create a parameter vector theta that is used to parameterize the rotation gates. A Hadamard gate and an RX rotation gate are applied to each qubit to form a parameterized quantum circuit. A sampler is created using the Sampler primitive of Qiskit, and a SamplerQNN object is created based on the circuit and the sampler. The quantum circuit is printed and plotted for visualization.

[0051] The parameterized quantum circuit is:

[0052]

[0053] wherein, represents the evolution operator of the entire quantum circuit; L represents the number of layers of the circuit; represents the parameterized quantum gate of the lth layer, and the parameters are .

[0054] By sampling the output probability distribution of the quantum circuit, random noise that meets quantum characteristics can be obtained. The output quantum state of the quantum circuit is measured to obtain random measurement results. The bit string b obtained by measurement is converted into a real noise value:

[0055]

[0056] wherein, N(t) is the noise value at time t; int(b) represents converting the bit string b into an integer; n is the number of qubits; and scale is the amplification coefficient of the noise.

[0057] As shown in Figure 2 , the quantum circuit diagram of an embodiment of the present application includes five qubits and corresponding gates. The Hadamard gate is used to place the qubit in a superposition state, so that the measurement result has randomness. The RX rotation gate is a rotation gate around the X axis, and the rotation angle is an adjustable parameter.

[0058] The quantum noise is added to the extracted fundamental frequency and the first three resonance peak frequencies to obtain the disturbed fundamental frequency and the disturbed resonance peak frequency. The formula for adding noise is as follows:

[0059] The disturbed fundamental frequency is: ;

[0060] The disturbed resonance peak frequency is: ;

[0061] wherein, is the disturbed fundamental frequency at time t; is the i-th formant frequency after the perturbation at time t; , represents the quantum noise sequence generated by the QNN, the value at time t. , is a parameter to control the base frequency and formant frequency noise intensity.

[0062] The forward method of the QNN is used to calculate the probability distribution of the output. For each probability distribution, a measurement is taken and mapped to a noise value. The generated noise sequence is added to the corresponding formant frequency, and np.clip is used to ensure that the frequency value is non-negative. The weight coefficients are updated by training the quantum neural network algorithm, and the weight coefficients determine the proportion of noise sequence added.

[0063] based on the base frequency after the perturbation generate an excitation signal; based on the formant frequency after the perturbation construct a formant filter; pass the excitation signal through the formant filter to obtain a reconstructed speech signal. Unlike directly adding noise to the audio signal, adding quantum noise to the formant frequency of the speech signal precisely interferes with the key features of the speech, increasing the difficulty of attack on the deepfake model.

[0064] Further, using the linear predictive coding (LPC) method, each frame of speech signal is filtered to simulate the filtering effect of the vocal tract. White noise is generated as an excitation source, and the synthesized speech signal is obtained by filtering. The synthesized audio signal is normalized to prevent signal overload.

[0065] Further, the parameters of the above quantum circuit are optimized by training. Specifically, the perceptual difference between the original speech signal and the reconstructed speech signal is calculated.

[0066] By comparing the Mel frequency cepstrum coefficient (MFCC) of the original speech signal and the reconstructed speech signal, the difference between the two is calculated:

[0067]

[0068] T is the total number of time frames; is the parameter vector of the quantum circuit; is the MFCC feature of the original audio at time t; is the MFCC feature of the processed audio at time t. The MFCC features of the original audio and the synthesized audio are extracted, and the average absolute difference of the MFCC features is calculated as a measure of perceptual difference.

[0069] Calculate the change in the formant frequency before adding noise and the formant frequency after the perturbation as the degree of interference:

[0070]

[0071] where F i (t) is the original i-th formant frequency; F i (t) is the perturbed i-th formant frequency, depending on the parameters .

[0072] The target loss function is obtained by combining the perceptual difference and the interference degree in a weighted manner;

[0073]

[0074] λ ∈ [0, 1] is a weight parameter to balance the influence between P and D. The quantum circuit parameters are optimized using classical optimization algorithms (such as gradient descent, Adam) to minimize the loss function . .

[0075] In each optimization iteration, the parameters are updated according to the gradient of the loss function :

[0076]

[0077] where is the learning rate; is the gradient of the loss function with respect to the parameter .

[0078] By defining a loss function, the perceptual difference between the noise-added audio and the original audio and the interference degree to the deepfake model are measured, and the optimal quantum circuit parameters are found through optimization algorithms. The noise added to the formant frequency under this optimal parameter can balance the auditory difference and the interference effect. The COBYLA method (evolutionary algorithm based on constrained optimization) is used to optimize without giving the gradient. By setting the maximum number of iterations, the time of the optimization process is limited. By optimizing the parameters of the quantum circuit, the generated quantum noise maximizes the interference effect while minimizing the impact on the human ear, achieving a balance between interference and sound quality.

[0079] The processed speech file is saved, and the waveform plots of the original speech signal and the perturbed speech signal, as well as the changes in the fundamental frequency and the formant frequencies, are drawn to help analyze and understand the processing effect. Specifically, matplotlib is used to draw the time-domain waveforms of the original speech signal and the perturbed speech signal; the curves of the fundamental frequency (F0) and the first three formant frequencies (F1, F2, F3) over time are drawn; the original audio and the audio after adding noise are displayed respectively, which facilitates comparison.

[0080] As Figure 3The original audio waveform of an embodiment of the present application is shown, the fundamental frequency and formant frequency are stable, and the voice features are clear. The abscissa represents time (s), and the ordinate represents amplitude. After adding noise to the audio, the fundamental frequency and formant frequency are disturbed, the fundamental frequency is canceled or reduced in some time periods, and the formant frequency fluctuates obviously, as shown in Figure 4 .

[0081] A high-quality fake voice is generated using a DeepFake model, and the fundamental frequency and formant frequency are similar to those of the original voice, as shown in Figure 5 , Figure 6-7 The fundamental frequency and formant frequency before and after adding noise using the DeepFake model for forgery, the blue part represents the fundamental frequency (F0), the red part represents the formant 1 frequency (F1) at the same time, the green part represents the formant 2 frequency (F2) at the same time, and the yellow part represents the formant 3 frequency (F3) at the same time. It can be seen that the fake voice with quantum formant frequency generated by deepfake has obvious noise, the fundamental frequency is almost canceled, the human voice is distorted, and the formant frequency is unstable.

[0082] As shown in Figure 8 , the original audio spectrum is shown, Figure 9 , and the audio spectrum after adding noise. Among them, the abscissa represents time (Time), and the ordinate represents frequency (Hz). The color in the figure represents different decibels (dB) of the audio from light to dark. The original audio spectrum and the audio spectrum after adding noise have no obvious difference, and ordinary listeners are almost difficult to detect obvious differences when listening.

[0083] The original audio is fused with the added quantum formant noise, and the original voice and the voice containing quantum formant noise are forged using deepfake. The quality of the fake voice containing quantum formant noise is obviously reduced, the noise is prominent, and the auditory perception difference is significant. As shown in Figure 10 , the deepfake model forgeries of the audio waveforms before and after adding noise are compared.

[0084] The method of anti-voice deep forgery proposed in the present application adds optimized quantum noise generated by a quantum neural network to the formant frequency of the voice signal. It can effectively interfere with the learning and generation of voice features by the deep forgery model. Through the design of the loss function and the adjustment of the optimization algorithm, it is ensured that the voice after adding noise has the smallest difference in hearing with the original voice, and the listener does not obviously change the processed voice. In the case of minimizing the difference in human ear perception, the interference to the deep forgery model is maximized.

[0085] According to another aspect of the present application, an anti-voice deep forgery device based on quantum formant disturbance is also provided, comprising a feature extraction module, a noise disturbance module, and a signal reconstruction module.

[0086] a feature extraction module configured to read the original speech signal, extract the fundamental frequency and formant frequencies of the original speech signal;

[0087] a noise disturbance module configured to generate quantum noise based on the parameterized quantum circuit; add the quantum noise to the extracted fundamental frequency and formant frequencies to obtain a disturbed signal;

[0088] a signal reconstruction module configured to preprocess the disturbed signal to obtain a reconstructed speech signal.

[0089] Further, according to another aspect of the present application, there is also provided an anti-speech deepfake device based on quantum formant disturbance, comprising a processor and a memory for storing processor-executable instructions. Wherein the processor is configured to implement the anti-speech deepfake method based on quantum formant disturbance as described above when executing the executable instructions. It should be pointed out here that the number of processors can be one or more.

[0090] The memory, as a kind of computer readable storage medium, can be used to store software programs, computer executable programs and various modules, such as the programs or modules corresponding to the anti-speech deepfake method based on quantum formant disturbance of the embodiments of the present application. The processor executes the software programs or modules stored in the memory, thereby performing various functional applications and data processing of the anti-speech deepfake device based on quantum formant disturbance.

[0091] According to another aspect of the present application, there is also provided a non-volatile computer readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the anti-speech deepfake method based on quantum formant disturbance as described above.

[0092] Application prospect of the present application: personal privacy protection: providing security protection for personal voice data, preventing malicious cloning or tampering.

[0093] Secure communication: in the communication that needs high confidentiality, prevent the speech from being intercepted and counterfeited.

[0094] Digital copyright: protect the copyright of audio works, prevent unauthorized copying and dissemination.

[0095] In summary, the anti-deepfake technology based on quantum formant disturbance proposed in the present application combines the advantages of quantum computing and speech signal processing, providing an innovative and effective solution to the security and ethical problems caused by deepfake, which has important theoretical significance and application value.

[0096] Having described various embodiments of the application, it is to be understood that the above description is meant not to limit and not to encompass all of the possible embodiments. Many modifications and variations of this application can be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. It is intended that the scope of the application be defined by the scope of the patent and by the claims as allowed by the patent office, which can include adaptations and modifications. It is further intended that each of the individual elements or variations of this application be deemed to be disclosed herein.

Claims

1. A method for anti-voice deepfake based on quantum resonance peak perturbation, characterized in that, The method comprises: reading an original speech signal, extracting a fundamental frequency and a formant frequency of the original speech signal; generating quantum noise based on a parameterized quantum circuit, wherein an output quantum state of the parameterized quantum circuit is measured to obtain a random measurement result, and a bit value obtained by the measurement is converted into a real noise value; adding the quantum noise to the extracted fundamental frequency and formant frequency to obtain a perturbed signal; preprocessing the perturbed signal to obtain a reconstructed speech signal.

2. The method of claim 1, wherein, When extracting the fundamental frequency and the formant frequency of the original speech signal, the fundamental frequency of each time frame and the first three formant frequencies of the current time frame in the original speech signal are extracted.

3. The method of claim 1, wherein, The perturbed signal comprises a perturbed fundamental frequency and a perturbed formant frequency. An excitation signal is generated based on the perturbed fundamental frequency. A formant filter is constructed based on the perturbed formant frequency.

4. The method of claim 1, wherein, The parameterized quantum circuit is: wherein, represents the evolution operator of the entire quantum circuit; L represents the number of layers of the circuit; represents the parameterized quantum gate of the l-th layer, with parameters .

5. The method of claim 3, wherein, An excitation signal is generated based on the perturbed fundamental frequency, and a formant filter is constructed based on the perturbed formant frequency, and the excitation signal is passed through the formant filter to obtain the reconstructed speech signal.

6. The method of claim 5, wherein, Optimizing the parameterized quantum circuit comprises: calculating a perceptual difference between the original speech signal and the reconstructed speech signal; calculating a change amount of the formant frequency before adding noise and the perturbed formant frequency as an interference degree; combining the perceptual difference and the interference degree in a weighted manner to obtain a target loss function; minimizing the loss function.

7. An apparatus for counteracting voice deepfake based on quantum resonance peak perturbation, characterized in that, The method comprises: a feature extraction module, a noise perturbation module, and a signal reconstruction module; The feature extraction module is configured to read an original speech signal and extract a fundamental frequency and a formant frequency of the original speech signal; The noise perturbation module is configured to generate quantum noise based on a parameterized quantum circuit, wherein an output quantum state of the parameterized quantum circuit is measured to obtain a random measurement result, and a bit value obtained by the measurement is converted into a real noise value; and the quantum noise is added to the extracted fundamental frequency and formant frequency to obtain a perturbed signal; The signal reconstruction module is configured to preprocess the perturbed signal to obtain a reconstructed speech signal.

8. An apparatus for counteracting deepfake speech based on quantum resonance peak perturbation, comprising: The method comprises: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the method of any one of claims 1 to 6 when executing the executable instructions.

9. A non-transitory computer readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by the processor, implement the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice signal quantum encryption communication system based on random number

    CN108768542A

  • Synthetic voice processing method and device, storage medium and electronic equipment

    CN112289298A