A method and system for anti-audio cloning of AI voice based on pseudo-acoustic masking
By introducing pseudo-timbres into the audio signal and utilizing the acoustic masking effect, the problem of unstable defense effect in existing technologies is solved, achieving efficient defense against AI voice cloning without affecting sound quality, and possessing high versatility and robustness.
Patent Information
- Application Number
- CN202411873345.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing speech cloning defense technologies suffer from unstable defense effects, poor sound quality, or inadequate detection results. In particular, noise addition methods lead to a decrease in sound quality, watermarks are easily damaged, and feature perturbation methods are difficult to adapt to different speech cloning algorithms.
By introducing pseudo-timbres into audio signals and utilizing the acoustic masking effect, the pseudo-timbres are made imperceptible to the human ear and embedded in inaudible frequency bands. A pseudo-timbre feature dictionary is constructed and distorted modulation is applied to interfere with the AI voice cloning model.
Without affecting sound quality, it achieves effective defense against AI voice cloning, possesses high versatility and robustness, and can resist various voice cloning attacks.
Smart Images

Figure CN119724229B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of audio data processing, and particularly relates to a method and system for pseudo-tone color adversarial AI voice cloning based on acoustic masking. BACKGROUND
[0002] With the progress of artificial intelligence (AI) technology, voice cloning technology has developed rapidly in recent years. Voice cloning technology uses deep learning algorithms to train a large amount of voice data to generate cloned voices that are highly similar to the voices of specific individuals. This technology has broad application prospects in voice interaction, virtual assistants, entertainment, and education. However, the misuse of voice cloning technology also poses risks such as privacy security and identity forgery. For example, malicious users can use cloned voices to commit telecommunications fraud, spread false information, or even affect the authenticity of evidence in the judicial field. Therefore, how to prevent AI models from cloning sensitive audio information has become a technical problem that needs to be solved urgently.
[0003] There are some defense technologies and products against voice cloning on the market, mainly including noise addition method, sound watermarking method, and adversarial method based on feature perturbation.
[0004] The noise addition method adds background noise or adversarial noise to the original audio, making it difficult for AI voice cloning models to accurately identify the features of the audio. Adversarial noise is generally generated by a generative adversarial network (GAN), which can interfere with AI models without significantly affecting human hearing. However, the noise addition method often needs to introduce high-intensity noise into the audio to achieve the desired defense effect, which will significantly degrade the audio quality and affect user experience. At the same time, with the progress of AI noise reduction technology, some noise reduction algorithms can effectively remove noise, thereby bypassing the defense and weakening its adversarial effect.
[0005] Sound watermarking technology embeds imperceptible watermark features in audio for subsequent detection of whether the audio has been cloned. Watermarks are usually embedded in audio through high-frequency or low-frequency micro changes and can only be detected in synthesized audio. However, the detection of sound watermarking relies on the integrity of the audio, and once the audio is compressed or converted, the watermark is easily damaged, thereby affecting its detection effect. In addition, the audio after watermark embedding can still be trained and copied by AI models, and it is not fundamentally prevented from being cloned. In fact, sound watermarking only works when the fake audio causes serious public opinion.
[0006] The feature perturbation-based adversarial method refers to perturbing audio features to make the AI model unable to correctly restore the features when generating cloned speech. This method interferes with specific frequency bands of the audio signal, such as modulating the fundamental frequency or harmonic characteristics, so that the cloned generated speech is distorted. However, since such a method needs to frequently adjust the interference features, it is difficult to adapt to different speech cloning algorithms, and the defense effect is unstable. SUMMARY
[0007] To solve the above problems, the present application provides a method and system for anti-AI speech cloning of pseudo timbre based on acoustic masking, which introduces pseudo timbre into the audio signal and utilizes the acoustic masking effect of the human auditory system to make these pseudo timbres imperceptible to the human ear, thereby achieving effective interference with the AI speech cloning model.
[0008] The first aspect of the present application provides a method for anti-AI speech cloning of pseudo timbre based on acoustic masking, mainly comprising:
[0009] Step S1, constructing a pseudo timbre based on the pitch features of a given audio file material;
[0010] Step S2, determining the masking threshold of the target audio according to the intensity and spectral features of the target audio, and marking the frequency band with power lower than the masking threshold as an inaudible frequency band;
[0011] Step S3, embedding the pseudo timbre into the inaudible frequency band of the target audio to obtain a synthesized audio.
[0012] Preferably, step S1 further comprises:
[0013] Step S11, extracting the pitch features in the audio file material, the pitch features at least including the fundamental frequency, formant and harmonic;
[0014] Step S12, calculating the mel spectrum of the audio file material;
[0015] Step S13, extracting the mel frequency cepstral coefficient from the mel spectrum as the key representation of the pseudo timbre, and constructing a pseudo timbre feature dictionary formed by the combination of the mel frequency cepstral coefficient and the modulated pitch features.
[0016] Preferably, in step S13, the modulation of the pitch features includes adjusting the phase and amplitude of the pitch features.
[0017] Preferably, step S2 further comprises:
[0018] Step S21, converting the audio signal of the target audio to the frequency domain to obtain a spectrum;
[0019] Step S22, dividing the spectrum into multiple frequency bands and determining the masking threshold of each frequency band;
[0020] Step S23, obtaining the frequency band index whose power is lower than the masking threshold.
[0021] Preferably, step S3 further comprises:
[0022] Step S30, distortion modulating the target audio.
[0023] Preferably, step S3 further comprises:
[0024] Step S31, traversing the frequency band index of each inaudible frequency band of the target audio, replacing the inaudible frequency band in the frequency spectrum of the target audio with the corresponding part of the pseudo timbre;
[0025] Step S32, inverse transforming the frequency spectrum of the modified target audio to obtain the synthesized audio embedded with the pseudo timbre.
[0026] The second aspect of the present application provides a system for pseudo timbre-based anti-AI voice cloning, mainly comprising:
[0027] a pseudo timbre construction module, configured to construct a pseudo timbre based on the pitch features of a given audio file material;
[0028] an inaudible frequency band determination module, configured to determine the masking threshold of a target audio according to the intensity and spectral features of the target audio, and mark the frequency band whose power is lower than the masking threshold as an inaudible frequency band;
[0029] an audio synthesis module, configured to embed the pseudo timbre into the inaudible frequency band of the target audio to obtain a synthesized audio.
[0030] Preferably, the pseudo timbre construction module comprises:
[0031] a pitch measurement unit, configured to extract pitch features in the audio file material, the pitch features at least including fundamental frequency, formant and harmonic;
[0032] a mel spectrum calculation unit, configured to calculate the mel spectrum of the audio file material;
[0033] a pseudo timbre feature dictionary construction unit, configured to extract mel frequency cepstral coefficients from the mel spectrum as key representations of the pseudo timbre, and construct a pseudo timbre feature dictionary formed by the combination of the mel frequency cepstral coefficients and the modulated pitch features.
[0034] Preferably, the inaudible frequency band determination module comprises:
[0035] a spectrum acquisition unit, configured to convert the audio signal of the target audio to the frequency domain to obtain a frequency spectrum;
[0036] The frequency band masking threshold calculation unit is configured to divide the frequency spectrum into a plurality of frequency bands and determine a masking threshold for each frequency band.
[0037] The frequency band index acquisition unit is configured to acquire the frequency band index of the frequency band with a power lower than the masking threshold.
[0038] Preferably, the audio synthesis module comprises:
[0039] The replacement unit is configured to replace the inaudible frequency band in the frequency spectrum of the target audio with the corresponding part of the pseudo timbre.
[0040] The Fourier transform unit is configured to inversely transform the spectrum of the modified target audio to obtain the synthesized audio with the pseudo timbre embedded.
[0041] The present application can actively defend against AI voice cloning without affecting the sound quality, has high universality and robustness, and can resist various voice cloning attacks. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 The flowchart of a preferred embodiment of the method for resisting AI voice cloning by the pseudo timbre based on acoustic masking according to the present application.
[0043] Figure 2 The principle diagram of the pseudo timbre synthesis based on the masking threshold according to the present application. DETAILED DESCRIPTION
[0044] To make the purposes, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described in more detail below with reference to the drawings in the embodiments of the present application. In the drawings, the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The described embodiments are part of the embodiments of the present application, not all of the embodiments of the present application. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application. The embodiments of the present application will be described in detail below with reference to the drawings.
[0045] The first aspect of the present application provides a method for resisting AI voice cloning by a pseudo timbre based on acoustic masking, as shown in Figure 1 The method mainly comprises the following steps:
[0046] Step S1: Constructing a pseudo timbre based on the pitch features of a given audio file material.
[0047] The step obtains a pseudo timbre by processing the audio file material, so as to be fused with the target audio. In some optional embodiments, step S1 further comprises:
[0048] Step S11, extracting the pitch features in the audio file material, the pitch features at least including the fundamental frequency, the formant and the harmonic;
[0049] Step S12, calculating the mel spectrum of the audio file material;
[0050] Step S13, extracting the mel frequency cepstral coefficient from the mel spectrum as the key representation of the pseudo timbre, and constructing a pseudo timbre feature dictionary formed by the combination of the mel frequency cepstral coefficient and the modulated pitch features.
[0051] Firstly, in step S11, the audio file material is analyzed to extract its pitch features such as the fundamental frequency, the formant and the harmonic. In step S12, the mel spectrum mel_spec of the audio is calculated to capture the time-frequency features of the audio. Then in step S13, the mel frequency cepstral coefficient is extracted from the mel spectrum as the key representation of the timbre, and combined with the pitch features in step S1 to form a pseudo timbre feature dictionary for subsequent pseudo timbre embedding.
[0052] In some optional embodiments, the modulation of the pitch features in step S13 includes adjusting the phase and amplitude of the pitch features.
[0053] The audio file material referred to in step S1 can be the same as the target audio in step S2, or can be other audio files different from the target audio in step S2. In this embodiment, when the same audio file as the target audio is used, the audio modulation can be performed thereon. Accordingly, the phase and amplitude of the audio signal can be adjusted slightly here, so that the modulated pseudo timbre is inconsistent with the target audio, thereby generating false features. After these false features replace the real features of the original target audio in step S3, the AI model can produce a deviation when learning the audio features, so that the original sound cannot be accurately cloned.
[0054] Step S2, determining the masking threshold of the target audio according to the intensity and spectral features of the target audio, and marking the frequency band with power lower than the masking threshold as an inaudible frequency band.
[0055] After the pseudo timbre is given in step S1, step S2 is used to determine which data in the target audio can be replaced by the pseudo timbre in step S1. Specifically, step S2 is based on the masking effect curve of the human ear at different frequency bands to find the frequency bands in the original target audio that are difficult to be perceived by the human ear, and mark them as inaudible frequency bands. These frequency bands can be modified at will.
[0056] In some optional embodiments, step S2 further comprises:
[0057] Step S21, converting the audio signal of the target audio to the frequency domain to obtain a frequency spectrum;
[0058] Step S22, dividing the frequency spectrum into multiple frequency bands and determining the masking threshold of each frequency band;
[0059] Step S23, obtaining the frequency band index whose power is lower than the masking threshold.
[0060] In this embodiment, after loading the target audio file, the audio signal and the sampling rate are obtained, and after converting the audio signal collected in each period to the frequency domain, the frequency spectrum is obtained, Figure 2 The relationship between the frequency spectrum and the sound intensity is given, Figure 2 In the figure, the abscissa is the frequency value of the audio signal, and the ordinate is the sound intensity of the audio signal, which can be regarded as the energy value of the frequency. When the sound intensity of a frequency is less than a certain threshold, the human ear cannot hear it, as shown by the dashed line in the figure. For example, at a frequency of 0.02 kHz, the sound intensity of 20 Hz frequency needs to reach 70 decibels to be heard. If the sound intensity is lower than 70 decibels, it cannot be heard. In addition, when the sound signal of a certain frequency has a large energy, that is, the sound signal intensity of a certain frequency is large, the masking threshold near the frequency will be greatly increased. For example, at a position of about 0.3 kHz, the energy of the sound signal of the frequency is large, about 60 decibels. At this time, the masking threshold of the signal near the frequency (0.3 kHz) will be increased, as shown by the solid line in the figure. Figure 2
[0061] Accordingly, based on the acoustic masking effect of the human ear and considering the curve of the relationship between the volume and the frequency, a masking curve shown in the figure is constructed in the frequency domain. Figure 2 Accordingly, based on the acoustic masking effect of the human ear and considering the curve of the relationship between the volume and the frequency, a masking curve shown in the figure is constructed in the frequency domain.
[0062] It should be further explained that the intensity of the pseudo tone color of each frequency band can be fine-tuned through weighted mapping to ensure the best interference effect without affecting the perception of the human ear. Through the weighted and mapping operations, the system can optimize the amplitude of the pseudo tone color in different frequency bands, so that the output audio has high quality and high defense.
[0063] Step S3, embedding the pseudo tone color into the inaudible frequency band of the target audio to obtain a synthesized audio.
[0064] In some optional embodiments, step S3 further includes:
[0065] Step S30, performing distortion modulation on the target audio.
[0066] To enhance the universality and anti-attack ability of the defense, this embodiment introduces a distortion layer before the pseudo-tone color embedding, which is embedded into the audio processing for modulation of the audio features in the Mel spectrum stage. Through slight distortion modulation, the complexity of the audio features is further increased, making it difficult for the AI voice cloning model to accurately restore the original features of the audio. The processing of the distortion layer mainly includes modulating the feature part of the target audio that is not replaced by the pseudo-tone color, ensuring that it can interfere in different cloning attacks. The distortion layer modulation effect significantly improves the defense effect without affecting the human ear's hearing.
[0067] In some optional embodiments, step S3 further comprises:
[0068] Step S31, traverse the frequency band index of each inaudible frequency band of the target audio, and replace the inaudible frequency band in the frequency spectrum of the target audio with the corresponding part of the pseudo-tone color;
[0069] Step S32, inverse transform the frequency spectrum of the modified target audio to obtain the synthesized audio embedded with the pseudo-tone color.
[0070] In step S31, it should be noted that the corresponding part of the pseudo-tone color refers to the frequency spectrum data belonging to the same frequency band as the target audio, and the sound intensity of these frequency spectrum data should be below the masking threshold. Since the audio file material used to generate the pseudo-tone color is usually consistent with the target audio, during the processing of the target audio and the audio file material, when it is determined that the intensity of certain frequency bands of the frequency spectrum data of the target audio is below the masking threshold, the intensity of the corresponding frequency bands of the audio file material is also below the masking threshold. Therefore, in step S31, only the frequency band index is needed to replace the data of the corresponding frequency band of the target audio with the data of the corresponding frequency band of the pseudo-tone color. Conversely, when the audio file material is inconsistent with the target audio, it should be ensured that the pseudo-tone color constructed from the audio file material is all inaudible audio, i.e. below the masking threshold.
[0071] Finally, after the audio synthesis is completed, it can be analyzed by short-time Fourier transform (STFT) to verify the embedding effect of the pseudo-tone color. The STFT analysis provides the frequency domain feature map of the audio, which can visually display the distribution of the pseudo-tone color in the audio frequency spectrum and ensure that the generated defense audio meets the design requirements.
[0072] The second aspect of the present application provides a system for anti-AI voice cloning based on acoustic masking and pseudo-tone color, which mainly comprises:
[0073] a pseudo-tone color construction module for constructing a pseudo-tone color based on the pitch features of a given audio file material;
[0074] The inaudible frequency band determination module is configured to determine a masking threshold of the target audio according to the intensity and spectral characteristics of the target audio, and mark a frequency band with power lower than the masking threshold as an inaudible frequency band.
[0075] The audio synthesis module is configured to embed the pseudo timbre into the inaudible frequency band of the target audio to obtain a synthesized audio.
[0076] In some optional embodiments, the pseudo timbre construction module comprises:
[0077] The pitch measurement unit is configured to extract pitch characteristics in the audio file material, the pitch characteristics comprising at least a fundamental frequency, a formant and a harmonic;
[0078] The mel-spectrum calculation unit is configured to calculate a mel-spectrum of the audio file material;
[0079] The pseudo timbre feature dictionary construction unit is configured to extract mel-frequency cepstral coefficients from the mel-spectrum as key representations of the pseudo timbre, and construct a pseudo timbre feature dictionary formed by the mel-frequency cepstral coefficients and modulated pitch characteristics.
[0080] In some optional embodiments, the inaudible frequency band determination module comprises:
[0081] The spectrum acquisition unit is configured to convert an audio signal of the target audio into a frequency domain to obtain a spectrum;
[0082] The per-frequency band masking threshold calculation unit is configured to divide the spectrum into a plurality of frequency bands, and determine a masking threshold of each frequency band;
[0083] The frequency band index acquisition unit is configured to acquire a frequency band index of a frequency band with power lower than the masking threshold.
[0084] In some optional embodiments, the audio synthesis module comprises:
[0085] The replacement unit is configured to traverse the frequency band index of each inaudible frequency band of the target audio, and replace the inaudible frequency band in the spectrum of the target audio with a corresponding part of the pseudo timbre;
[0086] The Fourier transform unit is configured to inversely transform the spectrum of the modified target audio to obtain a synthesized audio with the pseudo timbre embedded.
[0087] The above merely provides a specific implementation of the present application, but the protection scope of the present application is not limited to this. Any changes or replacements within the technical scope disclosed by the present application can be easily thought of by those skilled in the art, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for combating AI speech cloning using pseudo-timbres based on acoustic masking, characterized in that, include: Step S1: Construct pseudo-timbres based on the tonal features of the given audio file material; Step S2: Determine the masking threshold of the target audio based on the intensity and spectral characteristics of the target audio, and mark the frequency bands with power lower than the masking threshold as inaudible frequency bands; Step S3: Embed the pseudo-timbre into the inaudible frequency band of the target audio to obtain the synthesized audio; Step S1 further includes: Step S11: Extract the pitch features from the audio file material, wherein the pitch features include at least the fundamental frequency, formants, and harmonics; Step S12: Calculate the Mel spectrum of the audio file material; Step S13: Extract the Mel frequency cepstral coefficients from the Mel spectrum as key representations of the pseudo-timbre, and construct a pseudo-timbre feature dictionary formed by the combination of Mel frequency cepstral coefficients and modulated pitch features.
2. The method for combating AI speech cloning based on acoustic masking pseudo-timbre as described in claim 1, characterized in that, In step S13, the modulation of the pitch feature includes adjusting the phase and amplitude of the pitch feature.
3. The method for combating AI speech cloning based on acoustic masking pseudo-timbre as described in claim 1, characterized in that, Step S2 further includes: Step S21: Convert the audio signal of the target audio to the frequency domain to obtain the spectrum; Step S22: Divide the spectrum into multiple frequency bands and determine the masking threshold for each frequency band; Step S23: Obtain the frequency band index with power lower than the masking threshold.
4. The method for combating AI speech cloning based on acoustic masking pseudo-timbre as described in claim 1, characterized in that, Step S3 further includes: Step S30: Distortion modulation is applied to the target audio.
5. The method for combating AI speech cloning based on acoustic masking pseudo-timbre as described in claim 1, characterized in that, Step S3 further includes: Step S31: Traverse the frequency band index of each inaudible frequency band of the target audio and replace the inaudible frequency band in the spectrum of the target audio with the corresponding part of the pseudo timbre; Step S32: Perform an inverse transform on the spectrum of the modified target audio to obtain a synthesized audio with embedded pseudo-timbres.
6. A system for combating AI voice cloning using pseudo-voice based on acoustic masking, characterized in that, include: The pseudo-timbre construction module is used to construct pseudo-timbres based on the tonal features of a given audio file. The inaudible frequency band determination module is used to determine the masking threshold of the target audio based on the intensity and spectral characteristics of the target audio, and to mark frequency bands with power below the masking threshold as inaudible frequency bands; An audio synthesis module is used to embed the pseudo-timbre into the inaudible frequency band of the target audio to obtain synthesized audio; The pseudo-timbre construction module includes: A pitch measurement unit is used to extract pitch features from audio file materials, wherein the pitch features include at least the fundamental frequency, formants, and harmonics; Mel spectrum calculation unit, used to calculate the Mel spectrum of audio file material; The pseudo-timbre feature dictionary construction unit is used to extract Mel frequency cepstral coefficients from the Mel spectrum as key representations of pseudo-timbre, and to construct a pseudo-timbre feature dictionary formed by the combination of Mel frequency cepstral coefficients and modulated pitch features.
7. The system for combating AI voice cloning based on acoustic masking pseudo-timbre as described in claim 6, characterized in that, The inaudible frequency band determination module includes: The spectrum acquisition unit is used to convert the audio signal of the target audio to the frequency domain to obtain the spectrum; The masking threshold calculation unit for each frequency band is used to divide the spectrum into multiple frequency bands and determine the masking threshold for each frequency band. The frequency band index acquisition unit is used to acquire the frequency band indexes whose power is lower than the masking threshold.
8. The system for combating AI voice cloning based on acoustic masking pseudo-timbre as described in claim 6, characterized in that, The audio synthesis module includes: The replacement unit is used to traverse the frequency band index of each inaudible frequency band of the target audio and replace the inaudible frequency band in the spectrum of the target audio with the corresponding part of the pseudo timbre; The Fourier transform unit is used to perform an inverse transform on the spectrum of the modified target audio to obtain a synthesized audio with embedded pseudo-timbres.
Citation Information
Patent Citations
Microphone for resisting AI voice cloning by pseudo tone based on acoustic shielding
CN119724228A