A general testing method for audio forgery algorithms

By constructing a universal testing method to evaluate the threats of audio forgery algorithms, the problem that existing technologies cannot comprehensively evaluate audio forgery algorithms in the real world is solved. This enables a comprehensive evaluation of voiceprint recognition, auditory perception, and detection models, and provides a comprehensive security threat assessment of audio forgery algorithms.

CN119673176BActive Publication Date: 2025-09-30ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411832268.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-09-30
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

Existing audio forgery algorithm evaluation methods cannot effectively assess threats in the real world, especially the rapid development of high-precision forged audio technology, which is difficult to fully understand. Existing tools require a large number of voice samples or are only evaluated on training sets, and cannot quickly respond to updates in new technologies.

Method used

A general testing method is constructed to download voice data from open source corpora, generate forged audio, and add channel perturbations to simulate real-world scenarios. This method combines voiceprint recognition, auditory evaluation, and forged audio detection models to evaluate the threats posed by audio forgery algorithms.

Benefits of technology

A comprehensive evaluation of audio forgery algorithms in the real world is achieved, evaluating their deceptiveness to voiceprint recognition models, their concealment to human auditory perception, and their anti-detection capabilities by detection models, providing a comprehensive security threat assessment of audio forgery algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119673176B_ABST
    Figure CN119673176B_ABST
Patent Text Reader

Abstract

The present invention discloses a universal testing method for audio forgery algorithms. The universal testing method for audio forgery algorithms proposed in the present invention, by adding channel perturbation data enhancement to forged audio in S3, simulates the signal changes of forged audio in the real world, and can evaluate the actual threat of a given audio forgery algorithm in the real world. By testing the forged audio set based on the voiceprint recognition model, it can provide the deceptive threat of a given audio forgery algorithm to the real-world voiceprint recognition model and voiceprint recognition service. By evaluating the forged audio set based on the voice hearing test model, it can evaluate the hidden threat of a given audio forgery algorithm to human auditory perception. By detecting the forged audio set based on the forged audio detection model, it can effectively evaluate the anti-detection ability of a given audio forgery algorithm to the detection model, further revealing the comprehensive security threat of audio forgery algorithms in the real world.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of forged audio detection, and in particular to a universal testing method for audio forgery algorithms. Background Art

[0002] Throughout human evolution, voice has been a crucial medium for identifying individuals. In modern society, voiceprint biometrics has been further applied in a variety of authentication scenarios, including social media platforms (such as WeChat) and mobile banking (such as JPMorgan Chase and HSBC), as well as personalized services offered by smart speakers (such as Amazon Alexa and Alibaba Tmall Genie) and voice assistants (such as Apple Siri and Samsung Bixby). However, the rapid development of audio forgery technologies capable of accurately replicating the target speaker's voice has severely challenged the reliability of voice as a biometric method. Malicious exploitation of these technologies has resulted in significant economic losses and even interfered with human cognition.

[0003] Given the increasing prevalence of audio forgery, evaluating such techniques has become crucial for understanding their public impact and developing countermeasures. Early studies evaluated the security of GMM-based voiceprint recognition systems using tools such as Festvox. Shirvanian et al. quantified the vulnerability of five mobile voice authentication applications. Recent studies have also tested early deep learning-based audio forgery techniques (such as AutoVC and SV2TTS) on real-world platforms such as Azure, WeChat, and Alexa. However, these studies either evaluated only on speakers seen in the training set or required speech samples ranging from 5 minutes to 1 hour to fine-tune the models. With the rapid development of AIGC (generative artificial intelligence content), audio forgery technology has been iteratively updated. Recent technologies such as ElevenLabs, VALL-E, and NaturalSpeech3 can accurately replicate the voice of any target speaker using only a single audio clip. This rapid development makes it difficult to fully understand the true risks of audio forgery. Summary of the Invention

[0004] To address the above issues, the present invention provides a universal testing method for audio forgery algorithms. This universal testing framework can comprehensively assess the threats posed by a given audio forgery algorithm in the real world. The present invention is implemented through the following technical solutions:

[0005] The present invention discloses a universal testing method for audio forgery algorithms, comprising:

[0006] S1: Download speech audio containing multiple speakers from the open source corpus and construct a reference speaker speech dataset;

[0007] S2: Randomly select the original audio of the speaker in the open source corpus from the speaker speech dataset as the reference speech, and combine the attacker's forged text with the given audio forging algorithm to generate the original forged audio;

[0008] S3: Considering the real test scenario, different channel perturbations are added to the original forged audio obtained in S2 to obtain a set of forged audio that simulates the real scenario;

[0009] S4: Deploy multiple voiceprint recognition models, input the forged audio set obtained in S3 into the deployed voiceprint recognition model, and output a set of similarity scores;

[0010] S5: Deploy the speech listening evaluation model, input the forged audio set obtained in S3 into the deployed listening evaluation model, and output a set of audio intelligibility and naturalness scores;

[0011] S6: Deploy the forged audio detection model, input the forged audio set obtained in S3 into the deployed forged audio detection model, and output a probability score representing whether the corresponding audio is real speech.

[0012] As a further improvement, the open source corpus described in the present invention is VCTK or LibriSpeech or Mozilla Common Voice.

[0013] As a further improvement, in S3 of the present invention, adding different channel perturbations specifically includes the following:

[0014] Direct transmission in digital space without any modification;

[0015] Add random white noise to the audio;

[0016] Perform audio algorithm compression and re-decompression on the audio;

[0017] Simulate physical space playback for audio and adjust different background noises;

[0018] Add different spatial reverb effects to audio.

[0019] As a further improvement, in S4 of the present invention, the voiceprint recognition models deployed include open source voiceprint models, commercial voiceprint service models, and integrated voiceprint function models. The specific process of S4 is:

[0020] Record and register the speaker's original audio into the voiceprint recognition model;

[0021] The forged audio set obtained by S3 is input into the corresponding voiceprint recognition model;

[0022] The voiceprint recognition model outputs a set of scores that represent the similarity of speech.

[0023] As a further improvement, in S5 of the present invention, the quality distribution range of the output audio intelligibility and naturalness score set is 0-5.

[0024] As a further improvement, in S6 of the present invention, the forged audio detection model includes a passive forged audio detection model and an active audio watermark detection model.

[0025] As a further improvement, the specific steps of the passive forged audio detection model described in the present invention are:

[0026] Mix the fake audio set obtained in S3 and the real speech audio obtained in S1 to form a test set;

[0027] Input the test set into the passive forged audio detection model to obtain a probability score of whether the audio is real speech;

[0028] The probability score is compared with the default threshold for the passive forged audio detection model to obtain the detection accuracy.

[0029] As a further improvement, in S6 of the present invention, the specific steps for the active audio watermark detection model are:

[0030] Add a watermark to the forged audio set obtained in S3 to obtain a forged audio set with the watermark added;

[0031] Mix the watermarked forged audio set with the real speech obtained by S1 to form a test set;

[0032] Input the test set into the active audio watermark detection to obtain the probability score of whether the audio is real speech;

[0033] The probability score is compared with the default threshold of the given watermark detection model to obtain the detection accuracy.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] 1) This paper proposes a universal testing method for audio forgery algorithms. By adding channel perturbation data augmentation to forged audio in S3, it simulates the signal changes of forged audio in the real world and can evaluate the actual threat posed by a given audio forgery algorithm in the real world.

[0036] 2) The general testing method for audio forgery algorithms proposed in this invention can provide the deceptive threat of a given audio forgery algorithm to real-world voiceprint recognition models and voiceprint recognition services by testing the forged audio set based on the voiceprint recognition model in S4.

[0037] 3) The general testing method for audio forgery algorithms proposed in this invention can evaluate the hidden threat posed by a given audio forgery algorithm to human auditory perception by evaluating a set of forged audios based on the speech hearing test model in S5.

[0038] 4) The general testing method for audio forgery algorithms proposed in this invention can effectively evaluate the anti-detection ability of a given audio forgery algorithm against the detection model by detecting a forged audio set based on the forged audio detection model in S6, further revealing the comprehensive security threats posed by audio forgery algorithms in the real world. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is an algorithm flow chart of a general testing method for audio forgery algorithms provided in this embodiment. DETAILED DESCRIPTION

[0040] The present invention will be further described below with reference to specific examples. It should be understood that these examples are only intended to assist those of ordinary skill in the art in their understanding of the principles and knowledge of the present invention, and are not intended to limit the scope of the present invention and should not be considered to limit the application scenarios of the present invention. It should also be understood that after reading the contents taught by the present invention, those skilled in the art may make various changes or modifications to the present invention, but the deformation, changes and conversions made to the embodiments based on the principles and purpose of the present invention also fall within the scope defined by the claims appended hereto. It is also apparent that this specification is given as an example only with preferred embodiments, and it is not necessary to exhaust all embodiments in detail.

[0041] Figure 1 This is an algorithm flow chart of a general testing method for audio forgery algorithms provided in this embodiment.

[0042] The present invention discloses a universal testing method for audio forgery algorithms, comprising:

[0043] S1: Download speech audio containing multiple speakers from open source corpora and construct a reference speaker speech dataset from them; the open source corpus is VCTK, LibriSpeech, or Mozilla Common Voice.

[0044] S2: Randomly select the original audio of the speaker in the open source corpus from the speaker speech dataset as the reference speech, and combine the attacker's forged text with the given audio forging algorithm to generate the original forged audio;

[0045] S3: Considering a real-world test scenario, different channel perturbations are added to the original forged audio obtained in S2 to obtain a set of forged audio that simulates a real-world scenario. Adding different channel perturbations specifically includes:

[0046] Direct transmission in digital space without any modification;

[0047] Add random white noise to the audio;

[0048] Perform audio algorithm compression and re-decompression on the audio;

[0049] Simulate physical space playback for audio and adjust different background noises;

[0050] Add different spatial reverb effects to audio.

[0051] Specifically:

[0052] S3-1, adding random noise to the original audio in the digital domain, where the signal-to-noise ratio includes 5dB, 10dB, 30dB and 40dB;

[0053] S3-2, compressing and reconnecting the original audio in the digital domain, including compression and decompression processing using the FLAC format and the MP3 format;

[0054] S3-3, perform noise perturbation on the original audio in the physical domain, add a physical noise source with intensities of 10dB, 30dB, and 40dB;

[0055] S3-4. Replay the original audio in different physical spaces to collect the audio after spatial reverberation.

[0056] S4: Deploy multiple voiceprint recognition models, input the forged audio set obtained in S3 into the deployed voiceprint recognition model, and output a set of similarity scores;

[0057] The deployed voiceprint recognition models include open source voiceprint models, commercial voiceprint service models, and integrated voiceprint function models. The specific S4 process is:

[0058] Record and register the speaker's original audio into the voiceprint recognition model;

[0059] The forged audio set obtained by S3 is input into the corresponding voiceprint recognition model;

[0060] The voiceprint recognition model outputs a set of scores that represent the similarity of speech.

[0061] Specifically:

[0062] S4-1. Input the speaker's original audio into the voiceprint recognition model ECAPA-TDNN, ResNet, and Resemblyzer to generate a registered voiceprint embedding code.

[0063] S4-2. Input the forged audio set obtained in S3 into the above-mentioned voiceprint recognition model to generate a voiceprint embedding code corresponding to the forged audio;

[0064] S4-3. The voiceprint recognition model calculates the cosine similarity between the forged audio voiceprint embedding code and the registered voiceprint embedding code, and outputs a score set representing the similarity.

[0065] S5: Deploy the speech listening evaluation model, input the forged audio set obtained in S3 into the deployed listening evaluation model, and output a set of audio intelligibility and naturalness scores; the quality distribution range of the output audio intelligibility and naturalness score set is 0-5.

[0066] S5-1, input the forged audio set obtained in S3 into the listening evaluation model;

[0067] S5-2. The listening evaluation model outputs scores for audio intelligibility and naturalness.

[0068] S6: Deploy the forged audio detection model. Input the forged audio collection obtained in S3 into the deployed forged audio detection model, which outputs a probability score indicating whether the corresponding audio is authentic speech. The forged audio detection model includes a passive forged audio detection model and an active audio watermark detection model.

[0069] The specific steps for the passive forged audio detection model are:

[0070] Mix the fake audio set obtained in S3 and the real speech audio obtained in S1 to form a test set;

[0071] Input the test set into the passive forged audio detection model to obtain a probability score of whether the audio is real speech;

[0072] The probability score is compared with the default threshold for the passive forged audio detection model to obtain the detection accuracy.

[0073] The specific steps for the active audio watermark detection model are:

[0074] Add a watermark to the forged audio set obtained in S3 to obtain a forged audio set with the watermark added;

[0075] Mix the watermarked forged audio set with the real speech obtained by S1 to form a test set;

[0076] Input the test set into the active audio watermark detection to obtain the probability score of whether the audio is real speech;

[0077] The probability score is compared with the default threshold of the given watermark detection model to obtain the detection accuracy.

[0078] The above description is only a preferred embodiment of the present invention. Although the present invention has disclosed the preferred embodiment as above, it is not intended to limit the present invention. Any person skilled in the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solution of the present invention without departing from the scope of the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.

Claims

1. A general testing method for audio forgery algorithms, characterized by: include: S1: Download speech audio from open source corpora containing multiple speakers and construct a reference speaker real speech dataset; S2: Randomly select the original audio of the speaker in the open source corpus from the speaker speech dataset as the reference speech, and combine the attacker's forged text with the given audio forging algorithm to generate the original forged audio; S3: Considering the real test scenario, different channel perturbations are added to the original forged audio obtained in S2 to obtain a set of forged audio that simulates the real scenario; S4: Deploy multiple voiceprint recognition models, input the forged audio set obtained in S3 into the deployed voiceprint recognition model, and output a set of similarity scores; S5: Deploy the speech listening evaluation model, input the forged audio set obtained in S3 into the deployed listening evaluation model, and output a set of audio intelligibility and naturalness scores; S6: Deploy the forged audio detection model, input the forged audio set obtained in S3 into the deployed forged audio detection model, and output a probability score representing whether the corresponding audio is real speech; In S3, adding different channel perturbations specifically includes the following: Direct transmission in digital space without any modification; Add random white noise to the audio; Perform audio algorithm compression and re-decompression on the audio; Simulate physical space playback for audio and adjust different background noises; Add different spatial reverberation effects to audio; In the above S4, the voiceprint recognition models deployed include open source voiceprint models, commercial voiceprint service models, and integrated voiceprint function models. The specific process of the above S4 is: Record and register the speaker's original audio into the voiceprint recognition model; The forged audio set obtained by S3 is input into the corresponding voiceprint recognition model; The voiceprint recognition model outputs a set of scores that represent speech similarity; In the above-mentioned S5, the quality distribution range of the output audio intelligibility and naturalness score set is 0-5.

2. The universal testing method for audio forgery algorithms according to claim 1, characterized in that: The open source corpus is VCTK or LibriSpeech or Mozilla Common Voice.

3. The universal testing method for audio forgery algorithms according to claim 1, characterized in that: In the aforementioned S6, the forged audio detection model includes a passive forged audio detection model and an active audio watermark detection model.

4. The universal testing method for audio forgery algorithms according to claim 3, characterized in that: The specific steps of the passive forged audio detection model are as follows: Mix the fake audio set obtained in S3 and the real speech dataset of the reference speaker obtained in S1 to form a test set; Input the test set into the passive forged audio detection model to obtain a probability score of whether the audio is real speech; The probability score is compared with the default threshold for the passive forged audio detection model to obtain the detection accuracy.

5. The universal testing method for audio forgery algorithms according to claim 4, characterized in that: In S6, the specific steps for the active audio watermark detection model are: Add a watermark to the forged audio set obtained in S3 to obtain a forged audio set with the watermark added; Mix the watermarked forged audio set with the real speech dataset of the reference speaker obtained in S1 to form a test set; Input the test set into the active audio watermark detection to obtain the probability score of whether the audio is real speech; The probability score is compared with the default threshold of the given watermark detection model to obtain the detection accuracy.