Anti-forgery voiceprint authentication method, device, equipment, medium and program product

CN122799862APending Publication Date: 2026-09-22CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330322.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0002]随着音色转换和音色克隆技术的迅速发展,借助于伪造声音以欺骗声纹认证系统的操作也大量出现,从而对人们的信息和财产安全产生了威胁

Benefits of technology

[0017]本申请实施例还提供了一种计算机程序产品,所述计算机程序产品包括计算机程序;所述计算机程序被电子设备的处理器执行时,能够实现如前任一所述的抗伪造声纹认证方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799862A_ABST
    Figure CN122799862A_ABST
Patent Text Reader

Abstract

The application discloses an anti-forgery voiceprint authentication method and device, equipment, medium and program product, comprising: obtaining to-be-authenticated audio; extracting an acoustic feature set from the to-be-authenticated audio; performing a forged audio discrimination process on the acoustic feature set to obtain a first discrimination result; performing a speaker discrimination process on a voiceprint feature set corresponding to the acoustic feature set to obtain a second discrimination result; and determining an anti-forgery voiceprint authentication result for the to-be-authenticated audio based on the first discrimination result and the second discrimination result. Through the scheme, all-round accurate authenticity authentication of to-be-authenticated audio can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio biometric recognition technology, and in particular to an anti-counterfeiting voiceprint authentication method, device, equipment, medium and program product. Background Technology

[0002] With the rapid development of voice conversion and cloning technologies, operations that use forged voices to deceive voiceprint authentication systems have proliferated, posing a threat to people's information and property security. Therefore, how to comprehensively verify the authenticity of users' audio data has become an urgent technical problem to be solved. Summary of the Invention

[0003] Based on the above technical problems, this application provides an anti-counterfeiting voiceprint authentication method, apparatus, device, medium, and program product.

[0004] The technical solution provided in this application is as follows:

[0005] This application first provides an anti-counterfeiting voiceprint authentication method, the method comprising:

[0006] Obtain the audio to be authenticated;

[0007] Extract a set of acoustic features from the audio to be authenticated;

[0008] Perform forged audio identification processing on the acoustic feature set to obtain a first identification result;

[0009] Speaker identification processing is performed on the voiceprint feature set corresponding to the acoustic feature set to obtain a second identification result;

[0010] Based on the first authentication result and the second authentication result, the result of anti-spoofing voiceprint authentication for the audio to be authenticated is determined.

[0011] This application embodiment also provides an anti-counterfeiting voiceprint authentication device, the anti-counterfeiting voiceprint authentication device comprising:

[0012] The acquisition module is used to acquire the audio to be authenticated;

[0013] The authentication module is used to extract an acoustic feature set from the audio to be authenticated; perform forged audio authentication processing on the acoustic feature set to obtain a first authentication result; and perform speaker authentication processing on the voiceprint feature set corresponding to the acoustic feature set to obtain a second authentication result.

[0014] The determining module is used to determine the result of anti-spoofing voiceprint authentication for the audio to be authenticated based on the first authentication result and the second authentication result.

[0015] This application also provides an electronic device, which includes a processor and a memory; wherein the memory stores a computer program; when the computer program is executed by the processor, it can implement the anti-counterfeiting voiceprint authentication method as described above.

[0016] This application also provides a computer-readable storage medium storing a computer program; when the computer program is executed by a processor of an electronic device, it can implement the anti-counterfeiting voiceprint authentication method as described above.

[0017] This application also provides a computer program product, which includes a computer program; when the computer program is executed by the processor of an electronic device, it can implement the anti-counterfeiting voiceprint authentication method as described above.

[0018] The anti-spoofing voiceprint authentication method provided in this application, after acquiring the audio to be authenticated and extracting the acoustic feature set, performs anti-spoofing audio identification processing on the acoustic feature set to obtain a first identification result, thus realizing anti-spoofing audio identification processing of the audio to be authenticated; and performs speaker identification processing on the voiceprint feature set corresponding to the acoustic feature set to obtain a second identification result, thus realizing speaker identification processing of the speech to be authenticated; based on this, the result of anti-spoofing voiceprint authentication for the audio to be authenticated is determined based on the first and second identification results. In this way, not only is anti-spoofing audio identification processing and speaker identification processing of the audio to be authenticated realized simultaneously, but also the organic integration of the first and second identification results is achieved, thereby improving the comprehensiveness and completeness of the data in the result of anti-spoofing voiceprint authentication, and thus achieving the accuracy of anti-spoofing detection and speaker identification processing of the audio to be authenticated, which can meet the comprehensive identification needs of authenticity and similarity identification of the audio to be authenticated. Attached Figure Description

[0019] Figure 1 A flowchart illustrating the anti-counterfeiting voiceprint authentication method provided in this application embodiment;

[0020] Figure 2 A statistical diagram illustrating the first and second identification results provided for embodiments of this application;

[0021] Figure 3 This is another schematic diagram of the anti-counterfeiting voiceprint authentication method provided in the embodiments of this application;

[0022] Figure 4 This is a schematic diagram of the structure for voiceprint feature extraction provided in an embodiment of this application;

[0023] Figure 5This is a schematic diagram of the structure of the audio authentication system provided in the embodiments of this application;

[0024] Figure 6 This is a schematic diagram of the anti-counterfeiting voiceprint authentication device provided in the embodiments of this application;

[0025] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0027] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0028] With the rapid development of artificial intelligence and voiceprint recognition technologies, voiceprint authentication has become an efficient and convenient method of identity verification. However, with the widespread application of these authentication methods, voice conversion and voice cloning technologies have also developed rapidly. In this context, the voices of target speakers forged using voice conversion and voice cloning technologies can frequently deceive voiceprint authentication systems. Therefore, how to identify the similarity between audio data and the voice data of the target object, while simultaneously determining the authenticity of the audio data, has become an urgent technical problem to be solved.

[0029] Based on the above technical problems, this application provides an anti-forgery voiceprint authentication method. Figure 1 This is a flowchart illustrating the anti-counterfeiting voiceprint authentication method provided in the embodiments of this application, as shown below. Figure 1 As shown, the process may include the following steps:

[0030] Step 101: Obtain the audio to be authenticated.

[0031] In one implementation, the audio to be authenticated may include real-time audio data acquired in real time by an audio acquisition device, and may also include historical audio data acquired from the storage space of a terminal device or a network-side device.

[0032] Step 102: Extract the acoustic feature set from the audio to be authenticated.

[0033] In the field of speech signal processing, acoustic features are features extracted from speech signals that reflect frequency domain characteristics. The acoustic features of audio signals can be used in applications such as speech recognition, speech synthesis, and voiceprint recognition.

[0034] In one implementation, the acoustic feature set may include features such as the frequency, amplitude, and phase of the audio to be authenticated.

[0035] In one implementation, the acoustic feature set can be obtained in the following way:

[0036] The acoustic feature set is obtained by extracting features from the audio to be authenticated using a pre-trained neural network model.

[0037] Step 103: Perform fake audio identification processing on the acoustic feature set to obtain the first identification result.

[0038] In one implementation, the first identification result may indicate the probability that the audio to be authenticated is real audio or fake audio, or the first identification result may indicate that the audio to be authenticated is real audio or fake audio; for example, fake audio can be obtained by processing the original audio using sound cloning and / or timbre conversion technology.

[0039] Phono conversion refers to transforming the acoustic characteristics of an audio signal, such as timbre, for example, changing the timbre of speaker A to that of speaker B. In practical applications, signal processing combined with deep learning-based neural network technology is typically used to achieve timbre conversion. Phono cloning, on the other hand, is a technique that generates a voice that mimics the voice of speaker C by simulating or reconstructing the vocal features of speaker C. This process involves signal processing, machine learning, and deep learning technologies.

[0040] In one implementation, the forged audio detection process can be achieved in any of the following ways:

[0041] The features in the acoustic feature set are filtered by a data model that can capture the audio features of real speech to obtain the filtering results, and then the first identification result is determined based on the filtering results. For example, if the number of audio features of real speech in the acoustic feature set represented by the filtering results is less than a first threshold, the first identification result can be determined as: fake audio.

[0042] The acoustic feature set is processed by anti-spoofing-aware speaker verification (SASV) to obtain the first identification result. The goal of SASV is to defend the speaker verification system against audio attacks forged by synthetic technology. It captures the artifacts of synthetic audio in the acoustic feature set through the counter-measure (CM) module, thereby realizing the identification of the acoustic feature set and filtering out spoofing attacks corresponding to forged audio, thus improving the security and robustness of the speaker verification system.

[0043] Specifically, SASV can not only track the fluency of the audio to be authenticated, but also the generation process and phase characteristics of the audio to be authenticated; and by identifying and analyzing whether the audio to be authenticated contains specific patterns or traces corresponding to the forged audio, it can effectively detect and resist timbre switching and timbre cloning attacks.

[0044] Correspondingly, anti-spoofing refers to methods and techniques for preventing various deception attacks. In voiceprint recognition, anti-spoofing refers to measures to prevent the use of synthesized speech, recorded voices, or other forgery methods to deceive the voiceprint authentication system. These techniques can improve the security and reliability of the voiceprint authentication system.

[0045] Step 104: Perform speaker identification processing on the voiceprint feature set corresponding to the acoustic feature set to obtain the second identification result.

[0046] In one implementation, the voiceprint in the voiceprint feature set may include the unique sound features generated when an individual speaks; specifically, the voiceprint may change with the unique features of an individual's vocal cord structure, throat shape, and pronunciation habits, and voiceprint recognition technology uses these voiceprint features to identify the speaker's identity.

[0047] In one implementation, the second authentication result can indicate whether the audio to be authenticated is the audio of the target object; for example, the target object can be a pre-defined person or animal.

[0048] In one implementation, the second identification result can be obtained in the following way:

[0049] The acoustic feature set is processed to obtain a voiceprint feature set, and the target voiceprint feature is obtained. The degree of matching between the voiceprint feature in the voiceprint feature set and the target voiceprint feature is determined. If the degree of matching is greater than or equal to a second threshold, the second identification result can be: the audio to be authenticated is the audio of the target object. If the degree of matching is less than the second threshold, the second identification result can be: the audio to be authenticated is not the audio of the target object. The target voiceprint feature can include the voiceprint feature corresponding to the real speech of the target object that is preset. For example, the degree of matching can be reflected in the form of a matching score.

[0050] Step 105: Based on the first authentication result and the second authentication result, determine the result of the anti-spoofing voiceprint authentication for the audio to be authenticated.

[0051] In one implementation, the result of anti-spoofing voiceprint authentication may include: the voice to be authenticated is a spoofed voice, the voice to be authenticated is a real voice and belongs to the target object, and the voice to be authenticated is a real voice but does not belong to the target object.

[0052] In one implementation, the result of anti-spoofing voiceprint authentication can be obtained in the following way:

[0053] The first identification result is quantified to obtain the first data, and the second identification result is quantified to obtain the second data. Then, based on the joint statistical state of the first data and the second data, the result of anti-counterfeiting voiceprint authentication is determined.

[0054] Figure 2 This is a statistical diagram illustrating the first and second identification results provided in an embodiment of this application; Figure 2 In the two-dimensional coordinate system shown, the horizontal axis represents the voiceprint similarity score, which corresponds to the first identification result or the first data, and the vertical axis represents the authenticity detection score, which corresponds to the second identification result or the second data. The voiceprint similarity score increases along the direction of the horizontal axis arrow, and the larger the value, the higher the voiceprint similarity score. The higher the voiceprint similarity score, the greater the probability that the audio to be authenticated is the audio of the target object. The authenticity detection score decreases along the direction of the vertical axis arrow, and the smaller the value, the higher the degree of forgery. The smaller the value, the greater the probability that the audio to be authenticated is forged.

[0055] Taking the single dimension of voiceprint similarity score as an example, the first voiceprint similarity score 201 can correspond to the first audio. Since the first voiceprint similarity score 201 is concentrated and its overall value is relatively large, the first audio can be determined to be the audio of the target object based on the first voiceprint similarity score. The second voiceprint similarity score 202 can correspond to the second audio. Since the second voiceprint similarity score is concentrated and its overall value is relatively small, the second audio can be determined to be the audio of the target object based on the second voiceprint similarity score. As for the third voiceprint similarity score 203, its value is basically evenly distributed within a wide range of the horizontal axis. Therefore, it is impossible to determine whether the third audio corresponding to it is the audio of the target object.

[0056] Taking the single dimension of authenticity detection score as an example, for the first authenticity score 204, since its distribution is concentrated and its overall value is relatively large, the fourth audio corresponding to the first authenticity score can be determined to be real audio; the second authenticity score 205 has a similar distribution and value to the first authenticity score 204, so the fifth audio corresponding to it can be determined to be real speech; the third authenticity score 206 has a relatively concentrated distribution and its overall value is relatively small, so the sixth audio corresponding to it can be determined to be fake speech.

[0057] However, the accuracy of the results obtained by identifying audio data from any single dimension is insufficient. Therefore, in this embodiment, the first and second data are jointly statistically analyzed from two dimensions: the authenticity detection score and the voiceprint similarity score, to obtain the anti-spoofing voiceprint authentication result. For example, regarding the first result 207, although its voiceprint similarity score is evenly distributed across the entire horizontal axis, its authenticity detection score is relatively small. Therefore, the speech corresponding to this result can be determined to be spoofed speech. Regarding the second result 208, its voiceprint similarity score is distributed within a smaller horizontal axis range, but its authenticity detection score is relatively large. Therefore, the speech corresponding to this result is real speech but not the speech of the target object. Regarding the third result 209, its voiceprint similarity score is relatively concentrated and has a large value, and its authenticity detection score is also large and relatively concentrated. Therefore, the speech corresponding to this result is real speech and is the speech of the target object. It should be noted that the anti-spoofing voiceprint authentication result can include the first to the third results.

[0058] One approach involves first using a separately developed synthesis detection module to distinguish between real and synthesized audio. Then, a SASV-based synthesis detection module is cascaded with a pre-trained voiceprint authentication system. This first rejects synthesized speech, then rejects real speech from non-target speakers, thus completing voiceprint authentication. The problem with this approach is that developing a separate anti-spoofing module requires dedicated feature extractors and classifiers, which typically increases computational demands and resource utilization, making it unsuitable for portable scenarios. Furthermore, since anti-spoofing authentication and speaker verification are separate tasks, additional computation is required during the verification phase.

[0059] Figure 3 This is another schematic diagram of the anti-counterfeiting voiceprint authentication method provided in an embodiment of this application. For example... Figure 3 As shown, this method can be implemented through coordination between the client and the server. The client collects the user's audio data and sends it to the server for identification and authentication. The server can be equipped with a robust voiceprint authentication model that can detect deepfake voices. This model analyzes and verifies the received audio data to identify and resist timbre switching and timbre cloning attacks, thereby achieving audio identification with anti-spoofing capabilities.

[0060] Specifically, the method may include two phases: a registration phase 301 and an authentication and verification phase 302. Wherein:

[0061] For the registration phase 301, users register through the client and submit real user recordings to the server. The server then performs acoustic feature extraction and voiceprint extraction in sequence to create the user's voiceprint template, which is stored in the database for subsequent comparison.

[0062] For the authentication and anti-spoofing stage 302, the client sends the voice to be authenticated to the server. The server performs acoustic feature extraction and voiceprint extraction on the voice to be authenticated. Then, the server system compares the voiceprint features extracted from the voice to be authenticated with the voiceprint template, and scores the voiceprint similarity based on the comparison results, thereby obtaining the voiceprint authentication score corresponding to the first authentication result. At the same time, the server can also analyze the deception traces, including synthetic features and abnormal patterns, in the voice to be authenticated through a voice anti-spoofing algorithm to obtain a second authentication result. The second authentication result and the voiceprint authentication score are then fused to obtain the anti-spoofing voiceprint authentication result for the audio to be authenticated. Through this result, it is possible not only to determine whether the audio to be authenticated is the audio of the target object, but also to determine whether the audio to be authenticated is genuine audio.

[0063] The above process and structure enable simultaneous judgment of the authenticity and similarity of the audio to be authenticated, thereby simplifying the complexity of identifying and authenticating the audio and improving the flexibility of the identification and authentication process.

[0064] As can be seen from the above, the anti-spoofing voiceprint authentication method provided in this application, after obtaining the audio to be authenticated and extracting the acoustic feature set, performs anti-spoofing audio identification processing on the acoustic feature set to obtain a first identification result, thus realizing anti-spoofing audio identification processing of the audio to be authenticated; and performs speaker identification processing on the voiceprint feature set corresponding to the acoustic feature set to obtain a second identification result, thus realizing speaker identification processing of the speech to be authenticated; based on this, the result of anti-spoofing voiceprint authentication for the audio to be authenticated is determined based on the first identification result and the second identification result. In this way, not only is anti-spoofing audio identification processing and speaker identification processing of the audio to be authenticated realized simultaneously, but also the organic integration of the first identification result and the second identification result is realized, thereby improving the comprehensiveness and completeness of the data in the result of anti-spoofing voiceprint authentication, and thus realizing the accuracy of anti-spoofing detection and speaker identification processing of the audio to be authenticated, and meeting the comprehensive identification needs of authenticity identification and similarity identification of the audio to be authenticated.

[0065] In the technical solution provided in this application embodiment, speaker identification can be performed through the voiceprint feature set, thereby enabling the solution to have the speaker's identity authentication function. By performing forged audio identification, the introduction of an audio anti-spoofing mechanism is realized, thereby enabling the solution to more effectively identify synthetic voices or other deception attacks. In other words, the technical solution provided in this application can reject both non-target sample speech and speech of the target object generated by a deepfake system.

[0066] Based on the foregoing embodiments, the anti-spoofing voiceprint authentication method provided in this application performs spoofing audio identification processing on the acoustic feature set to obtain a first identification result, which can be achieved in the following ways:

[0067] An out-of-distribution (OOD) detection is performed on the features in the acoustic feature set to obtain the first discrimination result.

[0068] In one implementation, OOD detection can be achieved in the following way:

[0069] The detection model capable of performing OOD detection identifies features in the acoustic feature set based on the target feature set, and obtains a first identification result; for example, the target feature set may include the set of acoustic features of real speech captured by the detection model during training based on real samples; for example, the detection model may include a Lupin voiceprint authentication model.

[0070] Since the acoustic feature set carried by fake audio can differ from the acoustic feature set of real speech (i.e., the target feature set), the fake audio identification process can be modeled as OOD detection by leveraging these features.

[0071] In practical applications, speech synthesis and forgery technologies are constantly evolving, and attack methods using forged speech are becoming increasingly diverse and unpredictable, posing new challenges to the audio authentication process. Existing classification-based audio forgery detection methods can only handle known attacks, and their ability to detect new and unknown synthesized speech is insufficient.

[0072] The anti-spoofing voiceprint authentication method provided in this application obtains a first identification result by performing OOD detection on the features in the acoustic feature set. When the acoustic feature set corresponding to real speech remains stable, it can still achieve accurate detection of the speech to be identified generated by constantly updated audio synthesis and spoofing technologies through OOD detection, thereby improving the accuracy of the first identification result and improving the stability of spoofing audio identification processing.

[0073] Furthermore, by leveraging the concept of OOD detection, in practical applications, the process of identifying and processing fake audio can be continuously adjusted and optimized based on the collected real audio data and the acoustic feature set of fake speech. This allows for timely updates to the recognition strategy, improving the ability to respond to new types of attacks. It can also handle the identification and processing of fake speech generated by constantly changing attack methods and synthesis techniques, thus giving the process of identifying and processing fake audio dynamically adaptable and learning. Moreover, this flexible learning mechanism enables the process of identifying and processing fake audio to adapt to the detection needs of complex and ever-changing audio security threats.

[0074] Meanwhile, through OOD detection, the technical solution provided in this application can not only filter out forged speech attacks generated by known forged audio methods, but also filter and identify forged speech generated by unknown synthesis algorithms or other unknown attack methods, thus making the solution generalizable. Therefore, this solution is of great significance in improving the ability of voiceprint authentication systems to identify deception and is suitable for application scenarios with high security requirements.

[0075] Based on the foregoing embodiments, the anti-counterfeiting voiceprint authentication method provided in this application involves performing OOD detection on the features in the acoustic feature set to obtain a first identification result, which can be achieved through the following steps:

[0076] Step A1: Determine the set of quantization results corresponding to the set of acoustic features.

[0077] In one implementation, the quantization results in the quantization result set may include unnormalized class scores of acoustic features in the acoustic feature set; for example, the quantization results may be logits of acoustic features in the acoustic feature set.

[0078] In one implementation, the set of quantization results can be determined in the following way:

[0079] The acoustic features in the acoustic feature set are processed by a fully connected layer or a global flat pooling layer to obtain a set of quantization results.

[0080] Step A2: Determine the probability distribution state corresponding to the quantization result set.

[0081] In one implementation, the probability distribution state may include the k-th quantization result in the quantization result set, and the probability of its occurrence or existence in the quantization result set; wherein k may be an integer greater than or equal to 1.

[0082] In one implementation, the probability distribution state can be determined in the following way:

[0083] The Softmax function is used to perform parameterized class distribution processing on the data in the quantization result set, thereby obtaining the probability distribution state.

[0084] Step A3: Obtain the fitted audio data based on the probability distribution state fitting.

[0085] In one implementation, the fitted audio data may include audio data generated after encoding and decoding based on a probability distribution state.

[0086] In one implementation, the fitted audio data can be obtained in the following way:

[0087] The probability distribution is fitted using maximum likelihood estimation to obtain the fitted audio data.

[0088] Step A4: Perform OOD detection on the fitted audio data to obtain the first identification result.

[0089] In one implementation, the first identification result can be obtained in the following way:

[0090] The fitted audio data and the audio to be certified are compared to obtain the comparison result, and OOD detection is performed on the comparison result. For example, if the comparison result indicates that the similarity between the fitted audio data and the audio to be certified is greater than the third threshold, the first identification result can be determined as: the audio to be certified is real speech; if the comparison result indicates that the similarity is less than or equal to the third threshold, the first identification result can be determined as: the audio to be certified is fake speech.

[0091] It should be noted that forged speech carries unnatural and non-fluent acoustic features. Therefore, the similarity between the fitted speech data obtained by fitting in the above way and the forged speech will inevitably be less than the third threshold.

[0092] As can be seen from the above, the anti-counterfeiting voiceprint authentication method provided in this application determines the probability distribution state of the quantization result set after determining the quantization result set corresponding to the acoustic feature set. In this way, the distribution state of the features in the acoustic feature set can be accurately quantified through the probability distribution state. Furthermore, the fitted audio data is obtained based on the probability distribution state, and OOD detection is performed on the fitted audio data to obtain the first identification result. Through the above method, with the help of OOD detection, the fitted audio data can be accurately screened, thereby improving the accuracy of the first identification result.

[0093] Based on the foregoing embodiments, the anti-forgery voiceprint authentication method provided in this application includes a joint energy-based model (JEM) comprising a generative model and a classification model that are jointly generated.

[0094] Accordingly, performing OOD detection on the features in the acoustic feature set to obtain the first discrimination result can be achieved in the following way:

[0095] The first identification result is obtained by performing OOD detection on the features in the acoustic feature set using a generative model.

[0096] JEM is a deep learning framework that combines generative and discriminative models. Based on the principles of Energy-Based Models (EBM), it aims to perform both generation and discrimination tasks within a unified framework. JEM addresses the discrimination and generation problems by defining the joint probability of data and labels. It optimizes the generation task by minimizing the energy function corresponding to the data points. The definition of the energy function applies not only to loss optimization in discriminative models but also to providing generative capabilities for generative models; the discriminative model is also known as a classification model.

[0097] In mathematics and physics, the energy function is a function that describes the state of a system and is commonly used to evaluate the energy level of a system under different states. In machine learning and optimization, the energy function is also frequently used to measure the performance or loss of a model, thus serving as the objective function or loss function of an optimization algorithm.

[0098] In one implementation, the first identification result can be obtained in the following way:

[0099] The energy function is represented by the logits corresponding to the quantization result set. The joint density between the density function corresponding to the probability distribution state, the data to be identified, and the category identifier is defined. The energy data of the audio to be authenticated is determined by the generative model associated with the energy function. Then, OOD detection is performed on the energy data to obtain the first identification result. For example, if the energy data is greater than or equal to the fourth threshold, the first identification result can be determined as: the audio to be authenticated is fake audio. If the energy data is less than the fourth threshold, the first identification result can be determined as: the audio to be authenticated is real audio.

[0100] As can be seen from the above, in the anti-spoofing voiceprint authentication method provided in this application embodiment, the features in the acoustic feature set are subjected to OOD detection by the generative model in JEM to obtain the first identification result. This method can perform high-precision OOD detection on the audio to be authenticated in various scenarios, thereby improving the accuracy of the first identification result.

[0101] Based on the foregoing embodiments, in the anti-spoofing voiceprint authentication method provided in this application, before performing speaker identification processing on the voiceprint feature set corresponding to the acoustic feature set, the following steps may also be performed:

[0102] Step B1: Perform weighted fusion processing on the acoustic features in the acoustic feature set to obtain the weighted fusion result.

[0103] In one implementation, the weighted fusion result can be obtained in the following way:

[0104] Weight parameters are predetermined, and weighted fusion processing is performed on the acoustic features in the acoustic feature set based on the weight values ​​in the weight parameters to obtain the weighted fusion result.

[0105] Step B2: The weighted fusion result is processed by a multi-scale feature aggregation (MFA) model to extract features and obtain a set of voiceprint features.

[0106] Figure 4 This is a schematic diagram of the structure for voiceprint feature extraction provided in an embodiment of this application, as shown below. Figure 4 As shown, the audio to be authenticated is input into the Waveform Language Model (WavLM) for feature extraction, resulting in an acoustic feature set. Then, a weighted fusion process is performed on the acoustic feature set, and the weighted fusion result is sent to the MFA model for feature extraction, thereby aggregating acoustic features at different scales.

[0107] As can be seen from the above, the anti-counterfeiting voiceprint authentication method provided in this application obtains a weighted fusion result by performing weighted fusion processing on the acoustic features in the acoustic feature set. In this way, while highlighting the key features in the acoustic feature set, it can also balance the influence of the features in the acoustic feature set. Furthermore, by performing feature extraction processing on the weighted fusion result through the MFA model to obtain the voiceprint feature set, it is possible to accurately capture the relationship between the features in the weighted fusion result, thereby improving the accuracy of the voiceprint feature set.

[0108] Based on the foregoing embodiments, the anti-forgery voiceprint authentication method provided in this application includes a JEM comprising a generative model and a classification model that are jointly generated.

[0109] Accordingly, speaker identification processing is performed on the voiceprint feature set corresponding to the acoustic feature set to obtain the second identification result, which can be achieved in the following way:

[0110] By using a classification model, speaker identification processing is performed on the voiceprint feature set to obtain a second identification result.

[0111] In one implementation, the classification model can be trained by optimizing the classification loss Additive Angular Margin (AAM)-Softmax, and it can perform speaker classification; for example, the classification model can perform feature extraction of voiceprint features.

[0112] In one implementation, the second identification result can be obtained in the following way:

[0113] The similarity between features in the target feature set and features in the voiceprint feature set is calculated by a classification model, thereby obtaining the voiceprint similarity score in the aforementioned embodiment, and the second identification result is determined based on the voiceprint similarity score.

[0114] It should be noted that, in this embodiment of the application, the forged audio identification processing can be implemented by the generative model in JEM, and the speaker identification processing can be implemented by the classification model in JEM. That is to say, in this embodiment of the application, JEM can be used to perform forged audio identification processing and speaker identification processing on the audio to be authenticated respectively.

[0115] As can be seen from the above, the anti-spoofing voiceprint authentication method provided in this application performs speaker identification processing on the voiceprint feature set through the classification model in JEM to obtain a second identification result. Thus, through the above processing, with the help of the classification model of the JEM model, the voiceprint features in the voiceprint feature set can be captured more comprehensively, thereby improving the accuracy of the second identification result.

[0116] Based on the foregoing embodiments, the anti-spoofing voiceprint authentication method provided in this application involves performing speaker identification processing on the voiceprint feature set through a classification model to obtain a second identification result, which can be achieved through the following steps:

[0117] Step C1: Perform multi-classification processing on the features in the voiceprint feature set using a classification model to obtain the speaker category probability set of the audio to be authenticated.

[0118] In one implementation, the speaker category probability set may include a set of probabilities that the audio to be authenticated belongs to each of a plurality of speakers; for example, the probabilities in the speaker category probability set may be represented by the conditional probabilities between the audio to be authenticated and the speaker category identifier.

[0119] In one implementation, multi-classification processing can be achieved in the following way:

[0120] The logits corresponding to the voiceprint feature set are processed by the Softmax function included in the classification model to obtain the speaker category probability set.

[0121] Step C2: Determine the second identification result based on the speaker category probability set.

[0122] In one implementation, the speaker identifier corresponding to the highest probability in the speaker category probability set can be used to determine the second identification result. For example, if the speaker identifier is the identifier of the target user, the second identification result can be: the audio to be authenticated is the audio of the target user.

[0123] As can be seen from the above, the anti-spoofing voiceprint authentication method provided in this application performs multi-classification processing on the features in the voiceprint feature set through a classification model to obtain a speaker category probability set for the audio to be authenticated, and determines the second authentication result based on the speaker category probability set. Thus, by utilizing the above multi-classification processing, stable and comprehensive analysis of the features in the voiceprint feature set is achieved, thereby improving the stability of feature classification in the voiceprint feature set and consequently improving the accuracy of the second authentication result.

[0124] Based on the foregoing embodiments, in the anti-spoofing voiceprint authentication method provided in this application, forged audio processing is performed on the acoustic feature set to obtain a first identification result, which is achieved through the generative model in JEM; speaker identification processing is performed on the voiceprint feature set corresponding to the acoustic feature set to obtain a second identification result, which is achieved through the classification model in JEM.

[0125] Accordingly, the above method may also include the following steps:

[0126] Step D1: Obtain sample data.

[0127] In one implementation, the sample data may include real speech and fake speech; for example, the sample data may also include a authenticity identifier to distinguish whether the sample data is real speech.

[0128] In one implementation, the sample data may include the voices of multiple speakers; for example, the sample data may also include speaker identifiers to distinguish the speakers to which the sample data corresponds.

[0129] In one implementation, the sample identifier of an audio data set in the sample data may include two identifiers: an authenticity identifier and a speaker identifier.

[0130] Step D2: Jointly train the initial generative model and the initial classification model based on the sample data to obtain the generative model and the classification model.

[0131] Specifically, JEM addresses the classification problem for speaker identification and the generation problem for forged speech identification by defining the joint probability between sample data and sample identifiers. It optimizes the generation task of the generative model by minimizing the energy function corresponding to the sample data.

[0132] The definition of the energy function is applicable to loss optimization in classification models and can also provide generative capabilities for generative models.

[0133] In one implementation, the joint training of the initial classification model and the initial generation model can be achieved in the following way:

[0134] First, the optimization objective of JEM is determined. Second, the initial classification model is used to classify the voiceprint features corresponding to the sample data to obtain the first classification result. Then, the initial generation model is used to process the acoustic features corresponding to the sample data to generate the first generated audio. Next, the first classification result and the speaker identifier of the sample data are processed by the classification loss function included in the optimization objective to determine the first classification loss. The similarity between the first generated audio and the sample data is evaluated by the generation loss function included in the optimization objective to determine the first generation loss. Then, the parameters of the initial classification model and the initial generation model are adjusted by combining the first classification loss and the first generation loss to obtain the first adjusted classification model and the first adjusted generation model. The above method is then recursively executed to obtain the nth classification result and the nth generated audio, and to obtain the corresponding nth generation loss and the nth classification loss, to obtain the nth adjusted classification model and the nth adjusted generation model, until the generation loss and the classification loss meet the training requirements. Here, n is an integer greater than or equal to 2.

[0135] As can be seen from the above, the anti-spoofing voiceprint authentication method provided in this application, through joint training of the initial generation model and the initial classification model, enables the two models to complement each other, thereby allowing the two models to more comprehensively capture the real voiceprint features and deception traces in the sample data, improving the accuracy and robustness of JEM for audio identification and authentication; furthermore, through joint training, the initial generation model and the initial classification model can be jointly optimized, thus enabling the JEM obtained through the above training process to more accurately identify and authenticate spoofed speech, thereby improving the credibility of JEM.

[0136] Based on the foregoing embodiments, the anti-spoofing voiceprint authentication method provided in this application involves jointly training an initial generation model and an initial classification model based on sample data to obtain a generation model and a classification model, which can be achieved through the following steps:

[0137] Step E1: Perform feature extraction processing on the sample data to obtain the acoustic features of the sample.

[0138] Step E2: Use the initial classification model to perform speaker identification processing on the speaker features corresponding to the acoustic features of the samples, and obtain the sample classification results.

[0139] For example, the sample classification result can be obtained by obtaining the first identification result in the foregoing embodiments.

[0140] Step E3: Reconstruct the acoustic features of the sample using the initial generation model to obtain the sample generation result.

[0141] In one implementation, the reconstruction process can be the fitting operation described in the foregoing embodiments.

[0142] It should be noted that, in the embodiments of this application, the collection of modules including the generation model, the classification model, and the function of extracting acoustic features and voiceprint features can be integrated into an audio authentication system. Figure 5 This is a schematic diagram of the structure of the audio authentication system provided in an embodiment of this application. Figure 5 As shown, the sample data is input into a convolutional neural network (CNN) to extract the acoustic features of the samples and calculate the logits set corresponding to the features in the acoustic features of the samples. Then, the logits set is used to perform speaker classification processing and spoofing processing. Speaker classification processing can be implemented by a classification model, while spoofing processing can be implemented by a generative model.

[0143] Step E4: Based on the speaker identifiers and sample classification results of the sample data, adjust the parameters of the initial classification model to obtain the classification model. At the same time, based on the authenticity identifiers of the sample data and the sample generation results, adjust the parameters of the initial generation model to obtain the generation model.

[0144] In one implementation, the sample generation result can be achieved using the sample probabilities of the sample data. For example, the sample probability corresponding to sample point x in the sample data can be calculated using equation (1):

[0145]

[0146] Where Z(θ) and exp(-E) θ (x) represents the partition term and the exponent term, respectively, and θ can be related to... Figure 5 In the CNN correlation, Z(θ) is an unknown normalization constant.

[0147] Secondly, parameterized functions f can be used. θTo solve the K-class classification task, this function maps each sample point x to K real numbers (logits), and these logits are parameterized by the Softmax function to a class distribution, where K is an integer greater than or equal to 2; as shown in Equation (2):

[0148]

[0149] Among them, f θ (x)[y] represents f θ (x) The y-th index, i.e., the logits corresponding to the y-th class identifier.

[0150] For example, it can be done by adjusting f θ The obtained logits are reinterpreted to define the joint probability distribution p between the sample data and the labels. θ (x,y) and the probability distribution p of the sample data θ (x). Without changing f θ In this case, logits can be reused to define an energy model. Figure 5 The energy function in the equation is used for the joint distribution of data points x and labels y, where p θ (x,y) is shown in equation (3):

[0151]

[0152] For example, y is marginalized so that JEM obtains a non-normalized density model of sample point x, as shown in equation (4):

[0153]

[0154] At this point, LogSumExp of logits can be reused to define the energy function of the sample point x, as shown in equation (5):

[0155] E θ (x)=-LogSumExp y (f θ (x)[y])=-log∑ y exp(f θ (x)[y]) (5)

[0156] Through the above calculations, JEM uses logits to represent the energy function, defines the density function of the sample data and the joint density between the sample data and its label, reinterprets the classifier as a joint energy model, and reveals the generative model hidden in each classification model.

[0157] Therefore, the process of simultaneously training the initial generation model and the initial classification model is a process of simultaneously optimizing the initial generation model and the initial classification model. In the process of simultaneous optimization, combined with sampling of sample data and adversarial training operations, natural real speech can be reconstructed, thereby reducing the loss of acoustic features or voiceprint features carried by real speech during feature extraction. Furthermore, through adversarial training, the decision boundary between fake speech and real speech can be learned more effectively.

[0158] As described above, the anti-spoofing voiceprint authentication method provided in this application involves: extracting features from sample data to obtain sample acoustic features; performing speaker identification processing on the sample voiceprints corresponding to the sample acoustic features using an initial classification model to obtain sample classification results; reconstructing the sample acoustic features using an initial generation model to obtain sample generation results; and adjusting the parameters of the initial classification model based on the speaker representation and sample classification results to obtain a classification model. Simultaneously, adjusting the parameters of the initial generation model based on the authenticity of the sample data and the sample generation results to obtain a generation model. Thus, through the above process, targeted supervised training of the initial generation model and the initial classification model is achieved, thereby improving the accuracy of the generation model in authenticity identification and the classification model in speaker identification.

[0159] Based on the foregoing embodiments, the anti-spoofing voiceprint authentication method provided in this application can extract the acoustic feature set from the audio to be authenticated in the following ways:

[0160] The acoustic feature set is obtained by extracting features from the audio to be authenticated using WavLM.

[0161] For example, such as Figure 4 As shown, WavLM can include a CNN and multiple encoders to achieve comprehensive feature extraction of the audio to be authenticated.

[0162] As can be seen from the above, the anti-counterfeiting voiceprint authentication method provided in this application extracts features from the audio to be authenticated using WavLM to obtain an acoustic feature set, which can improve the integrity and comprehensiveness of the features in the acoustic feature set.

[0163] Compared to the SASV model in related technologies, the JEM architecture adopted in this application has the following advantages:

[0164] First, the JEM architecture possesses joint generation and speaker discrimination capabilities. SASV not only needs to accurately identify the speaker's identity but also needs to distinguish whether it is a speech spoofing attack. This requires the SASV model to not only have classification capabilities but also strong anomaly detection or spoofing identification capabilities. JEM's joint model architecture can perform classification and generation tasks simultaneously and can also measure the authenticity of audio data through an energy function. When faced with forged audio, the energy model will output a higher energy value, thereby being able to identify spoofing behavior.

[0165] Secondly, the JEM architecture possesses robustness against attacks. The speech spoofing attacks faced by the SASV model are essentially adversarial attacks against the system, that is, forged speech attempts to deceive the verification system. JEM, by virtue of the high energy output of its energy function for forged speech, exhibits strong robustness against forged speech attacks. This improves the robustness of JEM in recognizing forged speech, thereby reducing the probability of it being affected by forged speech.

[0166] Secondly, JEM features a unified and simplified architecture. Related technologies like SASV typically require separate training of the speaker verification model and the spoofing detection model, leading to system complexity and difficulty in optimization. JEM's advantage lies in its ability to simultaneously train both the generative and classification models within the same architecture, thereby reducing reliance on multiple models, simplifying the system architecture, and making the system more efficient and concise.

[0167] In summary, the technical solution provided by the embodiments of this application has at least the following technical advantages:

[0168] First, the technical solution provided in this application can effectively resist voice conversion and voice cloning attacks. Traditional voiceprint authentication systems are susceptible to voice conversion and voice cloning attacks. However, the technical solution provided in this application, by analyzing the synthesis characteristics and abnormal patterns of sound, can effectively detect and reject forged audio, thereby improving the solution's ability to resist forged voice attacks.

[0169] Secondly, this scheme has the advantages of joint optimization and information sharing. This scheme jointly optimizes the generative model and the classification model, taking into account the training objectives of both, thereby fully utilizing various feature information in the sample data.

[0170] Third, this solution possesses dynamic adaptability and generalization capabilities. By expanding the sample data based on constantly evolving attack methods and synthesis techniques, and adjusting and optimizing the parameters of the JEM model, the JEM model can update its identification strategy in a timely manner, improving its ability to respond to new attacks. This flexible learning mechanism enables it to adapt to complex and ever-changing security threats.

[0171] With the rapid development of deepfake technology, deepfake voice authentication poses a significant security threat to various industries. For example, in the financial, government, and enterprise sectors, traditional voiceprint authentication systems are vulnerable to deception attacks, such as voice synthesis attacks. Applying the technical solutions provided in this application to scenarios such as mobile phone unlocking, in-vehicle systems, smart homes, finance, government, and enterprises can improve the security and credibility of voice-forged authentication, enhancing system security in areas such as identity verification, security protection, forensic evidence collection, and even enterprise communications.

[0172] Based on the foregoing embodiments, this application also provides an anti-counterfeiting voiceprint authentication device. Figure 6 This is a schematic diagram of the anti-counterfeiting voiceprint authentication device provided in the embodiments of this application, as shown below. Figure 6 As shown, the anti-counterfeiting voiceprint authentication device 6 includes:

[0173] Module 601 is used to acquire the audio to be authenticated;

[0174] The identification module 602 is used to extract an acoustic feature set from the audio to be authenticated; perform forged audio identification processing on the acoustic feature set to obtain a first identification result; and perform speaker identification processing on the voiceprint feature set corresponding to the acoustic feature set to obtain a second identification result.

[0175] The determination module 603 is used to determine the result of anti-spoofing voiceprint authentication for the audio to be authenticated based on the first authentication result and the second authentication result.

[0176] In some embodiments, the identification module 602 is used to perform OOD detection on features in the acoustic feature set to obtain a first identification result.

[0177] In some embodiments, the identification module 602 is used to determine the quantization result set corresponding to the acoustic feature set; determine the probability distribution state of the quantization result set; fit the fitted audio data based on the probability distribution state; and perform OOD detection on the fitted audio data to obtain a first identification result.

[0178] In some embodiments, JEM includes a generative model and a classification model that are jointly generated; and an identification module 602 is used to perform OOD detection on features in the acoustic feature set through the generative model to obtain a first identification result.

[0179] In some embodiments, the identification module 602 is used to perform weighted fusion processing on the acoustic features in the acoustic feature set to obtain a weighted fusion result; and to perform feature extraction processing on the weighted fusion result through the MFA model to obtain a voiceprint feature set.

[0180] In some embodiments, JEM includes a generative model and a classification model that are jointly generated; the identification module 602 is used to perform speaker identification processing on the voiceprint feature set through the classification model to obtain a second identification result.

[0181] In some embodiments, the identification module 602 is used to perform multi-classification processing on the features in the voiceprint feature set through a classification model to obtain a speaker category probability set of the audio to be authenticated; and to determine a second identification result based on the speaker category probability set.

[0182] In some embodiments, a forged audio identification process is performed on the acoustic feature set to obtain a first identification result, which is achieved through the generative model in JEM; a speaker identification process is performed on the voiceprint feature set corresponding to the acoustic feature set to obtain a second identification result, which is achieved through the classification model of JEM; the acquisition module 601 is used to acquire sample data;

[0183] The aforementioned anti-spoofing audio authentication device also includes a training module, which is used to jointly train the initial generation model and the initial classification model based on sample data to obtain the generation model and the classification model.

[0184] In some embodiments, the training module is used to perform feature extraction processing on sample data to obtain sample acoustic features; perform speaker identification processing on the sample voiceprint features corresponding to the sample acoustic features through an initial classification model to obtain sample classification results; reconstruct the sample acoustic features through an initial generation model to obtain sample generation results; adjust the parameters of the initial classification model based on the speaker identifier and sample classification results of the sample data to obtain a classification model; and simultaneously, adjust the parameters of the initial generation model based on the authenticity identifier and sample generation results of the sample data to obtain a generation model.

[0185] In some embodiments, the identification module 602 is used to extract features from the audio to be authenticated using WavLM to obtain an acoustic feature set.

[0186] Based on the foregoing embodiments, this application also provides an electronic device. Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 7 As shown, the electronic device 7 includes a processor 701 and a memory 702; wherein, the memory stores a computer program; when the computer program is executed by the processor, it can implement the anti-counterfeiting voiceprint authentication method as described above.

[0187] Based on the foregoing embodiments, this application also provides a computer-readable storage medium storing a computer program; when the computer program is executed by the processor of an electronic device, it can implement the anti-counterfeiting voiceprint authentication method as described above.

[0188] Based on the foregoing embodiments, this application also provides a computer program product, which includes a computer program that, when executed by a processor of an electronic device, can implement the anti-counterfeiting voiceprint authentication method as described above.

[0189] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0190] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict.

[0191] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0192] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0193] It should be noted that the aforementioned computer-readable storage media can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various electronic devices that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0194] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0195] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0196] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware nodes. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0197] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0198] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0199] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0200] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for anti-forgery voiceprint authentication, characterized in that, The method includes: Obtain the audio to be authenticated; Extract a set of acoustic features from the audio to be authenticated; Perform forged audio identification processing on the acoustic feature set to obtain a first identification result; Speaker identification processing is performed on the voiceprint feature set corresponding to the acoustic feature set to obtain a second identification result; Based on the first authentication result and the second authentication result, the result of anti-spoofing voiceprint authentication for the audio to be authenticated is determined.

2. The method according to claim 1, characterized in that, The process of performing fake audio identification processing on the acoustic feature set to obtain a first identification result includes: OOD detection is performed on the features in the acoustic feature set to obtain the first identification result.

3. The method according to claim 2, characterized in that, The step of performing OOD detection on the features in the acoustic feature set to obtain the first identification result includes: Determine the set of quantization results corresponding to the set of acoustic features; Determine the probability distribution state of the quantization result set; The fitted audio data is obtained based on the probability distribution state fitting; The OOD detection is performed on the fitted audio data to obtain the first identification result.

4. The method according to claim 2 or 3, characterized in that, The joint energy model includes a jointly generated model and a classification model; the step of performing OOD detection on the features in the acoustic feature set to obtain the first discrimination result includes: The first identification result is obtained by performing OOD detection on the features in the acoustic feature set using the generative model.

5. The method according to claim 1, characterized in that, Before performing speaker identification processing on the voiceprint feature set corresponding to the acoustic feature set, the method further includes: A weighted fusion process is performed on the acoustic features in the acoustic feature set to obtain a weighted fusion result; The voiceprint feature set is obtained by performing feature extraction processing on the weighted fusion result using the MFA model.

6. The method according to claim 1, characterized in that, The joint energy model includes a jointly generated model and a classification model; the speaker identification process performed on the voiceprint feature set corresponding to the acoustic feature set to obtain a second identification result includes: The speaker identification process is performed on the voiceprint feature set using the classification model to obtain the second identification result.

7. The method according to claim 6, characterized in that, The step of performing speaker identification processing on the voiceprint feature set through the classification model to obtain the second identification result includes: The features in the voiceprint feature set are subjected to multi-classification processing by the classification model to obtain the speaker category probability set of the audio to be authenticated. The second identification result is determined based on the speaker category probability set.

8. The method according to claim 1, characterized in that, The method involves performing forged audio identification processing on the acoustic feature set to obtain a first identification result, which is achieved through the generative model in the joint energy model; performing speaker identification processing on the voiceprint feature set corresponding to the acoustic feature set to obtain a second identification result, which is achieved through the classification model of the joint energy model; the method further includes: Obtain sample data; The initial generative model and the initial classification model are jointly trained based on the sample data to obtain the generative model and the classification model.

9. The method according to claim 8, characterized in that, The step of jointly training the initial generative model and the initial classification model based on the sample data to obtain the generative model and the classification model includes: The sample data is subjected to feature extraction processing to obtain the sample acoustic features; The initial classification model is used to perform speaker identification processing on the sample voiceprint features corresponding to the sample acoustic features to obtain the sample classification result. The acoustic features of the sample are reconstructed using the initial generation model to obtain the sample generation result; Based on the speaker identifiers of the sample data and the sample classification results, the parameters of the initial classification model are adjusted to obtain the classification model. Simultaneously, based on the authenticity identifiers of the sample data and the sample generation results, the parameters of the initial generation model are adjusted to obtain the generation model.

10. The method according to claim 1, characterized in that, The extraction of the acoustic feature set from the audio to be authenticated includes: The acoustic feature set is obtained by extracting features from the audio to be authenticated using the waveform language model WavLM.

11. An anti-counterfeiting voiceprint authentication device, characterized in that, The anti-counterfeiting voiceprint authentication device includes: The acquisition module is used to acquire the audio to be authenticated; The authentication module is used to extract an acoustic feature set from the audio to be authenticated; perform forged audio authentication processing on the acoustic feature set to obtain a first authentication result; and perform speaker authentication processing on the voiceprint feature set corresponding to the acoustic feature set to obtain a second authentication result. The determining module is used to determine the result of anti-spoofing voiceprint authentication for the audio to be authenticated based on the first authentication result and the second authentication result.

12. An electronic device, characterized in that, The electronic device includes a processor and a memory; wherein the memory stores a computer program; when the computer program is executed by the processor, it is able to implement the anti-counterfeiting voiceprint authentication method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The readable storage medium stores a computer program; when the computer program is executed by the processor of the electronic device, it is able to implement the anti-counterfeiting voiceprint authentication method as described in any one of claims 1 to 10.

14. A computer program product, characterized in that, The computer program product includes a computer program; when the computer program is executed by the processor of an electronic device, it is capable of implementing the anti-counterfeiting voiceprint authentication method as described in any one of claims 1 to 10.