Identity Authentication Method, Device, Equipment and Storage Medium

By integrating attack audio recognition with identity authentication using neural networks, the system enhances the efficiency and accuracy of voice-based identity verification by leveraging a richer set of predictive features.

CN114974263BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210524001.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-07-15
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

In the existing identity authentication system based on voiceprint verification, attack audio recognition is separated from the identity authentication function, resulting in difficulty in improving authentication accuracy and efficiency.

Method used

Combining attack audio recognition with identity authentication, extract acoustic features through neural networks of various structures, fuse attack audio prediction results and voiceprint feature similarity for identity prediction.

Benefits of technology

Improve the accuracy and efficiency of identity authentication and effectively reduce the possibility of being attacked.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114974263B_ABST
    Figure CN114974263B_ABST
Patent Text Reader

Abstract

The present disclosure provides an identity authentication method, apparatus, device, and storage medium, relating to the field of computer technologies, and particularly to the fields of voice technologies, artificial intelligence technologies, and deep learning technologies. The implementation solution is as follows: obtain the acoustic features of the audio to be authenticated; based on the acoustic features, perform attack audio recognition on the audio to be authenticated to obtain the attack audio prediction result of the audio to be authenticated; based on the acoustic features, obtain the similarity of the voiceprint features between the audio to be authenticated and the registered audio; and perform identity prediction based on the attack audio prediction result and the similarity of the voiceprint features to obtain the identity authentication result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to the fields of voice technologies, artificial intelligence technologies, and deep learning technologies. Specifically, the present disclosure relates to an identity authentication method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), including both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0003] With the rapid development of technology, using voiceprint for verification has been widely applied in fields such as security, the Internet, and finance. Currently, a system for verifying based on voiceprint usually first verifies whether the audio is an attack audio. If it is not an attack audio, then the voiceprint features in the audio are extracted for verification.

[0004] The methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, any method described in this section should not be considered prior art merely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0005] The present disclosure provides an identity authentication method, apparatus, electronic device, computer-readable storage medium, and computer program product.

[0006] According to one aspect of the present disclosure, there is provided an identity authentication method, including: obtaining acoustic features of an audio to be authenticated; based on the acoustic features, performing attack audio recognition on the audio to be authenticated to obtain an attack audio prediction result of the audio to be authenticated; based on the acoustic features, obtaining a voiceprint feature similarity between the audio to be authenticated and a registered audio; and performing identity prediction based on the attack audio prediction result and the voiceprint feature similarity to obtain an identity authentication result.

[0007] According to another aspect of the present disclosure, there is provided a method for training an identity authentication model, including: obtaining a first sample data set, wherein each sample data in the first sample data set includes a test audio, a registered audio, and a corresponding identity authentication label; for each sample data in the first sample data set, performing the following operations: obtaining a first acoustic feature corresponding to the test audio and a second acoustic feature corresponding to the registered audio in the sample data; inputting the first acoustic feature into at least one pre-trained first feature extraction network respectively to obtain at least one first feature, wherein each first feature in the at least one first feature is used to determine whether the audio to be authenticated is an attack audio, and each first feature extraction network in the at least one first feature extraction network is constructed based on a neural network with a different structure; inputting the first acoustic feature and the second acoustic feature into at least one pre-trained second feature extraction network respectively to obtain at least one test audio feature corresponding to the first acoustic feature and a registered audio feature corresponding to each test audio feature in the at least one test audio feature, wherein each second feature extraction network in the at least one second feature extraction network is constructed based on a neural network with a different structure; inputting the at least one first feature into an attack audio prediction network to obtain an attack audio prediction result of the test audio; obtaining a feature similarity between each test audio feature in the at least one test audio feature and the registered audio feature corresponding to the test audio feature; inputting the attack audio prediction result and the at least one feature similarity corresponding to the at least one test audio feature into an identity prediction network to obtain an identity prediction result; and adjusting the parameters of the attack audio prediction network and the parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data.

[0008] According to another aspect of the present disclosure, there is provided an identity authentication device, including: a first acquisition unit configured to acquire an acoustic feature of an audio to be authenticated; an identification unit configured to perform attack audio identification on the audio to be authenticated based on the acoustic feature to obtain an attack audio prediction result of the audio to be authenticated; a second acquisition unit configured to acquire a voiceprint feature similarity between the audio to be authenticated and a registered audio based on the acoustic feature; and a prediction unit configured to perform identity prediction based on the attack audio prediction result and the voiceprint feature similarity to obtain an identity authentication result.

[0009] According to another aspect of the present disclosure, there is provided a training device for an identity authentication model, including: a third acquisition unit configured to acquire a first sample data set, wherein each sample data in the first sample data set includes a test audio, a registered audio, and a corresponding identity authentication label; an execution unit configured to perform operations of the following sub-units for each sample data in the first sample data set: a fifth acquisition sub-unit configured to acquire a first acoustic feature corresponding to the test audio and a second acoustic feature corresponding to the registered audio in the sample data; a first input sub-unit configured to input the first acoustic feature into at least one pre-trained first feature extraction network respectively to obtain at least one first feature, wherein each of the at least one first features is used to determine whether the audio to be authenticated is an attack audio, and each of the at least one first feature extraction networks is constructed based on neural networks with different structures; a second input sub-unit configured to input the first acoustic feature and the second acoustic feature together into at least one pre-trained second feature extraction network respectively to obtain at least one test audio feature corresponding to the first acoustic feature and a registered audio feature corresponding to each of the at least one test audio features, wherein each of the at least one second feature extraction networks is constructed based on neural networks with different structures; a third input sub-unit configured to input the at least one first feature into an attack audio prediction network to obtain an attack audio prediction result of the test audio; a sixth acquisition sub-unit configured to acquire a feature similarity between each of the at least one test audio features and the registered audio feature corresponding to the test audio feature; a fourth input sub-unit configured to input the attack audio prediction result and at least one feature similarity corresponding to the at least one test audio feature into an identity prediction network to obtain an identity prediction result; and an adjustment sub-unit configured to adjust parameters of the attack audio prediction network and parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data.

[0010] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above identity authentication method or the above training method of the identity authentication model.

[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above identity authentication method or the above training method of the identity authentication model.

[0012] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the above-mentioned identity authentication method or the training method of the above-mentioned identity authentication model.

[0013] According to one or more embodiments of the present disclosure, by combining attack audio recognition with identity authentication, the feature information on which identity prediction is based can be made more abundant, improving the authentication efficiency and the authentication accuracy at the same time.

[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings exemplarily show embodiments and constitute a part of the description, and are used together with the written description of the description to explain the exemplary implementation manners of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily the same elements.

[0016] Figure 1 The schematic diagram of an exemplary system in which the various methods described herein can be implemented according to an embodiment of the present disclosure is shown;

[0017] Figure 2 The flowchart of the identity authentication method according to an embodiment of the present disclosure is shown;

[0018] Figure 3 The schematic structural diagram of an identity authentication model according to an exemplary embodiment of the present disclosure is shown;

[0019] Figure 4 The flowchart of the training method of the identity authentication model according to an embodiment of the present disclosure is shown;

[0020] Figure 5 The schematic structural diagram of an attack audio recognition network according to an exemplary embodiment of the present disclosure is shown;

[0021] Figure 6 The structural block diagram of an identity authentication device according to an embodiment of the present disclosure is shown;

[0022] Figure 7 The structural block diagram of a training device of an identity authentication model according to an embodiment of the present disclosure is shown;

[0023] Figure 8 The structural block diagram of an exemplary electronic device capable of implementing the embodiment of the attack audio recognition network of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0025] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0026] The terms used in the description of various examples in the present disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.

[0027] With the development of artificial intelligence technology, in the fields of security, Internet, finance, smart cities, etc., it is becoming more and more popular to use voiceprint for identity authentication to log in to the system. Generally, an identity authentication (Auto Speaker Verification, ASV) system has the possibility of being broken by some means. Therefore, in order to reduce the possibility of being broken, a general system usually includes anti-attack (also called countermeasure, CM) functions such as attack audio recognition and liveness detection.

[0028] In the related art, attack audio recognition and identity authentication functions are processed separately, that is, first verify whether the audio is normal audio, rather than attack audio such as recorded playback, voice synthesis, splicing, or voice conversion; if it is determined that the audio is not attack audio, then extract the voiceprint features of the audio to compare with the registered audio, so as to draw a conclusion on whether the identity authentication is passed. Since the two systems of attack audio recognition and identity authentication are separated, it is difficult to further improve the accuracy and efficiency of identity authentication.

[0029] The embodiments of the present disclosure provide an identity authentication method, which can combine attack audio recognition with identity authentication, make the information on which identity prediction is based more abundant, and improve the authentication efficiency while also improving the authentication accuracy.

[0030] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0031] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein may be implemented according to embodiments of the present disclosure. Referring Figure 1 thereto, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.

[0032] In embodiments of the present disclosure, the server 120 may run one or more services or software applications that enable the execution of an authentication method or a training method of an authentication model.

[0033] In certain embodiments, the server 120 may also provide other services or software applications, which may include non-virtual environments and virtual environments. In certain embodiments, these services may be provided as web-based services or cloud services, such as provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.

[0034] In Figure 1 the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0035] Users may use the client devices 101, 102, 103, 104, 105, and / or 106 to obtain audio to be authenticated. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.

[0036] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smart phones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. The client devices are capable of executing various different applications such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.

[0037] Network 110 may be any type of network known to those skilled in the art, and it may support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0038] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that may be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.

[0039] The computing unit in server 120 can run one or more operating systems including any of the above - mentioned operating systems and any commercially available server operating system. Server 120 can also run any one of a variety of additional server applications and / or middle - layer applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0040] In some embodiments, server 120 can include one or more applications to analyze and combine data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 can also include one or more applications to display data feeds and / or real - time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.

[0041] In some embodiments, server 120 can be a server of a distributed system or a server combined with a blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to solve the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0042] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network - based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be relational databases, for example. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.

[0043] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key - value repositories, object repositories, or conventional repositories supported by a file system.

[0044] Figure 1System 100 can be configured and operated in various ways to enable the application of various methods and devices described according to the present disclosure.

[0045] According to an embodiment of the present disclosure, as Figure 2 shown, an identity authentication method is provided, including: Step S201, obtaining the acoustic features of the audio to be authenticated; Step S202, based on the acoustic features, performing attack audio recognition on the audio to be authenticated to obtain the attack audio prediction result of the audio to be authenticated; Step S203, based on the acoustic features, obtaining the voiceprint feature similarity between the audio to be authenticated and the registered audio; and Step S204, performing identity prediction based on the attack audio prediction result and the voiceprint feature similarity to obtain the identity authentication result.

[0046] Thus, by combining attack audio recognition with identity authentication, the feature information on which identity prediction is based can be made more abundant, improving the authentication efficiency while also improving the authentication accuracy.

[0047] Among them, the attack audio may refer to the audio obtained by processing the user's voice, or the audio synthesized or converted by a specific voice processing device, such as the audio obtained by recording the user's voice, the audio obtained by splicing multiple segments of the user's voice, the audio synthesized by a voice synthesizer and other devices, the audio obtained by conversion through a voice converter, etc. That is to say, the attack audio is not the directly collected user voice.

[0048] Since the directly collected real user voice (i.e., non-attack audio) is the audio without any post-processing, the real user voice may be superimposed with environmental noise, device noise, and signals generated by analog-to-digital conversion during the collection process; for the attack audio, such as the synthesized voice synthesized by a device, it usually does not have superimposed environmental noise; and for another example, the recording of the user's voice usually has more superimposed environmental noise, device noise, etc. compared to a normal audio. That is to say, generally, the difference between the attack audio and the normal audio lies in the differences in the superimposed environmental noise / device noise, etc. By extracting these differences, the attack audio and the normal audio can be distinguished.

[0049] In some embodiments, the audio to be authenticated may be an attack audio or a normal audio.

[0050] In some embodiments, some preprocessing operations can be first performed on the audio to be authenticated, for example, it may include removing the silent segments in the audio, frame segmentation, etc. Among them, frame segmentation can divide a segment of the audio to be authenticated into multiple audio frames according to a frame length of 25 ms and a frame shift of 10 ms. Subsequently, based on the preprocessing operation on the audio to be authenticated, the acoustic features of the audio to be authenticated are extracted.

[0051] In some embodiments, the acoustic features may include, but are not limited to, Mel Frequency Cepstral Coefficient (MFCC) features, Constant Q Cepstral Coefficients (CQCC) features, etc. In some embodiments, the above acoustic features can be extracted for each audio frame of the audio to be authenticated respectively.

[0052] In some embodiments, on the basis of extracting the above acoustic features, a sliding window (the window length is, for example, 5 frames) can be further adopted to perform mean normalization on the acoustic features of multiple consecutive frames, so as to eliminate the possible signal discontinuities existing at both ends of each frame.

[0053] In some embodiments, the above acoustic features can be further processed by differential processing, for example, first-order differential processing, second-order differential processing, etc., and the acoustic features after differential processing are used as the input features of the attack audio recognition network and the identity authentication network. Thus, through differential processing, richer acoustic features can be obtained, so as to provide more feature information for the attack audio recognition network and the identity authentication network, and further improve the accuracy of recognition and prediction.

[0054] In some embodiments, based on the acoustic features, performing attack audio recognition on the audio to be authenticated to obtain the attack audio prediction result of the audio to be authenticated may include: based on the acoustic features, respectively obtaining at least one first feature through at least one first feature extraction network, wherein each first feature in the at least one first feature is used to determine whether the audio to be authenticated is an attack audio, and each first feature extraction network in the at least one first feature extraction network is constructed based on neural networks with different structures; and obtaining the attack audio prediction result based on the at least one first feature.

[0055] Thus, the first feature is extracted through at least one feature extraction network with different structures, and the attack audio is recognized and predicted based on the at least one first feature, so as to improve the accuracy of prediction.

[0056] In some embodiments, the first feature extraction network can be constructed based on one or more of neural networks with different structures such as UBM-GMM, SVM, DNN, CNN, LSTM, Conformer, TDNN, etc. When multiple first feature extraction networks are used to extract the first features of the audio to be authenticated, each first feature extraction network can be constructed based on the above different neural networks respectively. Thus, by taking advantage of the different focuses of different neural networks on feature extraction, richer audio feature information (i.e., the first feature) can be obtained.

[0057] In some embodiments, the acoustic features corresponding to each frame of the audio to be authenticated can be sequentially input into the first feature extraction network to obtain the first feature corresponding to each frame, and then the first features corresponding to each frame are averaged to obtain the first feature of the audio to be authenticated.

[0058] In some embodiments, the attack audio prediction result can be obtained by inputting the first feature of the audio to be authenticated into a prediction network, which can be constructed based on one of neural networks such as DNN, CNN, and LSTM. By inputting the above first feature into the prediction network, the attack audio prediction result output by the prediction network can be obtained.

[0059] In some embodiments, the attack audio prediction result can be the probability that the audio to be authenticated is an attack audio.

[0060] In some embodiments, the attack audio prediction result can also be a two-dimensional vector, including the probability that the audio to be authenticated is an attack audio and the probability that the audio to be authenticated is not an attack audio, respectively.

[0061] In some embodiments, the number of at least one first feature extraction network can be at least two. Obtaining the attack audio prediction result based on at least one first feature can include: obtaining a first fusion feature vector of at least one first feature through a feature fusion network based on an attention mechanism; and inputting the first fusion feature vector into an attack audio prediction network to obtain the attack audio prediction result.

[0062] Thus, by fusing multiple first features through the attention mechanism, different weights can be given to different first features based on the trained feature fusion network, so as to realize the characterization of the importance degree of multiple first features. Predicting based on the feature vector obtained by the above feature fusion can further improve the prediction accuracy.

[0063] In some embodiments, when multiple first features of the audio to be authenticated are extracted by different first feature extraction networks respectively, before inputting the first features into the prediction network, the first features can be first fused through a feature fusion network to obtain a first fusion feature vector.

[0064] In some embodiments, the above feature fusion network can be a network based on an attention mechanism, which can calculate the following formula for multiple first features (denoted as h i , where i ∈ [1, N], and N is the number of first feature extraction networks) to obtain the first fusion feature vector h cm :

[0065] A = softmax(tanh(HT W1)W2)

[0066] h cm = HA

[0067] wherein, the matrix H consists of multiple first features, that is, H = [h1, h2, …, h N , W1 and W2 are respectively the weight matrices of the feature fusion network, and A is the weight coefficient vector obtained through calculation.

[0068] In some embodiments, obtaining the voiceprint feature similarity between the audio to be authenticated and the registered audio based on acoustic features may include: based on acoustic features, respectively obtaining at least one second feature through at least one second feature extraction network, wherein each second feature in the at least one second feature is used for identity authentication, and each second feature extraction network in the at least one second feature extraction network is constructed based on neural networks with different structures; and for each second feature in the at least one second feature, respectively obtaining the voiceprint feature similarity between the second feature and the voiceprint feature of the registered audio.

[0069] Thus, by extracting the second feature through one or more different feature extraction networks and respectively comparing the feature similarity with the registered audio based on at least one second feature, more feature information of the audio to be identified and the registered audio can be obtained through different networks, and combining this information can further improve the prediction accuracy.

[0070] In some embodiments, the second feature extraction network can be constructed based on one or more of neural networks with different structures such as UBM-GMM, SVM, DNN, CNN, LSTM, Conformer, TDNN, etc. When applying multiple second feature extraction networks to extract the second feature of the audio to be authenticated, each second feature extraction network can be respectively constructed based on the above different neural networks. Thus, by taking advantage of the different focuses of different neural networks on feature extraction, the feature similarity can be compared with the registered audio based on different focuses, thereby further improving the prediction accuracy.

[0071] In some embodiments, the acoustic features corresponding to each frame of the audio to be authenticated can be sequentially input into the second feature extraction network to obtain the second feature corresponding to each frame, and then the second features corresponding to each frame are averaged to obtain the second feature of the audio to be authenticated.

[0072] In some embodiments, when obtaining the second feature of the audio to be authenticated each time, the acoustic feature of the registered audio can be input into each second feature extraction network simultaneously to obtain the voiceprint feature of the registered audio corresponding to the second feature of each audio to be authenticated, so as to calculate the voiceprint feature similarity based on the second feature and its corresponding registered audio voiceprint feature respectively.

[0073] In some embodiments, before the registered audio is input into the second feature extraction network, preprocessing operations, acoustic feature extraction operations, differential operations, etc. similar to the above operations can also be performed.

[0074] In some embodiments, the voiceprint feature of the registered audio can be obtained by inputting it into each second feature extraction network in advance to obtain the voiceprint feature of the registered audio corresponding to each second feature extraction network. Thus, by pre-extracting and saving the voiceprint feature of the registered audio, computing resources can be saved and the efficiency of identity authentication can be improved.

[0075] In some embodiments, the voiceprint feature similarity between each second feature and its corresponding registered audio voiceprint feature can be calculated by calculating the cosine similarity between vectors, etc. When the second feature extraction is performed by multiple second feature extraction networks respectively, for the same audio to be authenticated, a voiceprint feature similarity score can be obtained based on each second feature extraction network, so as to form a multi-dimensional similarity score vector.

[0076] In some embodiments, performing identity prediction based on the attack audio prediction result and the voiceprint feature similarity to obtain the identity authentication result may include: generating a second fusion feature vector based on the attack audio prediction result and at least one voiceprint feature similarity corresponding to at least one second feature; and performing identity prediction based on the second fusion feature vector to obtain the identity authentication result of the audio to be authenticated.

[0077] Thus, by combining the attack audio prediction result with multiple voiceprint feature similarities to form a fusion feature vector and performing prediction based on this vector, the efficient combination of attack audio recognition and identity authentication can be realized, and the recognition efficiency and accuracy can be improved.

[0078] In some embodiments, the attack audio prediction result (for example, it can be a two-dimensional vector, respectively including the probability that the audio to be authenticated is an attack audio and the probability that the audio to be authenticated is not an attack audio) and at least one second feature (for example, it can be a one-dimensional or multi-dimensional voiceprint feature similarity score vector) can be concatenated to generate a second fusion feature vector, and prediction is performed based on this fusion feature vector to obtain the final identity authentication result.

[0079] In some embodiments, the corresponding prediction result can be obtained by inputting the second fusion feature vector into a prediction network, which can be constructed based on one of neural networks such as DNN, CNN, LSTM, etc. The prediction result of the prediction network can be a label for indicating whether the audio to be authenticated and the registered audio are of the same user; it can also be the probability that the audio to be authenticated and the registered audio are of the same user. When the probability is greater than a certain threshold, it can be determined that the audio to be authenticated passes the identity authentication.

[0080] Figure 3 FIG. 4 shows a schematic structural diagram of an identity authentication model according to an exemplary embodiment of the present disclosure.

[0081] In some exemplary embodiments, as Figure 3 shown, the identity authentication model 300 includes N first feature extraction networks (denoted as CM-1, CM-2,..., CM-N) and M second feature extraction networks (denoted as ASV-1, ASV-2,..., ASV-M), where both N and M are positive integers greater than 1, and where each first feature extraction network is constructed based on a neural network with a different structure, and each second feature extraction network is also constructed based on a neural network with a different structure.

[0082] The acoustic features of the audio to be authenticated can be input into the N first feature extraction networks. Correspondingly, through the N first feature extraction networks, the first features corresponding to each first feature extraction network (denoted as h1, h2,..., h N ) can be obtained respectively, and the above first features are input into the feature fusion network based on the attention mechanism to obtain the first fusion feature vector h cm ; subsequently, the first fusion feature vector h cm is input into the attack audio prediction network to obtain a two-dimensional vector s cm , where the two values in s cm respectively represent the probability that the audio to be authenticated is an attack audio and the probability that the audio to be authenticated is not an attack audio.

[0083] At the same time, the acoustic features of the audio to be authenticated and the acoustic features of the registered audio can be input into the M second feature extraction networks simultaneously. Correspondingly, through the M second feature extraction networks, the voiceprint feature similarity scores between the audio to be authenticated and the registered audio (denoted as s cv1 , s cv2 ,..., s cvM ) can be obtained respectively, thus constituting a voiceprint feature similarity vector s cv .

[0084] Subsequently, by combining the two-dimensional vector s cm and the voiceprint feature similarity vector s cvPerform splicing to obtain a second fused feature vector, and input it into a prediction network to obtain a prediction result indicating whether the identity authentication passes.

[0085] In some embodiments, such as Figure 4 As shown, a method for training an identity authentication model is provided, including: Step S401, obtaining a first sample data set, where each sample data in the first sample data set includes a test audio, a registered audio, and a corresponding identity authentication label; for each sample data in the first sample data set, perform the following operations: Step S402, obtaining a first acoustic feature corresponding to the test audio and a second acoustic feature corresponding to the registered audio in the sample data; Step S403, inputting the first acoustic feature into at least one pre-trained first feature extraction network respectively to obtain at least one first feature, where each first feature in the at least one first feature is used to determine whether the audio to be authenticated is an attack audio, and each first feature extraction network in the at least one first feature extraction network is constructed based on neural networks with different structures; Step S404, inputting the first acoustic feature and the second acoustic feature into at least one pre-trained second feature extraction network respectively to obtain at least one test audio feature corresponding to the first acoustic feature and a registered audio feature corresponding to each test audio feature in the at least one test audio feature, where each second feature extraction network in the at least one second feature extraction network is constructed based on neural networks with different structures; Step S405, inputting the at least one first feature into an attack audio prediction network to obtain an attack audio prediction result of the test audio; Step S406, obtaining a feature similarity between each test audio feature in the at least one test audio feature and the registered audio feature corresponding to the test audio feature; Step S407, inputting the attack audio prediction result and the at least one feature similarity corresponding to the at least one test audio feature into an identity prediction network to obtain an identity prediction result; and Step S408, adjusting the parameters of the attack audio prediction network and the parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data.

[0086] In some embodiments, the method for pre-training each first feature extraction network may specifically include the following operations. First, based on neural networks with different structures (such as UBM-GMM, SVM, DNN, CNN, LSTM, Conformer, TDNN, etc. neural networks), a complete one can be constructed respectively, as Figure 5 As shown, including an input layer, multiple hidden layers (hidden layer h1, hidden layer h2,..., hidden layer hn), and an output layer.

[0087] Subsequently, training data for training the attack audio recognition network is obtained (which can be open-source datasets such as Aishell, VoxCeleb, ASVspoof, etc.). Each training data includes a sample audio and a label indicating whether it is an attack audio. By inputting the acoustic features of each sample audio into the above attack audio recognition network, an attack audio prediction result is obtained, and a loss function (such as the cross-entropy loss function) is calculated based on this prediction result and the label corresponding to the sample audio. Then, the parameters of the attack audio recognition network are adjusted based on this loss function, thus completing the training of the attack audio recognition network.

[0088] When constructing the above identity authentication model, each pre-trained first feature extraction network can be obtained based on the corresponding attack audio recognition network above. Specifically, the output layer of the above-trained attack audio recognition network can be removed, and only the input layer and multiple hidden layers are retained as a first feature extraction network. In some embodiments, the output layer and one or more adjacent hidden layers of the above-trained attack audio recognition network can also be removed, and only the input layer and the remaining multiple hidden layers are retained as a first feature extraction network.

[0089] Similarly, the pre-trained second feature extraction network can also be obtained through a training method similar to the above method, which will not be elaborated here.

[0090] Thus, an identity authentication model is constructed based on the above at least one pre-trained first feature extraction network and at least one second feature extraction network. The test audio and registration audio in each sample data are respectively input into the above identity authentication model, and then the output prediction result is obtained. A loss function (such as the cross-entropy loss function) is calculated based on the prediction result and the label corresponding to the sample data. The parameters are adjusted by minimizing this loss function, thus finally completing the training of the model.

[0091] In some embodiments, the number of at least one pre-trained first feature extraction network is at least two. Inputting at least one first feature into the attack audio prediction network to obtain the attack audio prediction result of the test audio may include: inputting at least one first feature into the feature fusion network to obtain a first fusion feature vector; and inputting the first fusion feature vector into the attack audio prediction network to obtain the attack audio prediction result of the test audio. And wherein, adjusting the parameters of the attack audio prediction network and the identity prediction network based on the identity prediction result and the corresponding identity authentication label of the sample data includes: adjusting the parameters of the feature fusion network, the attack audio prediction network, and the identity prediction network based on the identity prediction result and the corresponding identity authentication label of the sample data.

[0092] Thus, the identity authentication model trained by the above method combines attack audio recognition with identity authentication, making the feature information on which identity prediction is based more abundant. While improving the authentication efficiency, it also improves the authentication accuracy.

[0093] To further improve the performance of the above identity authentication model in terms of identity authentication accuracy, relevant staff respectively applied the solutions in the relevant technologies (i.e., the solution where the attack audio recognition system and the identity verification system are independent of each other) and the method provided by the embodiments of the present disclosure to conduct experiments based on the same test set. Among them, the identity authentication model constructed according to the method provided by the present disclosure respectively includes three first feature extraction networks and three second feature extraction networks, and the attack audio recognition network and the prediction network respectively apply four-layer DNN neural networks.

[0094] Relevant technical personnel respectively tested the identity authentication system in the relevant technologies (including an independent attack audio recognition system and an identity verification system) and the identity authentication model of the present disclosure based on the test sets Dev and Eval of ASVspoof 2019 (Speaker Recognition Attack Competition). The test results are as follows:

[0095]

[0096] It can be seen from the above test results that the identity authentication model of the present disclosure can effectively reduce the equal error rate, that is, the method of the present disclosure can effectively improve the accuracy of voiceprint identity authentication.

[0097] In some embodiments, as Figure 6 shown, an identity authentication device 600 is provided, including: a first acquisition unit 610 configured to acquire the acoustic features of the audio to be authenticated; an identification unit 620 configured to perform attack audio recognition on the audio to be authenticated based on the acoustic features to obtain the attack audio prediction result of the audio to be authenticated; a second acquisition unit 630 configured to acquire the voiceprint feature similarity between the audio to be authenticated and the registered audio based on the acoustic features; and a prediction unit 640 configured to perform identity prediction based on the attack audio prediction result and the voiceprint feature similarity to obtain the identity authentication result.

[0098] Among them, the operations of the units 610 - 640 in the identity authentication device 600 are similar to the operations of steps S201 - S204 in the above identity authentication method, and will not be elaborated here.

[0099] In some embodiments, the recognition unit may include: a first acquisition subunit, configured to respectively obtain at least one first feature through at least one first feature extraction network based on acoustic features, where each of the at least one first features is used to determine whether the audio to be authenticated is an attack audio, and each of the at least one first feature extraction networks is constructed based on neural networks with different structures; and a second acquisition subunit, configured to obtain an attack audio prediction result based on the at least one first feature.

[0100] In some embodiments, the number of the at least one first feature extraction networks may be at least two, and the second acquisition subunit may include: an acquisition module, configured to obtain a first fusion feature vector of the at least one first feature through a feature fusion network based on an attention mechanism; and a first input module, configured to input the first fusion feature vector into an attack audio prediction network to obtain an attack audio prediction result.

[0101] In some embodiments, the second acquisition unit may include: a third acquisition subunit, configured to respectively obtain at least one second feature through at least one second feature extraction network based on acoustic features, where each of the at least one second features is used for identity authentication, and each of the at least one second feature extraction networks is constructed based on neural networks with different structures; and a fourth acquisition subunit, configured to respectively obtain the acoustic feature similarity between each of the at least one second features and the voiceprint feature of the registered audio for each of the at least one second features.

[0102] In some embodiments, the prediction unit may include: a generation subunit, configured to generate a second fusion feature vector based on the attack audio prediction result and the at least one acoustic feature similarity corresponding to the at least one second feature respectively; and a prediction subunit, configured to perform identity prediction based on the second fusion feature vector to obtain an identity authentication result of the audio to be authenticated.

[0103] In some embodiments, such as Figure 7As shown, a training device 700 for an identity authentication model is provided, including: a third acquisition unit 710 configured to acquire a first sample data set, where each sample data in the first sample data set includes a test audio, a registered audio, and a corresponding identity authentication label; an execution unit 720 configured to perform operations of the following sub-units for each sample data in the first sample data set: a fifth acquisition sub-unit 721 configured to acquire a first acoustic feature corresponding to the test audio and a second acoustic feature corresponding to the registered audio in the sample data; a first input sub-unit 722 configured to input the first acoustic feature into at least one pre-trained first feature extraction network respectively to acquire at least one first feature, where each first feature in the at least one first feature is used to determine whether the audio to be authenticated is an attack audio, and each first feature extraction network in the at least one first feature extraction network is constructed based on a neural network with a different structure; a second input sub-unit 723 configured to input the first acoustic feature and the second acoustic feature together into at least one pre-trained second feature extraction network respectively to acquire at least one test audio feature corresponding to the first acoustic feature and a registered audio feature corresponding to each test audio feature in the at least one test audio feature, where each second feature extraction network in the at least one second feature extraction network is constructed based on a neural network with a different structure; a third input sub-unit 724 configured to input the at least one first feature into an attack audio prediction network to acquire an attack audio prediction result of the test audio; a sixth acquisition sub-unit 725 configured to acquire a feature similarity between each test audio feature in the at least one test audio feature and the registered audio feature corresponding to the test audio feature; a fourth input sub-unit 726 configured to input the attack audio prediction result and at least one feature similarity corresponding to the at least one test audio feature into an identity prediction network to acquire an identity prediction result; and an adjustment sub-unit 727 configured to adjust parameters of the attack audio prediction network and parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data.

[0104] Among them, the operations of unit 710 - unit 720 and sub-units 721 - 727 in the training device 700 for the identity authentication model are similar to the operations of step S401 - step S408 in the above-mentioned training method for the identity authentication model, and will not be elaborated here.

[0105] In some embodiments, the number of at least one pre-trained first feature extraction network is at least two, and the third input subunit includes: a second input module configured to input at least one first feature into a feature fusion network to obtain a first fusion feature vector; and a third input module configured to input the first fusion feature vector into an attack audio prediction network to obtain an attack audio prediction result of the test audio; and wherein, the adjustment subunit is further configured to adjust the parameters of the feature fusion network, the parameters of the attack audio prediction network, and the parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data.

[0106] According to an embodiment of the present disclosure, there is also provided an electronic device, a readable storage medium, and a computer program product.

[0107] Reference Figure 8 , the structural block diagram of an electronic device 800 that can be used as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0108] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0109] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device capable of inputting information into the electronic device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include, but are not limited to, a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 807 can be any type of device capable of presenting information, and can include, but are not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but are not limited to, magnetic disks and optical discs. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but are not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0110] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the above authentication method or the training method of the above authentication model. For example, in some embodiments, the above authentication method or the training method of the above authentication model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the above authentication method or the training method of the above authentication model described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the above authentication method or the training method of the above authentication model in any other suitable manner (e.g., by means of firmware).

[0111] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0112] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0113] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0114] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0115] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0116] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0117] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0118] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the claims after authorization and their equivalent scope. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, the steps may be executed in an order different from that described in the present disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that emerge after the present disclosure.

Claims

1. An identity authentication method, comprising: Obtaining the acoustic features of the audio to be authenticated; Based on the acoustic features, performing attack audio recognition on the audio to be authenticated to obtain an attack audio prediction result of the audio to be authenticated, including: Inputting the acoustic features into a first feature extraction network to obtain first features output by the first feature extraction network; and Inputting the first features into an attack audio prediction network to obtain an attack audio prediction result output by the attack audio prediction network; Based on the acoustic features, obtaining the similarity of voiceprint features between the audio to be authenticated and the registered audio, where the similarity of voiceprint features includes the similarity between the second features of the audio to be authenticated and the voiceprint features of the registered audio, and the second features are obtained by performing feature extraction on the acoustic features through a second feature extraction network; and Based on the attack audio prediction result and the similarity of voiceprint features, performing identity prediction to obtain an identity authentication result, including: Inputting the attack audio prediction result and the similarity of voiceprint features into an identity prediction network to obtain an identity authentication result output by the identity prediction network.

2. The method according to claim 1, wherein, The performing attack audio recognition on the audio to be authenticated based on the acoustic features to obtain an attack audio prediction result of the audio to be authenticated includes: Based on the acoustic features, respectively obtaining at least one first feature through at least one first feature extraction network, where each of the at least one first features is used to determine whether the audio to be authenticated is an attack audio, and each of the at least one first feature extraction networks is constructed based on a neural network with a different structure; and Based on the at least one first feature, obtaining the attack audio prediction result.

3. The method according to claim 2, wherein, The number of the at least one first feature extraction networks is at least two, and the obtaining the attack audio prediction result based on the at least one first feature includes: Obtaining a first fusion feature vector of the at least one first feature through a feature fusion network based on an attention mechanism; and Inputting the first fusion feature vector into an attack audio prediction network to obtain the attack audio prediction result.

4. The method according to any one of claims 1-3, wherein, The obtaining the similarity of voiceprint features between the audio to be authenticated and the registered audio based on the acoustic features includes: Based on the acoustic features, respectively obtaining at least one second feature through at least one second feature extraction network, where each of the at least one second features is used for identity authentication, and each of the at least one second feature extraction networks is constructed based on a neural network with a different structure; and For each of the at least one second features, respectively obtaining the similarity of voiceprint features between the second feature and the voiceprint features of the registered audio.

5. The method according to claim 4, wherein, The performing identity prediction based on the attack audio prediction result and the similarity of voiceprint features to obtain an identity authentication result includes: Generating a second fusion feature vector based on the attack audio prediction result and at least one similarity of voiceprint features respectively corresponding to the at least one second feature; and Perform identity prediction based on the second fusion feature vector to obtain the identity authentication result of the audio to be authenticated.

6. A training method for an identity authentication model, comprising: Obtain a first sample data set, wherein each sample data in the first sample data set includes a test audio, a registered audio, and a corresponding identity authentication label; For each sample data in the first sample data set, perform the following operations: Obtain a first acoustic feature corresponding to the test audio and a second acoustic feature corresponding to the registered audio in the sample data; Input the first acoustic feature into at least one pre-trained first feature extraction network respectively to obtain at least one first feature, wherein each first feature in the at least one first feature is used to determine whether the audio to be authenticated is an attack audio, and each first feature extraction network in the at least one first feature extraction network is constructed based on a neural network with a different structure; Input the first acoustic feature and the second acoustic feature together into at least one pre-trained second feature extraction network respectively to obtain at least one test audio feature corresponding to the first acoustic feature and a registered audio feature corresponding to each test audio feature in the at least one test audio feature, wherein each second feature extraction network in the at least one second feature extraction network is constructed based on a neural network with a different structure; Input the at least one first feature into an attack audio prediction network to obtain the attack audio prediction result of the test audio; Obtain the feature similarity between each test audio feature in the at least one test audio feature and the registered audio feature corresponding to the test audio feature; Input the attack audio prediction result and at least one feature similarity corresponding to the at least one test audio feature into an identity prediction network to obtain an identity prediction result; and Adjust the parameters of the attack audio prediction network and the parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data.

7. The method according to claim 6, wherein, The number of the at least one pre-trained first feature extraction network is at least two, and the inputting the at least one first feature into an attack audio prediction network to obtain the attack audio prediction result of the test audio includes: Input the at least one first feature into a feature fusion network to obtain a first fusion feature vector; and Input the first fusion feature vector into the attack audio prediction network to obtain the attack audio prediction result of the test audio; and wherein, The adjusting the parameters of the attack audio prediction network and the parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data includes: Adjust the parameters of the feature fusion network, the parameters of the attack audio prediction network, and the parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data.

8. An identity authentication device, comprising: A first acquisition unit configured to acquire the acoustic feature of the audio to be authenticated; An identification unit, configured to perform attack audio identification on the audio to be authenticated based on the acoustic features, so as to obtain an attack audio prediction result of the audio to be authenticated, including: Inputting the acoustic features into a first feature extraction network to obtain first features output by the first feature extraction network; and Inputting the first features into an attack audio prediction network to obtain an attack audio prediction result output by the attack audio prediction network; A second acquisition unit, configured to obtain a voiceprint feature similarity between the audio to be authenticated and a registered audio based on the acoustic features, where the voiceprint feature similarity includes a similarity between second features of the audio to be authenticated and the voiceprint features of the registered audio, and the second features are obtained by performing feature extraction on the acoustic features through a second feature extraction network; and A prediction unit, configured to perform identity prediction based on the attack audio prediction result and the voiceprint feature similarity to obtain an identity authentication result, including: Inputting the attack audio prediction result and the voiceprint feature similarity into an identity prediction network to obtain an identity authentication result output by the identity prediction network.

9. The device according to claim 8, wherein The identification unit includes: A first acquisition subunit, configured to respectively obtain at least one first feature through at least one first feature extraction network based on the acoustic features, where each of the at least one first features is used to determine whether the audio to be authenticated is attack audio, and each of the at least one first feature extraction networks is constructed based on a neural network with a different structure; and A second acquisition subunit, configured to obtain the attack audio prediction result based on the at least one first feature.

10. The device according to claim 9, wherein, The number of the at least one first feature extraction networks is at least two, and the second acquisition subunit includes: An acquisition module, configured to obtain a first fusion feature vector of the at least one first feature through a feature fusion network based on an attention mechanism; and A first input module, configured to input the first fusion feature vector into an attack audio prediction network to obtain the attack audio prediction result.

11. The device according to any one of claims 8 - 10, wherein, The second acquisition unit includes: A third acquisition subunit, configured to respectively obtain at least one second feature through at least one second feature extraction network based on the acoustic features, where each of the at least one second features is used for identity authentication, and each of the at least one second feature extraction networks is constructed based on a neural network with a different structure; and A fourth acquisition subunit, configured to respectively obtain a voiceprint feature similarity between each of the at least one second features and the voiceprint features of the registered audio.

12. The apparatus according to claim 11, wherein, The prediction unit includes: A generation subunit, configured to generate a second fusion feature vector based on the attack audio prediction result and at least one voiceprint feature similarity corresponding to the at least one second feature; and A prediction subunit, configured to perform identity prediction based on the second fusion feature vector to obtain the identity authentication result of the audio to be authenticated.

13. A training device for an identity authentication model, comprising: A third acquisition unit, configured to acquire a first sample data set, wherein each sample data in the first sample data set includes a test audio, a registered audio, and a corresponding identity authentication label; An execution unit, configured to perform operations of the following sub-units for each sample data in the first sample data set: A fifth acquisition sub-unit, configured to acquire a first acoustic feature corresponding to the test audio and a second acoustic feature corresponding to the registered audio in the sample data; A first input sub-unit, configured to input the first acoustic feature into at least one pre-trained first feature extraction network respectively to obtain at least one first feature, wherein each first feature in the at least one first feature is used to determine whether the audio to be authenticated is an attack audio, and each first feature extraction network in the at least one first feature extraction network is constructed based on a neural network with a different structure; A second input sub-unit, configured to input the first acoustic feature and the second acoustic feature together into at least one pre-trained second feature extraction network respectively to obtain at least one test audio feature corresponding to the first acoustic feature and a registered audio feature corresponding to each test audio feature in the at least one test audio feature, wherein each second feature extraction network in the at least one second feature extraction network is constructed based on a neural network with a different structure; A third input sub-unit, configured to input the at least one first feature into an attack audio prediction network to obtain an attack audio prediction result of the test audio; A sixth acquisition sub-unit, configured to acquire a feature similarity between each test audio feature in the at least one test audio feature and the registered audio feature corresponding to the test audio feature; A fourth input sub-unit, configured to input the attack audio prediction result and at least one feature similarity corresponding to the at least one test audio feature into an identity prediction network to obtain an identity prediction result; and An adjustment sub-unit, configured to adjust parameters of the attack audio prediction network and parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data.

14. The apparatus according to claim 13, wherein, The number of the at least one pre-trained first feature extraction network is at least two, and the third input sub-unit includes: A second input module, configured to input the at least one first feature into a feature fusion network to obtain a first fusion feature vector; and A third input module, configured to input the first fusion feature vector into the attack audio prediction network to obtain an attack audio prediction result of the test audio; and wherein, The adjustment sub-unit is further configured to adjust parameters of the feature fusion network, parameters of the attack audio prediction network, and parameters of the identity prediction network based on the identity prediction result and the identity authentication label corresponding to the sample data.

15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to execute the method according to any one of claims 1-7.

17. A computer program product, comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Authentication model training method and device, and electronic equipment

    CN113035230A