Voiceprint feature extraction method, device, electronic device and storage medium

By generating the eigenvector of the posterior mean vector based on the Gaussian distribution, the problem of factor interference in the vocalprint feature extraction process is solved, and a higher extraction accuracy is achieved.

CN113889120BActive Publication Date: 2025-05-20BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111143893.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-28
Publication Date
2025-05-20
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

During the extraction of voiceprint features, it is affected by many factors such as the speaker himself and the environment, which affects the extraction accuracy.

Method used

By obtaining the initial vocalprint feature data, the initial eigenvector is generated, and the covariance matrix is ​​used to generate a posterior mean vector according to the Gaussian distribution as an updated eigenvector to improve the accuracy of vocalprint feature extraction.

Benefits of technology

This method can more accurately extract the voiceprint characteristics of the speaker, improve the accuracy of voiceprint characteristics extraction, and reduce the influence of factor interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113889120B_ABST
    Figure CN113889120B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voiceprint feature extraction method, device, electronic device and storage medium, which relate to the field of artificial intelligence technology, and in particular to speech recognition technology. The implementation scheme is as follows: a voiceprint feature extraction method includes: obtaining initial voiceprint feature data about a speaker; generating an initial feature vector of the speaker based on the initial voiceprint feature data; generating a covariance matrix corresponding to the initial feature vector; generating an updated feature vector of the speaker based on the initial feature vector and the covariance matrix, wherein the updated feature vector is a posterior mean vector of the initial feature vector according to a Gaussian distribution; and extracting the voiceprint feature of the speaker based on the updated feature vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to speech recognition technology, and specifically relates to a method, apparatus, electronic device, and storage medium for extracting voiceprint features. Background Art

[0002] Artificial intelligence is a discipline that studies enabling computers to simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0003] In speech recognition technology, voiceprint feature extraction is crucial. Voiceprint feature extraction generally can include various modules for extracting front-end features (also called low-level features), for extracting speaker features (also called high-level features), and for back-end classification. After the speech data is processed by the above modules, corresponding recognition results can be obtained.

[0004] However, in the process of voiceprint feature extraction, it is often affected by various factors such as the speaker himself and the environment, which has a certain impact on the accuracy of voiceprint feature extraction. How to eliminate these influences to more accurately extract the speaker's voiceprint features has become a hot research field in voiceprint feature extraction technology in recent years.

[0005] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, any method described in this section should not be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0006] The present disclosure provides a method, apparatus, electronic device, and storage medium for extracting voiceprint features.

[0007] According to one aspect of the present disclosure, there is provided a method for extracting voiceprint features, including: obtaining initial voiceprint feature data about a speaker; generating an initial feature vector of the speaker based on the initial voiceprint feature data; generating a covariance matrix corresponding to the initial feature vector; generating an updated feature vector of the speaker based on the initial feature vector and the covariance matrix, where the updated feature vector is the posterior mean vector of the initial feature vector according to the Gaussian distribution; and extracting the voiceprint feature of the speaker based on the updated feature vector.

[0008] According to another aspect of the present disclosure, a method for training a voiceprint feature extraction model is provided, including: providing sample initial voiceprint feature data about a predetermined speaker; generating a sample initial feature vector of the predetermined speaker based on the sample initial voiceprint feature data; generating a sample covariance matrix corresponding to the sample initial feature vector; generating an updated sample feature vector of the predetermined speaker based on the sample initial feature vector and the sample covariance matrix, wherein the updated sample feature vector is the posterior mean vector of the sample initial feature vector according to a Gaussian distribution; extracting the voiceprint feature of the predetermined speaker based on the updated sample feature vector; and obtaining network parameters for updating the voiceprint feature extraction model based on the voiceprint feature to train the voiceprint feature extraction model.

[0009] According to another aspect of the present disclosure, a voiceprint feature extraction device is provided, including: an acquisition unit configured to acquire initial voiceprint feature data about a speaker; a first generation unit configured to generate an initial feature vector of the speaker based on the initial voiceprint feature data; a second generation unit configured to generate a covariance matrix corresponding to the initial feature vector; a third generation unit configured to generate an updated feature vector of the speaker based on the initial feature vector and the covariance matrix, wherein the updated feature vector is the posterior mean vector of the initial feature vector according to a Gaussian distribution; and an extraction unit configured to extract the voiceprint feature of the speaker based on the updated feature vector.

[0010] According to another aspect of the present disclosure, a device for training a voiceprint feature extraction model is provided, including: a providing unit configured to provide sample initial voiceprint feature data about a predetermined speaker; a first sample generation unit configured to generate a sample initial feature vector of the predetermined speaker based on the sample initial voiceprint feature data; a second sample generation unit configured to generate a sample covariance matrix corresponding to the sample initial feature vector; a third sample generation unit configured to generate an updated sample feature vector of the predetermined speaker based on the sample initial feature vector and the sample covariance matrix, wherein the updated sample feature vector is the posterior mean vector of the sample initial feature vector according to a Gaussian distribution; a sample extraction unit configured to extract the voiceprint feature of the predetermined speaker based on the updated sample feature vector; and a network parameter acquisition unit configured to obtain network parameters for updating the voiceprint feature extraction model based on the voiceprint feature to train the voiceprint feature extraction model.

[0011] According to another aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is caused to execute the method as described above.

[0012] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.

[0013] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the method as described above.

[0014] According to one or more embodiments of the present disclosure, the voiceprint features of the speaker can be accurately extracted.

[0015] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0017] Figure 1 A schematic diagram of an exemplary system in which the various methods and apparatuses described herein can be implemented according to embodiments of the present disclosure is shown.

[0018] Figure 2 A flowchart of a voiceprint feature extraction method according to an embodiment of the present disclosure is shown.

[0019] Figure 3 A flowchart of a method for training a voiceprint feature extraction model according to an embodiment of the present disclosure is shown.

[0020] Figure 4 A schematic diagram of a voiceprint extraction model according to an embodiment of the present disclosure is shown.

[0021] Figure 5 A block diagram of a voiceprint feature extraction apparatus according to an embodiment of the present disclosure is shown.

[0022] Figure 6 A block diagram of a voiceprint feature extraction apparatus according to another embodiment of the present disclosure is shown.

[0023] Figure 7 A block diagram of a device for training a voiceprint feature extraction model according to an embodiment of the present disclosure is shown.

[0024] Figure 8A block diagram of an electronic device that can be applied to embodiments of the present disclosure is shown. Detailed implementation manners

[0025] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0026] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0027] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.

[0028] In the related art, in order to extract the features of a speaker, the methods generally used may include an unsupervised Gaussian mixture model (GMM), as well as supervised deep neural networks (DNN), convolutional neural networks (CNN), etc. The effects of speaker feature extraction that can be brought by the above methods are relatively limited, and there are bottlenecks in improving the accuracy of speaker feature extraction.

[0029] In view of the above problems, according to one aspect of the present disclosure, a method for extracting speaker features is provided. Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0030] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure is shown. Refer to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.

[0031] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable the execution of the voiceprint feature extraction method of the present disclosure.

[0032] In some embodiments, the server 120 may also provide other services or software applications that may include non-virtual environments and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.

[0033] In Figure 1 In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0034] Users may use client devices 101, 102, 103, 104, 105, and / or 106 to generate voice data for the voiceprint feature extraction method of the present disclosure. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.

[0035] Client devices 101, 102, 103, 104, 105, and / or 106 can include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices can run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android. Portable handheld devices can include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices can include head-mounted displays and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0036] Network 110 can be any type of network known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, public switched telephone network (PSTN), infrared network, wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0037] Server 120 can include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices of the server). In various embodiments, server 120 can run one or more services or software applications that provide the functions described below.

[0038] The computing unit in server 120 can run one or more operating systems including any of the above-mentioned operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0039] In some embodiments, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0040] In some embodiments, server 120 can be a server of a distributed system or a server incorporating a blockchain. Server 120 can also be a cloud server or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to address the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0041] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120 or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.

[0042] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.

[0043] Figure 1 System 100 can be configured and operated in various ways to enable the application of the various methods and apparatuses described in this disclosure.

[0044] Figure 2 shows a flowchart of a voiceprint feature extraction method 200 according to an embodiment of the present disclosure. As Figure 2 shown, the method 200 may include the following steps:

[0045] S202, obtain initial voiceprint feature data about the speaker;

[0046] S204, generate an initial feature vector of the speaker based on the initial voiceprint feature data;

[0047] S206, generate a covariance matrix corresponding to the initial feature vector;

[0048] S208, based on the initial feature vector and the covariance matrix, generate an updated feature vector of the speaker, where the updated feature vector is the posterior mean vector of the initial feature vector according to the Gaussian distribution; and

[0049] S210, extract the voiceprint feature of the speaker based on the updated feature vector.

[0050] According to the voiceprint feature extraction method of the present disclosure, in order to more accurately extract the voiceprint feature of the speaker, based on the generated initial feature vector of the speaker, by using the covariance matrix that can characterize the correlation between the initial feature vectors in different feature dimensions (for example, the corresponding feature dimensions related to the physiological characteristics of the speaker's own oral cavity, vocal cords, etc., the corresponding feature dimensions related to the environment and background during speaking, etc.), the posterior mean vector of the initial feature vector according to the Gaussian distribution can be obtained. This posterior mean vector can contain richer voiceprint feature information, so that the voiceprint feature of the speaker can be extracted as much as possible and fully. Therefore, this posterior mean vector, as an updated feature vector based on the initial feature vector, can be used to accurately extract the voiceprint feature of the speaker, thereby improving the accuracy of voiceprint feature extraction.

[0051] The following will describe each step of the voiceprint feature extraction method according to the present disclosure in detail.

[0052] In step S202, the speaker may refer to any speaker whose voiceprint feature is to be extracted. Initial voiceprint feature data about the speaker can be obtained.

[0053] According to some embodiments, the initial voiceprint feature data may be extracted from audio data about the speaker, and the initial voiceprint feature data may include a plurality of sub-feature data with context relationships, where the plurality of sub-feature data may correspond to a plurality of frames of the audio data.

[0054] Specifically, the audio data of a certain speaker can be obtained and framed. For example, a segment of audio data of a certain speaker can be framed into T frames of audio data (T is a natural number greater than 1), and corresponding voiceprint feature data can be extracted from each frame of audio data in the T frames of audio data to obtain T voiceprint feature data. Correspondingly, the T voiceprint feature data generally correspond to the initial voiceprint feature data of the speaker, and each voiceprint feature data corresponds to a sub-feature data.

[0055] Here, in order to obtain information related to the temporal variation of the speaker's voiceprint features, it is considered to frame the audio data into multiple audio frames, that is, T is greater than or equal to 2. Correspondingly, there is a temporal context relationship between the T audio frames, so there is also a temporal context relationship between the feature data extracted from them, that is, the T sub-feature data. For example, the T sub-feature data can reflect the voiceprint features of the speaker at a continuous number of time points (i.e., T frames), and such voiceprint features may be different for different speakers, so they can become unique voiceprint features.

[0056] In this way, information related to the temporal variation of the speaker's voiceprint features can be incorporated into the voiceprint extraction process, so as to facilitate the extraction of more information of the voiceprint features, thereby improving the accuracy of voiceprint feature extraction.

[0057] After the audio data is obtained, various preprocessing operations can also be performed, including noise removal (such as environmental noise, busy tone, ringtone, etc.) and data augmentation (such as aliasing echo, changing the rate (such as speeding up or slowing down the speech rate), time-domain and frequency-domain random masking, etc.).

[0058] The initial voiceprint feature data can be extracted from the audio data through various known voiceprint feature extraction techniques, such as MFCC (Mel Frequency Cepstral Coefficients), Fbank (filter bank), PLP (Perceptual Linear Prediction), etc. In addition, feature mean normalization (mean subtraction) can also be performed.

[0059] It should be noted that in the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0060] According to some embodiments, in step S204, the initial voiceprint feature data can be input into the first neural network to obtain an initial feature vector, where the initial feature vector includes multiple first elements corresponding to multiple sub-feature data and having a temporal context relationship.

[0061] The first neural network can be various known types of neural networks, such as TDNN (Time Delay Neural Network). Since a neural network can fit any kind of distribution, it is not necessary to preset the distribution to be satisfied in advance, but rely on the learning ability of the neural network itself to obtain the corresponding calculation result. By means of multiple mappings of the neural network to the feature data, the accuracy of calculating the initial feature vector can be improved.

[0062] For example, the first neural network can include multiple layers of TDNN (with ReLU (Rectified Linear Unit) activation function) and one fully connected layer (with ReLU activation function). In practical applications, for the selection of the first neural network, an appropriate neural network can be selected by considering the trade-off between effect and efficiency.

[0063] Correspondingly, the initial feature vector extracted from the initial voiceprint feature data can include T first elements corresponding to T sub-feature data, and the T first elements can also have a temporal context relationship.

[0064] According to some embodiments, in step S206, the initial feature vector can be input into the second neural network to obtain a covariance matrix corresponding to the initial feature vector, where the covariance matrix has multiple second elements corresponding to the multiple first elements of the initial feature vector, and each second element characterizes the correlation between the corresponding first element in different feature dimensions.

[0065] The second neural network can also be various known types of neural networks. For example, the second neural network can include two layers of fully connected layers (with ReLU activation function). According to the actual situation, the number of layers of the second neural network can be appropriately selected, for example, 2 to 5 layers.

[0066] In this article, the second neural network is also called the auxiliary fully connected layer because it takes the initial feature vector as input and auxiliarily outputs the covariance matrix to prepare for the subsequent application of the normalized exponential function (Softmax).

[0067] Correspondingly, the covariance matrix can have T second elements corresponding to the T first elements of the initial feature vector. Since covariance can be used to represent the correlation between data in different dimensions, each second element in the covariance matrix can characterize the correlation between the corresponding first element in the initial feature vector in different feature dimensions. Here, the different feature dimensions of the initial feature vector can refer to, for example, the corresponding feature dimensions related to the physiological characteristics of the speaker's own oral cavity, vocal cords, etc., and the corresponding feature dimensions related to the environment and background during speaking. For example, each second element in the covariance matrix can characterize the correlation between the pronunciation feature on the oral cavity and the pronunciation feature on the vocal cords in the initial feature vector.

[0068] In this way, by using the covariance matrix of the initial feature vectors, connections can be established between different feature dimensions of the initial feature vectors, rather than being limited to the feature information on each individual feature dimension, so that more abundant feature information about the speaker can be obtained.

[0069] In one example, after obtaining the covariance matrix, subsequent calculations related to the covariance matrix can be performed in a logarithmic manner, aiming to convert multiplication calculations into addition calculations, thereby reducing the computational amount.

[0070] According to some embodiments, in step S208, the Softmax (normalized exponential function) can be applied to the covariance matrix to obtain the Softmax value of the covariance matrix; and the Softmax value can be multiplied by the initial feature vectors to obtain updated feature vectors.

[0071] Since Softmax is similar to a normalization operation, after applying Softmax to the covariance matrix, multiplying the obtained Softmax value by the initial feature vectors is equivalent to performing a weighted average. Therefore, the posterior mean vector of the initial feature vectors according to the Gaussian distribution can be obtained.

[0072] Here, the posterior mean vector can be understood as an accurate value obtained based on the actual voiceprint feature data of the speaker. In other words, under the assumption that the voiceprint features of the speaker conform to the Gaussian distribution, the prior probability alone cannot reflect the specific form of this Gaussian distribution. In the case of obtaining the posterior mean vector as described above, the voiceprint features of the speaker can be accurately expressed by the Gaussian distribution, thus facilitating the accurate extraction of the speaker's voiceprint features.

[0073] In one example, in order to obtain this posterior mean vector, the prior mean and covariance can also be initialized (for example, initialized to 0) and added to the above weighted average calculation.

[0074] In addition, in the case where the initial voiceprint feature data includes multiple sub-feature data with a temporal context relationship, the above process of obtaining updated feature vectors can represent the temporal aggregation of multiple first elements with a temporal context relationship in the initial feature vectors, so as to facilitate obtaining more abundant feature information about the speaker.

[0075] According to some embodiments, in step S210, the updated feature vectors can be mapped through an embedding operation to generate the voiceprint features of the speaker. In this way, the voiceprint features of the speaker can be accurately extracted based on the updated feature vectors containing more feature information.

[0076] For example, the updated feature vector can be input into two fully-connected layers (with ReLU activation function) for mapping to generate the voiceprint feature of the speaker.

[0077] As described above, according to the voiceprint feature extraction method of the present disclosure, in order to more accurately extract the voiceprint feature of the speaker, based on the generated initial feature vector of the speaker, the covariance matrix capable of characterizing the correlation between the initial feature vectors in different feature dimensions (for example, the corresponding feature dimensions related to the physiological characteristics of the speaker's own oral cavity, vocal cords, etc., the corresponding feature dimensions related to the environment and background during speaking, etc.) can be used to obtain the posterior mean vector of the initial feature vector according to the Gaussian distribution. This posterior mean vector can contain richer voiceprint feature information, so that the voiceprint feature of the speaker can be extracted as much as possible and fully. Therefore, this posterior mean vector, as the updated feature vector based on the initial feature vector, can be used to accurately extract the voiceprint feature of the speaker, thereby improving the accuracy of voiceprint feature extraction.

[0078] According to an embodiment of the present disclosure, there is also provided a method for training a voiceprint feature extraction model.

[0079] Figure 3 FIG. 300 shows a flowchart of a method 300 for training a voiceprint feature extraction model according to an embodiment of the present disclosure.

[0080] As Figure 3 shown, the method 300 may include the following steps:

[0081] S302, providing sample initial voiceprint feature data about a predetermined speaker;

[0082] S304, generating a sample initial feature vector of the predetermined speaker based on the sample initial voiceprint feature data;

[0083] S306, generating a sample covariance matrix corresponding to the sample initial feature vector;

[0084] S308, generating an updated sample feature vector of the predetermined speaker based on the sample initial feature vector and the sample covariance matrix, wherein the updated sample feature vector is the posterior mean vector of the sample initial feature vector according to the Gaussian distribution;

[0085] S310, extracting the voiceprint feature of the predetermined speaker based on the updated sample feature vector; and

[0086] S312, obtaining network parameters for updating the voiceprint feature extraction model based on the voiceprint feature to train the voiceprint feature extraction model.

[0087] It can be noted that the operations of steps S304 to S310 in method 300 for training the voiceprint feature extraction model are similar to the operations of steps S204 to S210 in the voiceprint feature extraction method 200 described in combination with Figure 2 and are all used to extract the voiceprint features of the speaker.

[0088] The difference is that in method 300 for training the voiceprint feature extraction model, the data used for training in step S302 is for a predetermined speaker, that is, the sample initial voiceprint feature data, so that the model can be trained for a specific speaker. Accordingly, the speaker annotation needs to be performed on the audio data so as to be able to determine which specific speaker the audio data is for.

[0089] In addition, method 300 for training the voiceprint feature extraction model further includes updating the network parameters of the voiceprint feature extraction model in step S312. In other words, the voiceprint features extracted in step S310 are used to update the network parameters, so that multiple rounds of training can be iteratively performed until the network converges, thereby completing the training of the model. Accordingly, the voiceprint features extracted in step S310 can be input into, for example, a two-layer fully connected layer (with ReLU activation function) and then input into the output layer, calculate the loss through the CE (cross entropy) criterion, and finally update the network parameters in reverse according to SGD (stochastic gradient descent), and iterate multiple rounds until the network converges to complete the model training.

[0090] Figure 4 FIG. shows a schematic diagram of a voiceprint extraction model 400 according to an embodiment of the present disclosure. The following will be combined with Figures 2 to 4 to further illustrate the voiceprint feature extraction method according to the present disclosure and the method for training the voiceprint feature extraction model.

[0091] As Figure 4 shown, the voiceprint extraction model 400 may include an encoder 410, a Gaussian posterior inference module 420, and a decoder 430.

[0092] The encoder 410 may include a TDNN layer 411, a fully connected layer 412, and an auxiliary fully connected layer 413. Figure 4 It is schematically shown that the auxiliary fully connected layer 413 includes a first auxiliary fully connected layer 413-1 and a second auxiliary fully connected layer 413-2. However, the number of layers of the auxiliary fully connected layer 413 can be appropriately selected according to the actual situation. The TDNN layer 411 can also be a neural network including multiple layers, which can be appropriately selected according to the actual situation.

[0093] The TDNN layer 411 and the fully connected layer 412 together can correspond to the first neural network described in combination with Figure 2 and the auxiliary fully connected layer 413 can correspond to the combination described in Figure 2The second neural network described above. As Figure 4 shown, the initial voiceprint feature data X extracted from the audio data about the speaker can be input into the voiceprint extraction model 400. Among them, the initial voiceprint feature data X can include multiple sub-feature data x 1 , x 2 , … x T , and the multiple sub-feature data x 1 , x 2 , … x T can correspond to T frames after the audio data is framed.

[0094] It should be noted that during the training stage of the voiceprint extraction model 400, since the audio data is labeled with the speaker, the initial voiceprint feature data X is the sample initial voiceprint feature data corresponding to a specific speaker. After the training of the voiceprint extraction model 400 is completed, when using the voiceprint extraction model 400 for voiceprint feature extraction, the audio data can be about any speaker, so the initial voiceprint feature data X is also about any speaker.

[0095] As Figure 4 shown, the initial voiceprint feature data X = {x 1 , x 2 , …, x T} obtains the initial feature vector {z 1 , z 2 , …, z T} after passing through the TDNN layer 411 and the fully connected layer 412 (the first neural network). In addition, the initial feature vector {z 1 , z 2 , …, z T} obtains the covariance matrix log{L 1 , L 2 , …, L T} corresponding to the initial feature vector {z 1 , L 2 , …, L T} after passing through the auxiliary fully connected layer 413 (as described above, for the convenience of calculation, the logarithm method is used for the calculation of the covariance matrix).

[0096] Each first element z 1 , z 2 , …, z T} in the initial feature vector {z 1 , z 2 , …, z T} can contain feature information on different feature dimensions. For example, z 1 , z 2 , …, z TEach of them (e.g., taking z 1 as an example) may include, for example, feature information on corresponding feature dimensions related to the physiological characteristics of the speaker's oral cavity, vocal cords, etc. Therefore, the corresponding second elements {L 1 , L 2 , …, L T} in the covariance matrix (e.g., correspondingly taking L 1 as an example) can characterize, for example, the correlation between the pronunciation features on the oral cavity and the pronunciation features on the vocal cords.

[0097] The operation of the Gaussian posterior inference module 420 can correspond to generating an updated feature vector of the speaker as described above in combination with Figure 2 , Figure 3 , such as φs shown in Figure 4 . In the Gaussian posterior inference module 420, the covariance matrix log{L 1 , L 2 , …, L T} is Softmaxed, and then multiplied by the initial feature vector {z 1 , z 2 , …, z T} to obtain the updated feature vector φs. By multiplying the Softmax value of the covariance matrix log{L 1 , L 2 , …, L T} by the initial feature vector {z 1 , z 2 , …, z T}, that is, a weighted average operation is performed. Therefore, the obtained updated feature vector φs is the posterior mean vector of the initial feature vector according to the Gaussian distribution, which is a more accurate expression of the speaker's voiceprint features in the form of a Gaussian distribution.

[0098] In one example, the prior mean and covariance (μ p , L p ) can also be initialized to 0 and added to the above weighted average calculation, that is, L 0 = 0, z 0 = 0.

[0099] Meanwhile, when the initial voiceprint feature data X = {x 1 , x 2 , …, x T} includes multiple sub-feature data {x 1 , x 2 , …, x T} with a temporal context relationship, the above process of obtaining the updated feature vector φs can represent the operation of weighted averaging the initial feature vector {z 1 , z 2 , …, zT Multiple first elements z with temporal context in {...} 1 , z 2 , …, z T are temporally aggregated, so it can also be called temporal aggregation.

[0100] Note that during the training phase of the voiceprint extraction model 400, since the input initial voiceprint feature data X is sample initial voiceprint feature data about a specific speaker, correspondingly, the obtained initial feature vectors {z 1 , z 2 , …, z T}, covariance matrices log{L 1 , L 2 , …, L T} and updated feature vectors φs are also the sample initial feature vectors, sample covariance matrices, and updated sample feature vectors for this specific speaker, respectively.

[0101] The decoder 430 may include an embedding layer 431, a fully connected layer 432, and an output layer 433. The embedding layer 431 may map the updated feature vector φs to generate the voiceprint feature of the speaker.

[0102] Note that during the training phase of the voiceprint extraction model 400, it may further include a fully connected layer 432 and an output layer 433. The voiceprint feature extracted by the embedding layer 431 can be input into the fully connected layer 432 and then into the output layer 433 (for speaker classification), calculate the loss through the CE (cross entropy) criterion, and finally update the network parameters backward according to SGD (stochastic gradient descent), and iterate multiple rounds until the network converges. During the test phase, the voiceprint feature can be extracted by the embedding layer 431, and then calculate the pairwise similarity through cosine (cos).

[0103] The above has described the voiceprint feature extraction method of the present disclosure and the method for training the voiceprint feature extraction model in combination with Figures 2 to 4 . To more fully understand the present disclosure, the following briefly describes the model assumptions based on the principles of the present disclosure.

[0104] First, make assumptions about the prior distribution of the speaker's features:

[0105] 1) Model: z t = h + ∈ t

[0106] 2) Latent variable:

[0107] 3) Uncertainty:

[0108] Among them, zt is a variable generated for the model, representing the feature vector of the speaker; h is a latent variable, representing the base vector of the speaker; ∈ t is a residual variable, representing the difference vector of different speakers; μ and L are the mean and covariance in the Gaussian distribution, and z t , h, and ∈ t all follow Gaussian distributions.

[0109] The base vector h can reflect the average value of the population in several characteristic attributes (such as several characteristic dimensions related to the oral cavity, vocal cords, tongue, lips, etc.). The residual variable ∈ t can reflect the perturbation vector of the change in some of these characteristic attributes (such as the vibration frequency of the vocal cords). Based on this model, the characteristics of any speaker can be characterized.

[0110] Since the base vector h follows a Gaussian distribution, the posterior probability of the base vector h can be calculated for a specific population (such as 1000 Asians), thereby reflecting the specific form of the Gaussian distribution of this population with respect to the base vector h. That is, the initial feature vector obtained in the present disclosure is the posterior mean vector according to the Gaussian distribution.

[0111] After giving the input data about a specific population, the posterior probability distribution of the base vector h can be derived as:

[0112]

[0113] where:

[0114]

[0115]

[0116] Here, can be obtained through a neural network and where f enc () and g enc () can respectively correspond to the first neural network and the second neural network as described above.

[0117] In addition, the training process of the voiceprint extraction model can also be represented by the following pseudocode:

[0118]

[0119] Among them, θ can represent network parameters to be updated, such as weights and biases. θ can be calculated by multiplying the parameter g by the learning force, where the parameter g can reflect the amount of change or update learned from a batch of data each time, and the learning force can represent the magnitude of change. After obtaining the parameter g, it can be added to θ (such as weights and biases) to produce corresponding changes, thereby implementing the training process.

[0120] According to another aspect of the present disclosure, a voiceprint feature extraction device is also provided. Figure 5 The block diagram of a voiceprint feature extraction device 500 according to an embodiment of the present disclosure is shown.

[0121] As Figure 5 shown, the device 500 may include: an acquisition unit 502 configured to acquire initial voiceprint feature data about a speaker; a first generation unit 504 configured to generate an initial feature vector of the speaker based on the initial voiceprint feature data; a second generation unit 506 configured to generate a covariance matrix corresponding to the initial feature vector; a third generation unit 508 configured to generate an updated feature vector of the speaker based on the initial feature vector and the covariance matrix, where the updated feature vector is the posterior mean vector of the initial feature vector according to the Gaussian distribution; and an extraction unit 510 configured to extract the voiceprint feature of the speaker based on the updated feature vector.

[0122] According to some embodiments, the initial voiceprint feature data is extracted from audio data about the speaker, and the initial voiceprint feature data includes a plurality of sub-feature data having a temporal context relationship, where the plurality of sub-feature data corresponds to a plurality of frames of the audio data.

[0123] The operations performed by the above modules 502 to 510 correspond to steps S202 to S210 described with reference to Figure 2 Therefore, details of each aspect will not be elaborated herein.

[0124] Figure 6 The block diagram of a voiceprint feature extraction device 600 according to another embodiment of the present disclosure is shown. Figure 6 The modules 602 to 610 shown may respectively correspond to Figure 5 the modules 502 to 510 shown. In addition, the modules 604 to 610 may further include further sub-functional modules, which will be specifically described below.

[0125] According to some embodiments, the first generation unit 604 may include: a first sub-unit 6040 configured to input the initial voiceprint feature data into a first neural network to obtain an initial feature vector, where the initial feature vector includes a plurality of first elements having a temporal context relationship and corresponding to the plurality of sub-feature data.

[0126] According to some embodiments, the second generation unit 606 may include: a second sub-unit 6060 configured to input an initial feature vector into a second neural network to obtain a covariance matrix corresponding to the initial feature vector, wherein the covariance matrix has a plurality of second elements corresponding to a plurality of first elements of the initial feature vector, and each second element characterizes the correlation between the corresponding first element in different feature dimensions.

[0127] According to some embodiments, the third generation unit 608 may include: a third sub-unit 6080 configured to apply a normalized exponential function to the covariance matrix to obtain a normalized exponential value of the covariance matrix; and a fourth sub-unit 6082 configured to multiply the normalized exponential value by the initial feature vector to obtain an updated feature vector.

[0128] According to some embodiments, the extraction unit 610 may include: a mapping unit 6100 configured to map the updated feature vector through an embedding operation to generate a speaker's voiceprint feature.

[0129] According to another aspect of the present disclosure, there is also provided an apparatus for training a voiceprint feature extraction model. Figure 7 A block diagram of an apparatus 700 for training a voiceprint feature extraction model according to an embodiment of the present disclosure is shown.

[0130] As Figure 7 shown, the apparatus 700 may include: a providing unit 702 configured to provide sample initial voiceprint feature data about a predetermined speaker; a first sample generation unit 704 configured to generate a sample initial feature vector of the predetermined speaker based on the sample initial voiceprint feature data; a second sample generation unit 706 configured to generate a sample covariance matrix corresponding to the sample initial feature vector; a third sample generation unit 708 configured to generate an updated sample feature vector of the predetermined speaker based on the sample initial feature vector and the sample covariance matrix, wherein the updated sample feature vector is a posterior mean vector of the sample initial feature vector according to a Gaussian distribution; a sample extraction unit 710 configured to extract a voiceprint feature of the predetermined speaker based on the updated sample feature vector; and a network parameter acquisition unit 712 configured to obtain network parameters for updating the voiceprint feature extraction model based on the voiceprint feature to train the voiceprint feature extraction model.

[0131] According to another aspect of the present disclosure, there is also provided an electronic device including at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor executes the method as described above.

[0132] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0133] According to another aspect of the present disclosure, there is also provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the method as described above.

[0134] Reference Figure 8 , a block diagram of an electronic device 800 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0135] As Figure 8 shown, the device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0136] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, output unit 807, storage unit 808, and communication unit 809. The input unit 806 can be any type of device capable of inputting information to device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone, and / or remote control. The output unit 807 can be any type of device capable of presenting information, and can include but are not limited to a display, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 808 can include but are not limited to magnetic disks, optical discs. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, network card, infrared communication device, wireless communication transceiver, and / or chipset, such as Bluetooth TM devices, 1302.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0137] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the voiceprint feature extraction method. For example, in some embodiments, the voiceprint feature extraction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the voiceprint feature extraction method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the voiceprint feature extraction method in any other suitable manner (e.g., by means of firmware).

[0138] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0139] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0140] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0141] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0142] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0143] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0144] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.

[0145] In the technical solutions of this disclosure, the acquisition, storage, and application of the user's personal information all comply with the provisions of relevant laws and regulations and do not violate public order and good customs. The intention of this disclosure is to manage and process personal information data in a manner that minimizes the risk of inadvertent or unauthorized use and access. The risk is minimized by restricting data collection and deleting the data when it is no longer needed. It should be noted that all information related to personnel in this application is collected with the knowledge and consent of the personnel.

[0146] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, the steps may be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that emerge after the present disclosure.

Claims

1. A voiceprint feature extraction method, comprising: Acquire initial voiceprint feature data about a speaker, wherein the initial voiceprint feature data is extracted from audio data about the speaker, and the initial voiceprint feature data includes a plurality of sub-feature data having a temporal contextual relationship, wherein the plurality of sub-feature data corresponds to a plurality of frames of the audio data; Generating an initial feature vector of the speaker based on the initial voiceprint feature data, comprising: inputting the initial voiceprint feature data into a first neural network to obtain the initial feature vector, wherein the initial feature vector comprises a plurality of first elements having the temporal contextual relationship and corresponding to the plurality of sub-feature data; Generate a covariance matrix corresponding to the initial eigenvector, which includes: inputting the initial eigenvector into a second neural network to obtain the covariance matrix corresponding to the initial eigenvector, wherein the covariance matrix has a plurality of second elements corresponding to a plurality of first elements of the initial eigenvector, and each second element represents the correlation between the corresponding first element and different feature dimensions; Based on the initial feature vector and the covariance matrix, generating an updated feature vector of the speaker, wherein the updated feature vector is a posterior mean vector of the initial feature vector according to a Gaussian distribution; and The voiceprint feature of the speaker is extracted based on the updated feature vector.

2. The method according to claim 1, wherein: Generating an updated feature vector of the speaker includes: applying a normalized exponential function to the covariance matrix to obtain a normalized exponential value for the covariance matrix; and The normalized exponent value is multiplied by the initial eigenvector to obtain the updated eigenvector.

3. The method according to claim 1 or 2, wherein: The extracting the voiceprint feature of the speaker based on the updated feature vector includes: mapping the updated feature vector through an embedding operation to generate the voiceprint feature of the speaker.

4. A method for training a voiceprint feature extraction model, comprising: Providing sample initial voiceprint feature data about a predetermined speaker, wherein the sample initial voiceprint feature data is extracted from audio data about the speaker, and the sample initial voiceprint feature data includes a plurality of sub-feature data having a temporal contextual relationship, wherein the plurality of sub-feature data corresponds to a plurality of frames of the audio data; Generating a sample initial feature vector of the predetermined speaker based on the sample initial voiceprint feature data, comprising: inputting the sample initial voiceprint feature data into a first neural network to obtain the sample initial feature vector, wherein the sample initial feature vector comprises a plurality of first elements having the temporal contextual relationship and corresponding to the plurality of sub-feature data; Generate a sample covariance matrix corresponding to the sample initial feature vector, which includes: inputting the sample initial feature vector into a second neural network to obtain the sample covariance matrix corresponding to the sample initial feature vector, wherein the sample covariance matrix has a plurality of second elements corresponding to a plurality of first elements of the sample initial feature vector, and each second element represents the correlation between the corresponding first element and different feature dimensions; Based on the sample initial feature vector and the sample covariance matrix, generating an updated sample feature vector of the predetermined speaker, wherein the updated sample feature vector is a posterior mean vector of the sample initial feature vector according to a Gaussian distribution; extracting the voiceprint feature of the predetermined speaker based on the updated sample feature vector; and Based on the voiceprint features, network parameters for updating a voiceprint feature extraction model are acquired to train the voiceprint feature extraction model.

5. A voiceprint feature extraction device, comprising: an acquiring unit, configured to acquire initial voiceprint feature data about a speaker, wherein the initial voiceprint feature data is extracted from audio data about the speaker, and the initial voiceprint feature data includes a plurality of sub-feature data having a temporal contextual relationship, wherein the plurality of sub-feature data corresponds to a plurality of frames of the audio data; A first generating unit is configured to generate an initial feature vector of the speaker based on the initial voiceprint feature data, wherein the first generating unit includes: A first subunit is configured to input the initial voiceprint feature data into a first neural network to obtain the initial feature vector, wherein the initial feature vector includes a plurality of first elements having the temporal context relationship and corresponding to the plurality of sub-feature data; A second generating unit is configured to generate a covariance matrix corresponding to the initial eigenvector, wherein the second generating unit comprises: A second subunit is configured to input the initial feature vector into a second neural network to obtain the covariance matrix corresponding to the initial feature vector, wherein the covariance matrix has a plurality of second elements corresponding to a plurality of first elements of the initial feature vector, and each second element represents a correlation between the corresponding first element and different feature dimensions; A third generating unit is configured to generate an updated feature vector of the speaker based on the initial feature vector and the covariance matrix, wherein the updated feature vector is a posterior mean vector of the initial feature vector according to a Gaussian distribution; and The extraction unit is configured to extract the voiceprint feature of the speaker based on the updated feature vector.

6. The device according to claim 5, wherein: The third generating unit comprises: A third subunit is configured to apply a normalized index function to the covariance matrix to obtain a normalized index value with respect to the covariance matrix; and The fourth subunit is configured to multiply the normalized index value by the initial feature vector to obtain the updated feature vector.

7. The device according to claim 5 or 6, wherein: The extraction unit comprises: A mapping unit is configured to map the updated feature vector through an embedding operation to generate the voiceprint feature of the speaker.

8. A device for training a voiceprint feature extraction model, comprising: A providing unit is configured to provide sample initial voiceprint feature data about a predetermined speaker, wherein the sample initial voiceprint feature data is extracted from audio data about the speaker, and the sample initial voiceprint feature data includes a plurality of sub-feature data having a temporal contextual relationship, wherein the plurality of sub-feature data corresponds to a plurality of frames of the audio data; The first sample generating unit is configured to generate a sample initial feature vector of the predetermined speaker based on the sample initial voiceprint feature data, which includes: inputting the sample initial voiceprint feature data into a first neural network to obtain the sample initial feature vector, wherein the sample initial feature vector includes a plurality of first elements having the temporal context relationship and corresponding to the plurality of sub-feature data; A second sample generating unit is configured to generate a sample covariance matrix corresponding to the sample initial feature vector, which includes: inputting the sample initial feature vector into a second neural network to obtain the sample covariance matrix corresponding to the sample initial feature vector, wherein the sample covariance matrix has a plurality of second elements corresponding to a plurality of first elements of the sample initial feature vector, and each second element represents the correlation between the corresponding first element and different feature dimensions; a third sample generating unit, configured to generate an updated sample feature vector of the predetermined speaker based on the sample initial feature vector and the sample covariance matrix, wherein the updated sample feature vector is a posterior mean vector of the sample initial feature vector according to a Gaussian distribution; a sample extraction unit configured to extract a voiceprint feature of the predetermined speaker based on the updated sample feature vector; and The network parameter acquisition unit is configured to acquire network parameters for updating the voiceprint feature extraction model based on the voiceprint feature to train the voiceprint feature extraction model.

9. An electronic device, comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.

11. A computer program product comprising a computer program, wherein: The computer program implements the method according to any one of claims 1 to 4 when executed by a processor.

Citation Information

Patent Citations

  • Accurate vocal print recognizing method based on conference scene small sample conditions

    CN109994116A

  • Voiceprint recognition method, device and apparatus for original voice, and storage medium

    CN111524525A