Voiceprint model training, voiceprint extraction method, device, equipment and storage medium

Through the two-stage training method and multimodal feature fusion, the accuracy and robustness of voiceprint recognition technology in complex environments are improved, solving the problem of insufficient robustness and accuracy of voiceprint recognition in existing technologies.

CN114333849BActive Publication Date: 2025-09-26BIGO TECH PTE LTD

Patent Information

Application Number
CN202210158367.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-21
Publication Date
2025-09-26
Estimated Expiration
2042-02-21

AI Technical Summary

Technical Problem

Existing voiceprint recognition technology has poor robustness and accuracy in the face of environmental noise, channel mismatch, multiple speakers and impersonation attacks.

Method used

A two-stage training method is adopted. First, the voiceprint model and classification model are initially trained with user classification as the goal, and then the training is continued with the voiceprint of the same user as the goal. The multimodal feature fusion method is used to enhance the robustness of the voiceprint model.

Benefits of technology

The accuracy and robustness of the voiceprint model in normal scenarios and harsh environments are improved, and the resistance to environmental noise, cross-channel devices and counterfeit attacks is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333849B_ABST
    Figure CN114333849B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device, and storage medium for training and extracting a voiceprint model. The method comprises: determining a voiceprint model and a classification model; extracting at least two voice features from a voice signal, wherein the voice signal is labeled with the user to which it belongs; initially training the voiceprint model and the classification model based on the at least two voice features with the goal of classifying the user; if the initial training is completed, continuing to train the voiceprint model and the classification model based on the at least two voice features with the goal of classifying the user and constraining the voiceprints of the same user, wherein the voiceprint model is used to extract the voiceprint from the at least two voice features, and the classification model is used to predict the user to which the voice signal belongs and is discarded upon completion of continued training. Converging the voiceprint model with the goal of constraining the voiceprints of the same user can improve the performance of the voiceprint model in scenarios such as environmental noise, cross-channel devices, and counterfeit attacks, improve the accuracy of the voiceprint, and thus improve the robustness of the voiceprint model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular to a voiceprint model training, voiceprint extraction method, device, equipment and storage medium. Background Art

[0002] Voiceprint is a biometric feature that is not only specific but also relatively stable. It can be applied to voiceprint recognition (also known as speaker identification), that is, automatically identifying the identity of the speaker corresponding to the current voice based on the speaker's unique personality information contained in the voice.

[0003] Currently, the technologies for extracting voiceprints and performing voiceprint recognition mainly include template matching, Gaussian mixture model (GMM), joint factor analysis (JFA), and neural network (DNN). However, these methods have major robustness issues in the face of environmental noise, channel mismatch, multiple speakers, changes in the speaker themselves, and counterfeit attacks. The accuracy of the extracted voiceprints is low, making voiceprint recognition less stable in business operations. Summary of the Invention

[0004] The present invention provides a voiceprint model training, voiceprint extraction method, device, equipment and storage medium to solve the problem of how to improve the accuracy of voiceprints.

[0005] According to one aspect of the present invention, a method for training a voiceprint model is provided, comprising:

[0006] Determine the voiceprint model and classification model;

[0007] extracting at least two speech features from a speech signal, wherein the speech signal is annotated with an attributed user;

[0008] With the goal of classifying the user, initially training the voiceprint model and the classification model based on at least two of the speech features;

[0009] If the initial training is completed, the voiceprint model and the classification model are continued to be trained based on at least two of the voice features with the goal of classifying the users and constraining the voiceprints of the same user. The voiceprint model is used to extract voiceprints from at least two of the voice features, and the classification model is used to predict the user to whom the voice signal belongs, and is discarded when the continued training is completed.

[0010] According to another aspect of the present invention, a voiceprint extraction method is provided, comprising:

[0011] Collecting voice signals from users;

[0012] extracting at least two speech features from the speech signal;

[0013] Loading a voiceprint model trained based on the voiceprint model described in any embodiment of the present invention;

[0014] At least two of the speech features are input into the voiceprint model to extract the voiceprint of the user.

[0015] According to another aspect of the present invention, a device for training a voiceprint model is provided, comprising:

[0016] Model determination module, used to determine the voiceprint model and classification model;

[0017] A speech feature extraction module, configured to extract at least two speech features from a speech signal, wherein the speech signal is annotated with a user to which it belongs;

[0018] An initial training module, configured to initially train the voiceprint model and the classification model based on at least two of the speech features with the goal of classifying the user;

[0019] The continuing training module is used to continue training the voiceprint model and the classification model based on at least two of the speech features with the goal of classifying the user and constraining the voiceprint of the same user if the initial training is completed. The voiceprint model is used to extract the voiceprint from at least two of the speech features, and the classification model is used to predict the user to which the speech signal belongs, and is discarded when the continued training is completed.

[0020] According to another aspect of the present invention, a voiceprint extraction device is provided, comprising:

[0021] A voice signal acquisition module is used to collect voice signals from users;

[0022] A speech feature extraction module, configured to extract at least two speech features from the speech signal;

[0023] A voiceprint model loading module, used to load a voiceprint model trained based on the voiceprint model described in any embodiment of the present invention;

[0024] The voiceprint extraction module is used to input at least two of the speech features into the voiceprint model to extract the voiceprint of the user.

[0025] According to another aspect of the present invention, an electronic device is provided, comprising:

[0026] at least one processor; and

[0027] a memory communicatively connected to the at least one processor; wherein,

[0028] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the voiceprint model training method and voiceprint extraction method described in any embodiment of the present invention.

[0029] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and the computer program is used to enable a processor to implement the voiceprint model training method and voiceprint extraction method described in any embodiment of the present invention when executed.

[0030] In this embodiment, a voiceprint model and a classification model are determined; at least two voice features are extracted from a voice signal, and the voice signal is labeled with the user to which it belongs; with the goal of classifying the user, the voiceprint model and the classification model are initially trained based on at least two voice features; if the initial training is completed, the voiceprint model and the classification model are further trained based on at least two voice features with the goal of classifying the user and constraining the voiceprints of the same user. The voiceprint model is used to extract voiceprints from at least two voice features, and the classification model is used to predict the user to which the voice signal belongs, and is discarded when the continued training is completed. This embodiment uses at least two speech features as samples to train the voiceprint model, realizes the fusion of multimodal features, enhances the diversity of samples, and can improve the performance of the voiceprint model. The voiceprint model is trained in two stages. The first and second stages both take user classification as the training goal, which can ensure the performance of the voiceprint model in normal scenarios and the accuracy of the voiceprint. On this basis, the second stage converges the voiceprint model with the voiceprint of the same user as the training goal, which can improve the performance of the voiceprint model in scenarios such as environmental noise, cross-channel equipment, and counterfeit attacks, improve the accuracy of the voiceprint, and thus improve the robustness of the voiceprint model.

[0031] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0033] Figure 1 This is a flowchart of a method for training a voiceprint model according to the first embodiment of the present invention;

[0034] Figures 2A to 2BThis is a schematic diagram of the structure of a voiceprint model provided according to the first embodiment of the present invention;

[0035] Figures 3A to 3D 2 is a schematic structural diagram of a residual block provided according to the first embodiment of the present invention;

[0036] Figure 4 This is a flow chart of a voiceprint extraction method provided according to the second embodiment of the present invention;

[0037] Figure 5 This is a flow chart of a voiceprint extraction method provided according to the third embodiment of the present invention;

[0038] Figure 6 2 is a schematic structural diagram of a voiceprint model training device according to a fourth embodiment of the present invention;

[0039] Figure 7 This is a schematic diagram of the structure of a voiceprint extraction device provided according to the fifth embodiment of the present invention;

[0040] Figure 8 It is a structural diagram of an electronic device for implementing the voiceprint model training and voiceprint extraction method of an embodiment of the present invention. DETAILED DESCRIPTION

[0041] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0042] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0043] Example 1

[0044] Figure 1This is a flow chart of a method for training a voiceprint model provided in the first embodiment of the present invention. This embodiment is applicable to the case of training a voiceprint model based on user classification and constraining the voiceprints of the same user. The method can be executed by a voiceprint model training device, which can be implemented in the form of hardware and / or software. The voiceprint model training device can be configured in an electronic device. Figure 1 As shown, the method includes:

[0045] Step 101: Determine the voiceprint model and classification model.

[0046] In this embodiment, a voiceprint model and a classification model can be pre-built, wherein the voiceprint model is used to extract voiceprints from voice signals, the classification model is used to predict the user to which the voice signal belongs, and the classification model is used to assist in training the voiceprint model and is discarded when further training is completed.

[0047] Generally speaking, voiceprint models and classification models are deep learning models. The structures of voiceprint models and classification models are not limited to artificially designed neural networks. They can also be neural networks optimized by model quantization methods, neural networks searched for voiceprint characteristics by NAS (Neural Architecture Search) methods, etc. This embodiment does not impose any restrictions on this.

[0048] When training the voiceprint model, the voiceprint model and classification model are loaded into the memory for running, and the parameters of the voiceprint model and classification model are initialized.

[0049] Step 102: Extract at least two speech features from the speech signal.

[0050] In this embodiment, multiple voice signals can be collected in advance from public data sets and other channels and stored in the database as samples for training the voiceprint model. These voice signals have been labeled with the users to whom they belong (tags), so that the voiceprint model can be trained under the supervision of these users (tags).

[0051] Furthermore, for these speech signals, data enhancement can be performed by using time domain enhancement, frequency domain enhancement, adding noise, reverberation, etc.

[0052] When training the voiceprint model, a speech signal is read from a database, and at least two speech features are extracted from the speech signal frame by frame using at least two operators.

[0053] In one example, the speech features include filter bank features Fbanks, pitch features Pitch, and Mel Frequency Cepstral Coefficients (MFCC).

[0054] Among them, the filter bank feature Fbanks processes the speech signal in a way similar to the human ear, which can improve the performance of voiceprint recognition. In the specific implementation, the speech signal is converted from a time domain signal to a frequency domain signal using methods such as Fast Fourier Transform (FFT). The energy spectrum of the speech signal in the frequency domain is calculated, and the energy spectrum is filtered with Mel frequency Mel. The logarithm of the energy spectrum of the filtered sum is taken to obtain the filter bank feature Fbank.

[0055] The pitch feature Pitch is related to the fundamental frequency (F0) of the sound and reflects the pitch information, that is, the tone, which can be extracted using operators such as YIN.

[0056] The adjacent features of the filter bank features Fbanks are highly correlated (adjacent filter banks overlap), so by performing cepstral transformation on them, the Mel-frequency cepstral coefficients MFCC can be obtained. That is, the extraction of the Mel-frequency cepstral coefficients MFCC is based on the discrete cosine transform on the basis of the filter bank features FBank.

[0057] In this example, the dimension of the filter bank feature Fbanks is 64, the dimension of the pitch feature Pitch is 3, and the dimension of the Mel-frequency cepstral coefficient MFCC is 80.

[0058] Of course, the above-mentioned speech features are merely examples. When implementing this embodiment, other speech features may be set according to actual circumstances, and this embodiment does not limit this. Furthermore, in addition to the above-mentioned speech features, those skilled in the art may also adopt other speech features according to actual needs, and this embodiment does not limit this either.

[0059] Step 103: With the goal of classifying the user, initially train a voiceprint model and a classification model based on at least two speech features.

[0060] In this embodiment, the voiceprint model is trained in two stages with the assistance of the classification model. The first stage of one-stage training is called initial training, and the second stage of two-stage training is called continued training.

[0061] In the initial training, users are classified according to at least two speech features, and the classification result (category) is the user, that is, the user to whom the speech signal belongs is predicted based on at least two speech features. The training goal is to make the user labeled for the speech signal the same as the user predicted for the speech signal.

[0062] In one embodiment of the present invention, step 103 may include the following steps:

[0063] Step 1031: Input at least two speech features into a voiceprint model to extract a voiceprint.

[0064] In this embodiment, at least two speech features are input into the voiceprint model, and the voiceprint model processes the at least two speech features according to its own structure, thereby extracting the possible voiceprint of the user.

[0065] In one embodiment of the present invention, Figure 2A As shown, the voiceprint model includes a time delay neural network (TDNN), a first residual block Res_block_1, a second residual block Res_block_2, a third residual block Res_block_3, a fourth residual block Res_block_4, a self-attentive pooling layer (SAP), and a fully connected layer FC.

[0066] Among them, the first residual block Res_block_1, the second residual block Res_block_2, the third residual block Res_block_3, and the fourth residual block Res_block_4 are all residual blocks (Residualblock), and the residual block is an encapsulation of some structures containing residual networks.

[0067] Furthermore, the structures of the first residual block Res_block_1, the second residual block Res_block_2, the third residual block Res_block_3, and the fourth residual block Res_block_4 can be set according to the requirements of services such as voiceprint recognition. The structures of the first residual block Res_block_1, the second residual block Res_block_2, the third residual block Res_block_3, and the fourth residual block Res_block_4 can be the same or different, and this embodiment does not impose any restrictions on this.

[0068] For example, Figure 3A As shown, the first residual block Res_block_1 has 6 convolutional layers, and a shortcut connection shortcut (also known as a jump connection, i.e. Figure 3A ), such as Figure 3B As shown, the second residual block Res_block_2 has 8 convolutional layers, and a shortcut connection shortcut is inserted every two convolutional layers (i.e. Figure 3B ), such as Figure 3C As shown, the third residual block Res_block_3 has 12 convolutional layers, and a shortcut connection shortcut is inserted every two convolutional layers (i.e. Figure 3C ), such as Figure 3DAs shown, the fourth residual block Res_block_4 has 6 convolutional layers, and a shortcut connection shortcut is inserted every two convolutional layers (i.e. Figure 3D (connected by the curve in the figure).

[0069] In this example, the first residual block Res_block_1, the second residual block Res_block_2, the third residual block Res_block_3, and the fourth residual block Res_block_4 all use 3×3 filters. For the same output feature size, each layer has the same number of filters. If the feature size is halved, the number of filters is doubled to maintain the time complexity of each layer.

[0070] The dimensions of the outputs of the first residual block Res_block_1, the second residual block Res_block_2, the third residual block Res_block_3, and the fourth residual block Res_block_4 increase sequentially, namely 64, 128, 256, and 512, respectively.

[0071] When the input and output have the same dimensions (such as Figures 3A to 3D When connecting with solid curves in the figure, you can directly use the identity shortcut connection.

[0072] When the dimension increases (such as Figures 3B to 3D When the shortcut connection is connected by the dotted curve in the figure, two options are available. One option is that the shortcut connection still performs the identity mapping, but additionally fills the input with zeros to increase the dimension. This option does not introduce additional parameters. The other option is to perform a linear projection (convolution operation) on the shortcut connection to match the dimension (accomplished by 1×1 convolution).

[0073] In this embodiment, the voiceprint model has a simpler structure, fewer parameters, and a faster calculation speed. Accordingly, a smaller number of speech features are used.

[0074] For example, Figure 2A As shown, the at least two speech features include filter bank features Fbanks and pitch features Pitch.

[0075] Of course, in addition to the filter bank features Fbanks and the pitch features Pitch, it can also be the filter bank features Fbanks and the Mel-frequency cepstral coefficients MFCC, or the pitch features Pitch and the Mel-frequency cepstral coefficients MFCC, etc. This embodiment does not limit this.

[0076] In this example, at least two speech features (such as the filter bank feature Fbanks and the pitch feature Pitch) are fused into a new feature through the concatenation function, which is recorded as the first candidate feature.

[0077] Generally, the first candidate feature is a two-dimensional feature, and the residual block mostly processes three-dimensional features. Therefore, the first candidate feature can be input into the time delay neural network TDNN and converted into a three-dimensional feature, which is recorded as the second candidate feature.

[0078] The second candidate feature is input into the first residual block Res_block_1 for processing. The first residual block Res_block_1 maps the second candidate feature into a new feature, which is recorded as the third candidate feature.

[0079] The third candidate feature is input into the second residual block Res_block_2 for processing. The second residual block Res_block_2 maps the third candidate feature into a new feature, which is recorded as the fourth candidate feature.

[0080] The fourth candidate feature is input into the third residual block Res_block_3 for processing. The third residual block Res_block_3 maps the fourth candidate feature into a new feature, which is recorded as the fifth candidate feature.

[0081] The fifth candidate feature is input into the fourth residual block Res_block_4 for processing. The fourth residual block Res_block_4 maps the fifth candidate feature into a new feature, which is recorded as the sixth candidate feature.

[0082] The sixth candidate feature is input into the self-attention pooling layer SAP and aggregated into a new feature, which is recorded as the seventh candidate feature.

[0083] Among them, the self-attention pooling layer SAP uses the self-attention mechanism to aggregate the features at the speech frame level (the sixth candidate feature) into a structural design for sentence-level feature expression. This attention mechanism on the speech frame can make the extracted features more informative.

[0084] In the self-attention pooling layer SAP, a multilayer perceptron (MLP) transforms the speech frame level features (the sixth candidate feature) {x1, x2, ..., x T} is mapped to the hidden layer expression {h1,h2,…,h T}:

[0085] h t =tanh(Wx t +b)

[0086] The softmax function is used to obtain the weight w of each speech frame level feature (the sixth candidate feature) t :

[0087]

[0088] Among them, u is a learnable context vector expression, which is expressed by weight w t For the features at the speech frame level (the sixth candidate feature) {x1, x2, ..., x T} Perform weighted summation to obtain the sentence-level feature (the seventh candidate feature) e:

[0089]

[0090] The seventh candidate feature is input into the fully connected layer FC and mapped into the voiceprint.

[0091] In another embodiment of the present invention, Figure 2B As shown, the voiceprint model includes at least two branch networks ( Figure 2B The dotted box part in the figure), the fourth residual block Res_block_4, the self-attention pooling layer SAP, the fully connected layer FC, each branch network has a time delay neural network TDNN, the first residual block Res_block_1, the second residual block Res_block_2, and the third residual block Res_block_3.

[0092] In this embodiment, the voiceprint model has a more complex structure, a larger number of parameters, and better robustness. Accordingly, a larger number of speech features are used.

[0093] For example, Figure 2B As shown, the at least two speech features include filter bank features Fbanks, pitch features Pitch, and Mel-frequency cepstral coefficients MFCC.

[0094] In this embodiment, the number of branch networks is the same as the number of speech features. For each speech feature, each speech feature is input into each branch network, and the time delay neural network TDNN is called to convert the speech feature into a three-dimensional feature, which is recorded as the first reference feature. The first residual block Res_block_1 is called to map the first reference feature into a new feature, which is recorded as the second reference feature. The second residual block Res_block_2 is called to map the second reference feature into a new feature, which is recorded as the third reference feature. The third residual block Res_block_3 is called to map the third reference feature into a new feature, which is recorded as the fourth reference feature.

[0095] Since the number of branch networks is at least two, the number of fourth reference features outputted therefrom is also two. At this time, at least two fourth reference features are fused into a new feature, which is recorded as the fifth reference feature.

[0096] The fifth reference feature is input into the fourth residual block Res_block_4 and mapped into a new feature, which is recorded as the sixth reference feature.

[0097] The sixth reference feature is input into the self-attention pooling layer SAP and aggregated into a new feature, which is recorded as the seventh reference feature.

[0098] The seventh reference feature is input into the fully connected layer FC and mapped into the voiceprint.

[0099] Step 1032: Input the voiceprint into the classification model to predict the user to whom the voice signal belongs.

[0100] In this embodiment, the user is used as the classification result (category), and the voiceprint is input into the classification model. The classification model processes the voiceprint according to its own structure, thereby predicting the user to which the voice signal belongs.

[0101] In one example, the classification model is one or more fully connected layers (FC), which sequentially map voiceprints to probabilities of belonging to different users. Optionally, the user with the highest probability is selected as the user to which the voice signal belongs.

[0102] Step 1033: Calculate the difference between the marked user and the predicted user as the first loss value.

[0103] In this embodiment, for the same voice signal, there are both users with marked attributes and users with predicted attributes. At this time, a preset first loss function, such as the softmax function and its modified functions (such as A-Softmax, L-Softmax, AM-Softmax, ArcFace, etc.), can be called to calculate the difference between the users with marked attributes and the users with predicted attributes, and obtain the loss value LOSS, which is recorded as the first loss value. Then, the first loss value represents the loss of the voiceprint model between categories.

[0104] Taking ArcFace as an example, the difference between the labeled user and the predicted user can be calculated as the first loss value using the following formula:

[0105]

[0106] Among them, L AAM-softmax is the first loss value, N is the batch size, n is the number of users, y i is the annotated user, θ is the angle between the speech feature and the parameters in the classification model, m is the angle penalty, and s is the hyperparameter.

[0107] Step 1034: Update the classification model and voiceprint model according to the first loss value.

[0108] After completing forward propagation in the voiceprint model and classification model, backpropagation can be performed on the voiceprint model and classification model. The first loss value can be substituted into optimization algorithms such as SGD (stochastic gradient descent) and Adam (Adaptive momentum) to calculate and update the first gradients of the parameters in the voiceprint model and classification model, and then update the parameters in the voiceprint model and classification model according to the first gradients.

[0109] Step 1035 , determine whether the first loss value converges; if so, execute step 1036 ; if not, return to execute step 1031 .

[0110] Step 1036: Determine whether the initial training is completed.

[0111] In this embodiment, a first convergence condition can be set in advance for the first loss value as a condition for stopping training the voiceprint model and the classification model. For example, the number of iterations reaches a threshold, the amplitude of the change in the first loss value for multiple consecutive times is less than a certain threshold, and so on. In each round of iterative training, it is determined whether the first loss value meets the first convergence condition.

[0112] If the first convergence condition is met, it can be considered that the first loss value has converged and remained stable, confirming that the initial training is completed.

[0113] If the first convergence condition is not met, the next round of iterative training can be entered, and steps 1031 to 1034 can be re-executed, and the iterative training can be repeated until the first loss value converges.

[0114] Step 104: If the initial training is completed, the voiceprint model and the classification model are continuously trained based on at least two speech features with the goal of classifying the users and constraining the voiceprints of the same user.

[0115] If the first stage one-stage training (i.e. initial training) is completed, the second stage two-stage training (i.e. continued training) can be carried out. In the second stage two-stage training (i.e. continued training), the voiceprint model and classification model inherit the parameters of the first stage one-stage training (i.e. initial training).

[0116] In the initial training, the training goal is to make the user who labels the speech signal the same as the user who predicts the speech signal. Then, when the initial training is completed, the difference between the voiceprints of different categories (users) can be increased. However, for the same category (user), there are certain differences between their voiceprints. Then, in the continued training, the users are classified according to at least two speech features, and the classification result (category) is the user, that is, the user to whom the speech signal belongs is predicted based on at least two speech features. The training goal is to constrain the voiceprint of the same user, reduce the difference between the voiceprints of the same user, and continue to increase the difference between the voiceprints of different users under the condition that the user who labels the speech signal and the user who predicts the speech signal are the same.

[0117] In one embodiment of the present invention, step 104 may include the following steps:

[0118] Step 1041: Input at least two speech features into a voiceprint model to extract a voiceprint.

[0119] In a structure of a voiceprint model, the voiceprint model includes a time-delay neural network, a first residual block, a second residual block, a third residual block, a fourth residual block, a self-attention pooling layer, and a fully connected layer;

[0120] In this structure, at least two speech features are fused into a first candidate feature;

[0121] Inputting the first candidate feature into a time-delay neural network and converting it into a three-dimensional second candidate feature;

[0122] Input the second candidate feature into the first residual block and map it into the third candidate feature;

[0123] Input the third candidate feature into the second residual block and map it into the fourth candidate feature;

[0124] Inputting the fourth candidate feature into the third residual block and mapping it into the fifth candidate feature;

[0125] Inputting the fifth candidate feature into the fourth residual block and mapping it into the sixth candidate feature;

[0126] Input the sixth candidate feature into the self-attention pooling layer and aggregate it into the seventh candidate feature;

[0127] Input the seventh candidate feature into the fully connected layer and map it into the voiceprint;

[0128] Among them, the at least two speech features include filter bank features and pitch features.

[0129] In another voiceprint model structure, the voiceprint model includes at least two branch networks, a fourth residual block, a self-attention pooling layer, and a fully connected layer. Each branch network has a time-delay neural network, a first residual block, a second residual block, and a third residual block.

[0130] In this structure, each speech feature is input into each branch network, the time-delay neural network is called to convert the speech feature into a three-dimensional first reference feature, the first residual block is called to map the first reference feature into a second reference feature, the second residual block is called to map the second reference feature into a third reference feature, and the third residual block is called to map the third reference feature into a fourth reference feature.

[0131] fusing at least two fourth reference features into a fifth reference feature;

[0132] Inputting the fifth reference feature into the fourth residual block and mapping it into the sixth reference feature;

[0133] Input the sixth reference feature into the self-attention pooling layer and aggregate it into the seventh reference feature;

[0134] Input the seventh reference feature into the fully connected layer and map it into a voiceprint;

[0135] The at least two speech features include filter bank features, pitch features, and Mel-frequency cepstral coefficients.

[0136] Step 1042: Input the voiceprint into the classification model to predict the user to whom the voice signal belongs.

[0137] Step 1043: Calculate the difference between the marked user and the predicted user as the first loss value.

[0138] Taking ArcFace as an example, the difference between the labeled user and the predicted user can be calculated as the first loss value using the following formula:

[0139]

[0140] Among them, L AAM-softmax is the first loss value, N is the batch size, n is the number of users, y i is the annotated user, θ is the angle between the speech feature and the parameters in the classification model, m is the angle penalty, and s is the hyperparameter.

[0141] Step 1044: For the same user, calculate the difference between the user's voiceprints as the second loss value.

[0142] In this embodiment, for the same user, there are multiple voiceprints. At this time, the preset second loss function can be called to calculate the difference between the voiceprints of the user to obtain the loss value LOSS, which is recorded as the second loss value. Then, the second loss value represents the loss of the voiceprint model within the category.

[0143] Taking the Angular Prototypical loss function as an example, the difference between users' voiceprints can be calculated as the second loss value using the following formula:

[0144]

[0145] Among them, L AP is the second loss value, N is the number of users, S j,k is the similarity between the voiceprint of the jth user and the voiceprint of the kth user.

[0146] Furthermore, S j,k =ω·cos(x j,M ,c k )+b, where ω and b are hyperparameters, x j,M is the voiceprint of the jth user, c k is the voiceprint of the kth user.

[0147] Step 1045: Merge the first loss value and the second loss value into a third loss value.

[0148] In this embodiment, when continuing to train the classification model and the voiceprint model, the first loss value L is comprehensively considered. AAM-softmax and the second loss value L AP , the first loss value L AAM-softmax and the second loss value L AP Fusion is the third loss value LOSS total , that is, the third loss value LOSS total Represents the comprehensive loss of the voiceprint model between categories and within categories.

[0149] In a specific implementation, the fusion method can be linear fusion or nonlinear fusion, which is not limited in this embodiment.

[0150] For linear fusion, the first loss value L can be calculated based on the business needs. AAM-softmax Configure the appropriate first weight W AAM-softmax , for the second loss value L AP Configure the appropriate second weight W AP .

[0151] For example, in some businesses, the first loss value L between categories AAM-softmax Than the second loss value L within the category AP is important, then the first weight W AAM-softmax Greater than the second weight W AP .

[0152] During fusion, the first loss value L AAM-softmax Multiply by the preset first weight W AAM-softmax, obtain the first adjustment weight value, and set the second loss value L AP Multiply by the preset second weight W AP , obtain the second weighted value, calculate the sum of the first and second weighted values ​​as the third loss value LOSS total , the fusion process is expressed as follows:

[0153] LOSS total =W AAM-softmax L AAM-softmax +W AP L AP

[0154] Step 1046: Update the classification model and voiceprint model according to the third loss value.

[0155] After completing forward propagation in the voiceprint model and classification model, backpropagation can be performed on the voiceprint model and classification model, and the third loss value can be substituted into optimization algorithms such as SGD and Adam to calculate and update the second gradient of the parameters in the voiceprint model and classification model, and update the parameters in the voiceprint model and classification model according to the second gradient.

[0156] Step 1047 , determine whether the third loss value converges; if so, execute step 1048 ; if not, return to execute step 1041 .

[0157] Step 1048: Determine completion and continue training.

[0158] In this embodiment, a second convergence condition can be set in advance for the third loss value as a condition for stopping training the voiceprint model and the classification model. For example, the number of iterations reaches a threshold, the amplitude of the change in the third loss value for multiple consecutive times is less than a certain threshold, and so on. In each round of iterative training, it is determined whether the third loss value meets the second convergence condition.

[0159] If the second convergence condition is met, the third loss value can be considered to have converged and remained stable. The training is confirmed to be completed and the parameters of the voiceprint model and classification model are output respectively and persisted in the database, waiting to be applied to the business.

[0160] If the second convergence condition is not met, the next round of iterative training can be entered, and steps 1041 to 1046 can be re-executed, and the iterative training can be repeated until the third loss value converges.

[0161] In this embodiment, a voiceprint model and a classification model are determined; at least two voice features are extracted from a voice signal, and the voice signal is labeled with the user to which it belongs; with the goal of classifying the user, the voiceprint model and the classification model are initially trained based on at least two voice features; if the initial training is completed, the voiceprint model and the classification model are further trained based on at least two voice features with the goal of classifying the user and constraining the voiceprints of the same user. The voiceprint model is used to extract voiceprints from at least two voice features, and the classification model is used to predict the user to which the voice signal belongs, and is discarded when the continued training is completed. This embodiment uses at least two speech features as samples to train the voiceprint model, realizes the fusion of multimodal features, enhances the diversity of samples, and can improve the performance of the voiceprint model. The voiceprint model is trained in two stages. The first and second stages both take user classification as the training goal, which can ensure the performance of the voiceprint model in normal scenarios and the accuracy of the voiceprint. On this basis, the second stage converges the voiceprint model with the voiceprint of the same user as the training goal, which can improve the performance of the voiceprint model in scenarios such as environmental noise, cross-channel equipment, and counterfeit attacks, improve the accuracy of the voiceprint, and thus improve the robustness of the voiceprint model.

[0162] Example 2

[0163] Figure 4 This is a flow chart of a voiceprint extraction method provided in the second embodiment of the present invention. This embodiment is applicable to the case where a user's voiceprint is extracted based on a voiceprint model. The method can be executed by a voiceprint extraction device, which can be implemented in the form of hardware and / or software and can be configured in an electronic device. Figure 4 As shown, the method includes:

[0164] Step 401: Collect voice signals from the user.

[0165] In this embodiment, the operating systems in the electronic device include Windows, Android, iOS, etc. These operating systems can support running various applications, such as instant messaging tools, door lock applications, lock screen applications, payment applications, etc.

[0166] While protecting user data such as photos and contacts, these applications can verify the user's identity based on the user's voiceprint to determine whether the user has the authority to operate the user data.

[0167] During the stages of user voiceprint registration and voiceprint verification, the user may be asked to speak. While the user is speaking, the microphone is used to collect the user's voice signal.

[0168] In order to ensure the accuracy of verification, the length of the voice signal is set as a target value (such as 2 seconds). Therefore, during the user voiceprint registration, voiceprint verification, and other stages, if the length of the currently collected voice signal does not reach the target value, the voice signal can be warped forward and backward. That is, the voice signal is recorded in the form of an array, and some data can be extracted from the voice signal and added to the head or tail of the voice signal until the length of the voice signal reaches the target value.

[0169] Step 402: Extract at least two speech features from the speech signal.

[0170] In this embodiment, the user's speech signal may be uniformly sampled multiple times (5 times). In each sampling, at least two operators are used to extract at least two speech features from the speech signal frame by frame.

[0171] In one example, the speech features include filter bank features Fbanks, pitch features Pitch, and Mel-frequency cepstral coefficients MFCC.

[0172] Furthermore, for different business scenarios, the structure of the voiceprint model is different, and the voice features used are also different. Therefore, in a given business scenario, the structure of the voiceprint model is fixed, and the corresponding voice features can be extracted according to the structure of the voiceprint model.

[0173] For example, for Figure 3A The voiceprint model in the example can extract the filter bank features Fbanks and pitch features Pitch from the speech signal. Figure 3B The voiceprint model in ,can extract filter bank features Fbanks, pitch features Pitch, and Mel frequency cepstral coefficients MFCC from the speech signal.

[0174] Step 403: Load the voiceprint model.

[0175] In this embodiment, a voiceprint model (including structure and parameters) may be pre-trained, wherein the voiceprint model is used to extract voiceprints from speech signals.

[0176] During the stages of user voiceprint registration and voiceprint verification, the voiceprint model and its parameters can be loaded into the memory and run, waiting for voiceprint extraction.

[0177] Generally speaking, voiceprint models and classification models are deep learning models. The structures of voiceprint models and classification models are not limited to artificially designed neural networks. They can also be neural networks optimized by model quantization methods, neural networks searched for voiceprint characteristics by NAS (Neural Architecture Search) methods, etc. This embodiment does not impose any restrictions on this.

[0178] When training the voiceprint model, the voiceprint model and classification model are loaded into the memory for running, and the parameters of the voiceprint model and classification model are initialized.

[0179] In one embodiment of the present invention, the voiceprint model training method is as follows:

[0180] Determine the voiceprint model and classification model;

[0181] extracting at least two speech features from a speech signal, wherein the speech signal is annotated with a user;

[0182] With the goal of user classification, the voiceprint model and classification model are initially trained based on at least two speech features;

[0183] If the initial training is completed, the voiceprint model and classification model will be trained based on at least two speech features with the goal of classifying users and constraining the voiceprints of the same user. The voiceprint model is used to extract voiceprints from speech signals, and the classification model is used to predict the user to whom the speech signal belongs, and will be discarded when the training is completed.

[0184] In this embodiment, since the training method of the voiceprint model is basically similar to the application of the first embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the first embodiment, and this embodiment will not be described in detail here.

[0185] Step 404: Input at least two speech features into the voiceprint model to extract the user's voiceprint.

[0186] In each sampling, at least two speech features are input into the voiceprint model. The voiceprint model processes the at least two speech features according to its own structure and extracts the voiceprint from the speech signal as the user's voiceprint, where the number of the user's voiceprints is the same as the number of samples.

[0187] When a user registers a voiceprint, the voiceprint will be associated with the user's identity to realize the identity of the registered user.

[0188] During the stage of user voiceprint verification, the voiceprint waits to match the voiceprints of other registered identities to verify the user's identity.

[0189] In one embodiment of the present invention, the voiceprint model includes a time-delay neural network, a first residual block, a second residual block, a third residual block, a fourth residual block, a self-attention pooling layer, and a fully connected layer; in this embodiment, step 404 includes the following steps:

[0190] fusing at least two speech features into a first candidate feature;

[0191] Inputting the first candidate feature into a time-delay neural network and converting it into a three-dimensional second candidate feature;

[0192] Input the second candidate feature into the first residual block and map it into the third candidate feature;

[0193] Input the third candidate feature into the second residual block and map it into the fourth candidate feature;

[0194] Inputting the fourth candidate feature into the third residual block and mapping it into the fifth candidate feature;

[0195] Inputting the fifth candidate feature into the fourth residual block and mapping it into the sixth candidate feature;

[0196] Input the sixth candidate feature into the self-attention pooling layer and aggregate it into the seventh candidate feature;

[0197] Input the seventh candidate feature into the fully connected layer and map it into the voiceprint;

[0198] Among them, the at least two speech features include filter bank features and pitch features.

[0199] In another embodiment of the present invention, the voiceprint model includes at least two branch networks, a fourth residual block, a self-attention pooling layer, and a fully connected layer. Each branch network includes a time-delay neural network, a first residual block, a second residual block, and a third residual block. In this embodiment, step 404 includes the following steps:

[0200] Input each speech feature into each branch network, call the time-delay neural network to convert the speech feature into a three-dimensional first reference feature, call the first residual block to map the first reference feature into a second reference feature, call the second residual block to map the second reference feature into a third reference feature, and call the third residual block to map the third reference feature into a fourth reference feature;

[0201] fusing at least two fourth reference features into a fifth reference feature;

[0202] Inputting the fifth reference feature into the fourth residual block and mapping it into the sixth reference feature;

[0203] Input the sixth reference feature into the self-attention pooling layer and aggregate it into the seventh reference feature;

[0204] Input the seventh reference feature into the fully connected layer and map it into a voiceprint;

[0205] The at least two speech features include filter bank features, pitch features, and Mel-frequency cepstral coefficients.

[0206] In this embodiment, a voice signal is collected from a user; at least two voice features are extracted from the voice signal; a voiceprint model is loaded; and the at least two voice features are input into the voiceprint model to extract the user's voiceprint. During the voiceprint model training phase, this embodiment uses at least two voice features as samples to train the voiceprint model, achieving multimodal feature fusion, enhancing sample diversity, and improving the performance of the voiceprint model. The voiceprint model is trained in two stages. Both the first and second stages are trained with user classification as the goal, ensuring the performance of the voiceprint model in normal scenarios and the accuracy of the voiceprint. On this basis, the second stage converges the voiceprint model with the goal of constraining the voiceprint of the same user. This improves the performance of the voiceprint model in scenarios such as environmental noise, cross-channel devices, and impersonation attacks, improves the accuracy of the voiceprint, and thus improves the robustness of the voiceprint model.

[0207] Example 3

[0208] Figure 5 This is a flow chart of a voiceprint extraction method provided by the third embodiment of the present invention. This embodiment adds an operation of verifying the user's identity based on the above embodiment. Figure 5 As shown, the method includes:

[0209] Step 501: Collect voice signals from the user.

[0210] Step 502: Extract at least two speech features from the speech signal.

[0211] Step 503: Load the voiceprint model.

[0212] Step 504: Input at least two speech features into the voiceprint model to extract the user's voiceprint.

[0213] Step 505: Obtain the voiceprint of the registered identity as a reference voiceprint.

[0214] During the stage of user voiceprint verification, the voiceprint registered during the stage of user voiceprint registration can be obtained and recorded as a reference voiceprint, wherein the reference voiceprint is associated with the identity of the user.

[0215] Step 506: Calculate the similarity between the user's voiceprint and the reference voiceprint.

[0216] In a specific implementation, the cosine angle between the user's voiceprint and the reference voiceprint can be calculated as the similarity between the two.

[0217] In order to improve the accuracy of verifying user identity, multiple samples can be taken during the user registration voiceprint, voiceprint verification and other stages to obtain multiple voiceprints and multiple reference voiceprints of the user. At this time, the similarity between each voiceprint of the user and each reference voiceprint can be calculated, and the average value of these similarities can be calculated as the overall similarity between the multiple voiceprints of the user and the multiple reference voiceprints.

[0218] Step 507: Determine the user's identity based on the similarity.

[0219] In a specific implementation, the similarity indicates the degree of similarity between the current user and the user with the registered identity. The similarity can be used to determine whether the current user and the user with the registered identity are the same user.

[0220] Generally, if the similarity is greater than or equal to a threshold, it can be confirmed that the current user and the user with the registered identity are the same user, and the registered identity is assigned to the current user.

[0221] If the similarity is less than the threshold, it can be confirmed that the current user and the registered user are not the same user. The identity of the current user cannot be determined for the time being, and other reference voiceprints will be used for matching.

[0222] Example 4

[0223] Figure 6 This is a structural diagram of a voiceprint model training device provided by the fourth embodiment of the present invention. Figure 6 As shown, the device includes:

[0224] Model determination module 601, used to determine the voiceprint model and classification model;

[0225] A speech feature extraction module 602 is configured to extract at least two speech features from a speech signal, wherein the speech signal is annotated with an attributed user;

[0226] An initial training module 603 is configured to initially train the voiceprint model and the classification model based on at least two speech features with the goal of classifying the user;

[0227] The continued training module 604 is used to continue training the voiceprint model and the classification model based on at least two of the voice features with the goal of classifying the user and constraining the voiceprint of the same user if the initial training is completed. The voiceprint model is used to extract the voiceprint from at least two of the voice features, and the classification model is used to predict the user to which the voice signal belongs, and is discarded when the continued training is completed.

[0228] In one embodiment of the present invention, the initial training module 603 includes:

[0229] A voiceprint extraction module, configured to input at least two of the speech features into the voiceprint model to extract a voiceprint;

[0230] A user classification module, configured to input the voiceprint into the classification model and predict the user to which the voice signal belongs;

[0231] A first loss value calculation module, configured to calculate a difference between the marked user and the predicted user as a first loss value;

[0232] a first model updating module, configured to update the classification model and the voiceprint model according to the first loss value;

[0233] a first convergence determination module, configured to determine whether the first loss value has converged; if so, calling the first completion determination module; if not, returning to calling the voiceprint extraction module;

[0234] The first completion determination module is used to determine the completion of the initial training.

[0235] In one embodiment of the present invention, the continuing training module 604 includes:

[0236] A voiceprint extraction module, configured to input at least two of the speech features into the voiceprint model to extract a voiceprint;

[0237] A user classification module, configured to input the voiceprint into the classification model and predict the user to which the voice signal belongs;

[0238] A first loss value calculation module, configured to calculate a difference between the marked user and the predicted user as a first loss value;

[0239] a second loss value calculation module, configured to calculate the difference between the voiceprints of the users as a second loss value;

[0240] A third loss value fusion module, configured to fuse the first loss value and the second loss value into a third loss value;

[0241] a second model updating module, configured to update the classification model and the voiceprint model according to the third loss value;

[0242] a second convergence determination module, configured to determine whether the third loss value converges; if so, calling the second completion determination module; if not, returning to calling the voiceprint extraction module;

[0243] The second completion determination module is used to determine the completion of continuing training.

[0244] In one embodiment of the present invention, the voiceprint model includes a time-delay neural network, a first residual block, a second residual block, a third residual block, a fourth residual block, a self-attention pooling layer, and a fully connected layer;

[0245] The voiceprint extraction module is also used to:

[0246] fusing at least two of the speech features into a first candidate feature;

[0247] Inputting the first candidate feature into the time-delay neural network and converting it into a three-dimensional second candidate feature;

[0248] Inputting the second candidate feature into the first residual block and mapping it into a third candidate feature;

[0249] Inputting the third candidate feature into the second residual block and mapping it into a fourth candidate feature;

[0250] Inputting the fourth candidate feature into the third residual block and mapping it into a fifth candidate feature;

[0251] Inputting the fifth candidate feature into the fourth residual block and mapping it into a sixth candidate feature;

[0252] Inputting the sixth candidate feature into the self-attention pooling layer to aggregate it into a seventh candidate feature;

[0253] Inputting the seventh candidate feature into the fully connected layer and mapping it into a voiceprint;

[0254] Among them, at least two of the speech features include filter bank features and pitch features.

[0255] In another embodiment of the present invention, the voiceprint model includes at least two branch networks, a fourth residual block, a self-attention pooling layer, and a fully connected layer, each of the branch networks having a time-delay neural network, a first residual block, a second residual block, and a third residual block;

[0256] The voiceprint extraction module is further configured to:

[0257] Inputting each of the speech features into each of the branch networks, calling the time-delay neural network to convert the speech features into three-dimensional first reference features, calling the first residual block to map the first reference features into second reference features, calling the second residual block to map the second reference features into third reference features, and calling the third residual block to map the third reference features into fourth reference features;

[0258] fusing at least two of the fourth reference features into a fifth reference feature;

[0259] Inputting the fifth reference feature into the fourth residual block and mapping it into a sixth reference feature;

[0260] Inputting the sixth reference feature into the self-attention pooling layer to aggregate it into a seventh reference feature;

[0261] Inputting the seventh reference feature into the fully connected layer and mapping it into a voiceprint;

[0262] Among them, at least two of the speech features include filter bank features, pitch features, and Mel-frequency cepstral coefficients.

[0263] In one embodiment of the present invention, the first loss value calculation module is further configured to:

[0264] The difference between the marked user and the predicted user is calculated as the first loss value using the following formula:

[0265]

[0266] Among them, L AAM-softmax is the first loss value, N is the batch size, n is the number of users, y i is the user marked, θ is the angle between the speech feature and the parameters in the classification model, m is the angle penalty, and s is the hyperparameter.

[0267] In one embodiment of the present invention, the second loss value calculation module is further configured to:

[0268] The difference between the voiceprints of the same user is calculated as the second loss value using the following formula:

[0269]

[0270] Among them, L AP is the second loss value, N is the number of users, S j,k is the similarity between the voiceprint of the j-th user and the voiceprint of the k-th user.

[0271] In one embodiment of the present invention, the third loss value fusion module is further configured to:

[0272] Multiplying the first loss value by a preset first weight to obtain a first weight adjustment value;

[0273] Multiplying the second loss value by a preset second weight to obtain a second weighted value;

[0274] Calculating a sum of the first weight adjustment value and the second weight adjustment value as a third loss value;

[0275] The first weight is greater than the second weight.

[0276] The voiceprint model training device provided in the embodiment of the present invention can execute the voiceprint model training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the voiceprint model training method.

[0277] Example 5

[0278] Figure 7 This is a schematic diagram of the structure of a voiceprint extraction device provided in Example 5 of the present invention. Figure 7 As shown, the device includes:

[0279] The voice signal collection module 701 is used to collect voice signals from users;

[0280] A speech feature extraction module 702 is configured to extract at least two speech features from the speech signal;

[0281] Voiceprint model loading module 703, used to load the voiceprint model;

[0282] A voiceprint extraction module 704 is configured to input at least two of the speech features into the voiceprint model to extract the user's voiceprint;

[0283] The training method of the voiceprint model is as follows:

[0284] Determine the voiceprint model and classification model;

[0285] extracting at least two speech features from a speech signal, wherein the speech signal is annotated with an attributed user;

[0286] With the goal of classifying the user, initially training the voiceprint model and the classification model based on at least two of the speech features;

[0287] If the initial training is completed, the voiceprint model and the classification model are continued to be trained based on at least two of the voice features with the goal of classifying the users and constraining the voiceprints of the same user. The voiceprint model is used to extract voiceprints from at least two of the voice features, and the classification model is used to predict the user to whom the voice signal belongs, and is discarded when the continued training is completed.

[0288] In one embodiment of the present invention, the apparatus further comprises:

[0289] A reference voiceprint acquisition module is used to obtain the voiceprint of a registered identity as a reference voiceprint;

[0290] A similarity calculation module, configured to calculate the similarity between the user's voiceprint and the reference voiceprint;

[0291] An identity determination module is used to determine the identity of the user according to the similarity.

[0292] In one embodiment of the present invention, the voiceprint model includes a time-delay neural network, a first residual block, a second residual block, a third residual block, a fourth residual block, a self-attention pooling layer, and a fully connected layer;

[0293] The voiceprint extraction module 704 is further configured to:

[0294] fusing at least two of the speech features into a first candidate feature;

[0295] Inputting the first candidate feature into the time-delay neural network and converting it into a three-dimensional second candidate feature;

[0296] Inputting the second candidate feature into the first residual block and mapping it into a third candidate feature;

[0297] Inputting the third candidate feature into the second residual block and mapping it into a fourth candidate feature;

[0298] Inputting the fourth candidate feature into the third residual block and mapping it into a fifth candidate feature;

[0299] Inputting the fifth candidate feature into the fourth residual block and mapping it into a sixth candidate feature;

[0300] Inputting the sixth candidate feature into the self-attention pooling layer to aggregate it into a seventh candidate feature;

[0301] Inputting the seventh candidate feature into the fully connected layer and mapping it into a voiceprint;

[0302] Among them, at least two of the speech features include filter bank features and pitch features.

[0303] In another embodiment of the present invention, the voiceprint model includes at least two branch networks, a fourth residual block, a self-attention pooling layer, and a fully connected layer, each of the branch networks having a time-delay neural network, a first residual block, a second residual block, and a third residual block;

[0304] The voiceprint extraction module 704 is further configured to:

[0305] Inputting each of the speech features into each of the branch networks, calling the time-delay neural network to convert the speech features into three-dimensional first reference features, calling the first residual block to map the first reference features into second reference features, calling the second residual block to map the second reference features into third reference features, and calling the third residual block to map the third reference features into fourth reference features;

[0306] fusing at least two of the fourth reference features into a fifth reference feature;

[0307] Inputting the fifth reference feature into the fourth residual block and mapping it into a sixth reference feature;

[0308] Inputting the sixth reference feature into the self-attention pooling layer to aggregate it into a seventh reference feature;

[0309] Inputting the seventh reference feature into the fully connected layer and mapping it into a voiceprint;

[0310] Among them, at least two of the speech features include filter bank features, pitch features, and Mel-frequency cepstral coefficients.

[0311] The voiceprint extraction device provided in the embodiment of the present invention can execute the voiceprint extraction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the voiceprint extraction method.

[0312] Example 6

[0313] Figure 8 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0314] like Figure 8 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0315] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0316] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the voiceprint model training method and the voiceprint extraction method.

[0317] In some embodiments, the voiceprint model training method and the voiceprint extraction method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the voiceprint model training method and the voiceprint extraction method described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the voiceprint model training method and the voiceprint extraction method in any other appropriate manner (for example, by means of firmware).

[0318] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0319] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0320] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0321] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0322] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0323] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0324] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

Claims

1. A method for training a voiceprint model, characterized in that: include: Determine the voiceprint model and classification model; extracting at least two speech features from a speech signal, wherein the speech signal is annotated with an attributed user; With the goal of classifying the user, initially training the voiceprint model and the classification model based on at least two of the speech features; If the initial training is completed, the voiceprint model and the classification model are further trained based on at least two of the speech features, with the goal of classifying the user and constraining the voiceprints of the same user. The voiceprint model is used to extract the voiceprint from the at least two speech features, and the classification model is used to predict the user to which the voice signal belongs. The classification model is discarded when the training is completed. The method of continuing to train the voiceprint model and the classification model based on at least two speech features with the goal of classifying the user and constraining the voiceprints of the same user includes: Inputting at least two of the speech features into the voiceprint model to extract a voiceprint; Inputting the voiceprint into the classification model to predict the user to whom the voice signal belongs; Calculating a difference between the marked user and the predicted user as a first loss value; Calculating a difference between the voiceprints of the users as a second loss value; Merging the first loss value and the second loss value into a third loss value; Updating the classification model and the voiceprint model according to the third loss value; Determine whether the third loss value converges; if so, determine to complete and continue training; if not, return to executing the at least two steps of inputting the speech features into the voiceprint model to extract the voiceprint.

2. The method according to claim 1, characterized in that The method of initially training the voiceprint model and the classification model based on at least two speech features with the goal of classifying the user includes: Inputting at least two of the speech features into the voiceprint model to extract a voiceprint; Inputting the voiceprint into the classification model to predict the user to whom the voice signal belongs; Calculating a difference between the marked user and the predicted user as a first loss value; updating the classification model and the voiceprint model according to the first loss value; Determine whether the first loss value converges; if so, determine that the initial training is completed; if not, return to executing the step of inputting at least two of the speech features into the voiceprint model to extract the voiceprint.

3. The method according to claim 1 or 2, characterized in that The voiceprint model includes a time-delay neural network, a first residual block, a second residual block, a third residual block, a fourth residual block, a self-attention pooling layer, and a fully connected layer; The step of inputting at least two of the speech features into the voiceprint model to extract the voiceprint comprises: fusing at least two of the speech features into a first candidate feature; Inputting the first candidate feature into the time-delay neural network and converting it into a three-dimensional second candidate feature; Inputting the second candidate feature into the first residual block and mapping it into a third candidate feature; Inputting the third candidate feature into the second residual block and mapping it into a fourth candidate feature; Inputting the fourth candidate feature into the third residual block and mapping it into a fifth candidate feature; Inputting the fifth candidate feature into the fourth residual block and mapping it into a sixth candidate feature; Inputting the sixth candidate feature into the self-attention pooling layer to aggregate it into a seventh candidate feature; Inputting the seventh candidate feature into the fully connected layer and mapping it into a voiceprint; Among them, at least two of the speech features include filter bank features and pitch features.

4. The method according to claim 1 or 2, characterized in that The voiceprint model includes at least two branch networks, a fourth residual block, a self-attention pooling layer, and a fully connected layer. Each branch network has a time-delay neural network, a first residual block, a second residual block, and a third residual block. The step of inputting at least two of the speech features into the voiceprint model to extract the voiceprint comprises: Inputting each of the speech features into each of the branch networks, calling the time-delay neural network to convert the speech features into three-dimensional first reference features, calling the first residual block to map the first reference features into second reference features, calling the second residual block to map the second reference features into third reference features, and calling the third residual block to map the third reference features into fourth reference features; fusing at least two of the fourth reference features into a fifth reference feature; Inputting the fifth reference feature into the fourth residual block and mapping it into a sixth reference feature; Inputting the sixth reference feature into the self-attention pooling layer to aggregate it into a seventh reference feature; Inputting the seventh reference feature into the fully connected layer and mapping it into a voiceprint; Among them, at least two of the speech features include filter bank features, pitch features, and Mel-frequency cepstral coefficients.

5. The method according to claim 1 or 2, characterized in that The calculating, as a first loss value, a difference between the marked user and the predicted user includes: The difference between the marked user and the predicted user is calculated as the first loss value using the following formula: Among them, L AAM-softmax is the first loss value, N is the batch size, n is the number of users, y i is the user marked, θ is the angle between the speech feature and the parameters in the classification model, m is the angle penalty, and s is the hyperparameter.

6. The method according to claim 1, characterized in that The calculating the difference between the voiceprints of the users as a second loss value includes: The difference between the voiceprints of the same user is calculated as the second loss value using the following formula: Among them, L AP is the second loss value, N is the number of users, S j,k is the similarity between the voiceprint of the j-th user and the voiceprint of the k-th user.

7. The method according to claim 1, characterized in that The fusing the first loss value and the second loss value into a third loss value includes: Multiplying the first loss value by a preset first weight to obtain a first weight adjustment value; Multiplying the second loss value by a preset second weight to obtain a second weighted value; Calculating a sum of the first weight adjustment value and the second weight adjustment value as a third loss value; The first weight is greater than the second weight.

8. A voiceprint extraction method, characterized in that: include: Collecting voice signals from users; extracting at least two speech features from the speech signal; Loading a voiceprint model trained according to any one of claims 1 to 7; At least two of the speech features are input into the voiceprint model to extract the voiceprint of the user.

9. The method according to claim 8, characterized in that Also includes: Obtain the voiceprint of the registered identity as a reference voiceprint; Calculating the similarity between the user's voiceprint and the reference voiceprint; The identity of the user is determined based on the similarity.

10. A voiceprint model training device, characterized in that: include: Model determination module, used to determine the voiceprint model and classification model; A speech feature extraction module, configured to extract at least two speech features from a speech signal, wherein the speech signal is annotated with a user to which it belongs; An initial training module, configured to initially train the voiceprint model and the classification model based on at least two of the speech features with the goal of classifying the user; a continuing training module, configured to, upon completion of initial training, continue training the voiceprint model and the classification model based on at least two of the speech features, with the goal of classifying the user and constraining the voiceprint of the same user; the voiceprint model is configured to extract the voiceprint from the at least two speech features; the classification model is configured to predict the user to which the speech signal belongs, and is discarded upon completion of continuing training; Wherein, the continued training module includes: A voiceprint extraction module, configured to input at least two of the speech features into the voiceprint model to extract a voiceprint; A user classification module, configured to input the voiceprint into the classification model and predict the user to which the voice signal belongs; A first loss value calculation module, configured to calculate a difference between the marked user and the predicted user as a first loss value; a second loss value calculation module, configured to calculate the difference between the voiceprints of the users as a second loss value; A third loss value fusion module, configured to fuse the first loss value and the second loss value into a third loss value; a second model updating module, configured to update the classification model and the voiceprint model according to the third loss value; a second convergence determination module, configured to determine whether the third loss value converges; if so, calling the second completion determination module; if not, returning to calling the voiceprint extraction module; The second completion determination module is used to determine the completion of continuing training.

11. A voiceprint extraction device, characterized in that: include: A voice signal acquisition module is used to collect voice signals from users; A speech feature extraction module, configured to extract at least two speech features from the speech signal; A voiceprint model loading module, configured to load a voiceprint model trained according to any one of claims 1 to 7; The voiceprint extraction module is used to input at least two of the speech features into the voiceprint model to extract the voiceprint of the user.

12. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the voiceprint model training method described in any one of claims 1-7 or the voiceprint extraction method described in any one of claims 8-9.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is used to enable a processor to implement the voiceprint model training method described in any one of claims 1 to 7 or the voiceprint extraction method described in any one of claims 8 to 9 when executed.

Citation Information

Patent Citations

  • Voiceprint determination method and device

    CN111326161A

  • Voiceprint recognition model training method, recognition method, electronic equipment and storage medium

    CN111951791A

  • Model training method and device, identity recognition method and device and electronic equipment

    CN114049900A

Cited By

  • Voiceprint recognition method based on custom keyword

    CN115831129A

  • A voiceprint recognition method based on a custom keyword

    CN115831129B