Voiceprint Recognition Method, Training Method of Voiceprint Extraction Network and Related Devices

During the training process of the voiceprint extraction network, combining supervision losses and semi-supervised losses, and using massive voice data without speaker tags and a small amount of voice data with speaker tags for training, the problem of insufficient generalization capabilities and recognition accuracy of the existing voiceprint extraction network is solved, and higher recognition performance and accuracy are achieved.

CN114783414BActive Publication Date: 2025-06-17ANHUI IFLYTEK INTELLIGENT SYST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210307702.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2025-06-17
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

The existing voiceprint extraction network has shortcomings in generalization capabilities and recognition accuracy, and it is difficult to effectively distinguish different speakers.

Method used

During the training process of the voiceprint extraction network, combining supervision losses and semi-supervised losses, a large amount of speech data without speaker tags and a small amount of speech data with speaker tags are used for training, so as to improve the generalization ability and recognition accuracy of the network.

Benefits of technology

It significantly improves the recognition performance of the voiceprint extraction network and the accuracy of speaker recognition, and enhances the robustness of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114783414B_ABST
    Figure CN114783414B_ABST
Patent Text Reader

Abstract

The present application discloses a voiceprint recognition method, a training method for a voiceprint extraction network, and related devices. The voiceprint recognition method includes: extracting features from the audio to be recognized to obtain the audio features to be recognized; inputting the audio features to be recognized into the trained voiceprint extraction network to obtain the voiceprint features to be recognized; wherein, when training the voiceprint extraction network, it is based on voice data with speaker labels and voice data without speaker labels, and the total loss used when training the voiceprint extraction network is related to the supervised loss and the semi-supervised loss; determining the speaker corresponding to the audio to be recognized based on the voiceprint features to be recognized. Through the above design method, the present application can significantly improve the recognition performance of the voiceprint extraction network and improve the accuracy of speaker recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of voiceprint recognition, and particularly relates to a voiceprint recognition method, a training method for a voiceprint extraction network, and related devices. Background Art

[0002] Speaker recognition, also known as voiceprint recognition, is a technology for distinguishing the identity of speakers based on human speech. A voiceprint refers to the sound wave spectrum of speech information, which is essentially a combination of time-domain information and frequency-domain information. The theoretical basis of voiceprint recognition is that each voice has unique characteristics, and different people's voices can be effectively distinguished through these characteristics.

[0003] In recent years, in the field of speaker recognition, with the rise of deep learning methods, researchers have proposed various voiceprint extraction networks. However, the current voiceprint extraction networks have problems of low generalization ability and low recognition accuracy. Summary of the Invention

[0004] This application provides a voiceprint recognition method, a training method for a voiceprint extraction network, and related devices to improve the generalization ability of the voiceprint extraction network and the recognition accuracy.

[0005] To solve the above technical problems, a technical solution adopted in this application is: providing a voiceprint recognition method, including: extracting features from the audio to be recognized to obtain the audio features to be recognized; inputting the audio features to be recognized into the trained voiceprint extraction network to obtain the voiceprint features to be recognized; wherein, when training the voiceprint extraction network, it is based on the speech data with speaker labels and the speech data without speaker labels, and the total loss used when training the voiceprint extraction network is related to the supervised loss and the semi-supervised loss; determining the speaker corresponding to the audio to be recognized based on the voiceprint features to be recognized.

[0006] To solve the above technical problems, another technical solution adopted by this application is: to provide a training method for a voiceprint extraction network, including: constructing a first sample set and a second sample set; wherein, the first sample set contains multiple first audio samples, and each of the first audio samples is set with a speaker label; the second sample set contains multiple second audio samples, and each of the second audio samples is not set with a speaker label; respectively performing feature extraction on multiple first audio samples to obtain multiple first audio features, and respectively performing feature extraction on multiple second audio samples to obtain multiple second audio features; inputting multiple first audio features and multiple second audio features into a first voiceprint extraction network to obtain a first posterior probability corresponding to each audio feature, and inputting multiple first audio features and multiple second audio features into a second voiceprint extraction network to obtain a second posterior probability corresponding to each audio feature; obtaining the supervised loss based on the first posterior probability and the second posterior probability of all the first audio features, and obtaining the semi-supervised loss based on the first posterior probability and the second posterior probability of all the second audio features; obtaining a total loss based on the supervised loss and the semi-supervised loss, and respectively adjusting the parameters in the first voiceprint extraction network and the second voiceprint extraction network based on the total loss; in response to reaching the stop training condition, outputting one of the first voiceprint extraction network and the second voiceprint extraction network as the trained voiceprint extraction network.

[0007] To solve the above technical problems, another technical solution adopted by this application is: to provide a voiceprint recognition device, including: a first extraction module, configured to perform feature extraction on the audio to be recognized to obtain the audio feature to be recognized; a second extraction module, connected to the first extraction module, configured to input the audio feature to be recognized into the trained voiceprint extraction network to obtain the voiceprint feature to be recognized; wherein, when training the voiceprint extraction network, it is based on the voice data with speaker labels and the voice data without speaker labels, and the total loss adopted when training the voiceprint extraction network is related to the supervised loss and the semi-supervised loss; a determination module, connected to the second extraction module, configured to determine the speaker corresponding to the audio to be recognized based on the voiceprint feature to be recognized.

[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, including a memory and a processor coupled to each other, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the voiceprint recognition method in any of the above embodiments.

[0009] To solve the above technical problems, another technical solution adopted by this application is: to provide a storage device storing program instructions that can be run by a processor, and the program instructions are used to implement the voiceprint recognition method described in any of the above embodiments.

[0010] Different from the prior art, the beneficial effect of this application is that in the voiceprint recognition method provided by this application, a trained voiceprint extraction network is used to extract the voiceprint features to be recognized; among them, when training the voiceprint extraction network, a large amount of speaker-unlabeled voice data that is easy to obtain and has a low cost and a small amount of speaker-labeled voice data are used for training, and a combination of a supervised loss function and a semi-supervised loss function is used as the objective function for training. Compared with the prior art, in which only a small amount of speaker-labeled voice data is used for training, it has better robustness, can learn more speaker-related information in the large amount of speaker-unlabeled voice data, significantly improves the recognition performance of the voiceprint extraction network, and improves the accuracy of speaker recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:

[0012] Figure 1 It is a schematic flowchart of an embodiment of the voiceprint recognition method of this application;

[0013] Figure 2 It is a schematic structural diagram of an embodiment of the Xvector network model;

[0014] Figure 3 It is a schematic flowchart of an embodiment of the training method of the voiceprint extraction network of this application;

[0015] Figure 4 For Figure 3 It is a schematic structural diagram of an embodiment of the training method of the voiceprint extraction network in

[0016] Figure 5 For Figure 3 It is a schematic flowchart of an embodiment after step S206 in

[0017] Figure 6 It is a schematic structural diagram of an embodiment of the voiceprint recognition device of this application;

[0018] Figure 7 It is a schematic structural diagram of an embodiment of the electronic device of this application;

[0019] Figure 8 This is a schematic structural diagram of an embodiment of the storage device of the present application. Specific embodiments

[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0021] Please refer to Figure 1 , Figure 1 This is a schematic flowchart of an embodiment of the voiceprint recognition method of the present application. The voiceprint recognition method includes:

[0022] S101: Extract features from the audio to be recognized to obtain the audio features to be recognized.

[0023] Specifically, the above-mentioned audio features to be recognized may be FB features (i.e., frequency domain features). The specific implementation process of the above step S101 may be: performing pre-emphasis, framing, windowing, short-time Fourier transform (STFT), Mel filtering, mean removal, etc. on the audio to be recognized in sequence to obtain the audio features to be recognized. For example, the audio features to be recognized may be 48-dimensional FB features.

[0024] S102: Input the audio features to be recognized into the trained voiceprint extraction network to obtain the voiceprint features to be recognized; wherein, when training the voiceprint extraction network, it is based on the voice data with speaker labels and the voice data without speaker labels, and the total loss used when training the voiceprint extraction network is related to the supervised loss and the semi-supervised loss.

[0025] Specifically, the above-mentioned voiceprint extraction network may be an Xvector network model based on the structure of a time-delay neural network (TDNN). This model can extract the first-order and second-order statistical information of audio features in the time domain, and establish the correlation relationship between the time domain and the frequency domain through the delay mapping mechanism of the time-delay neural network. In the experimental results of many papers, the method of using TDNN to model the temporal structure of speech signals has significant advantages in the field of speaker recognition compared with the traditional full-variable system model. Please refer to Figure 2 , Figure 2 This is a schematic structural diagram of an embodiment of the Xvector network model. The Xvector network model generally includes multiple layers of frame-level TDNN layers (for example, Figure 2 includes five TDNNs, namely TDNN1,..., TDNN5), a statistical pooling layer (for example, Figure 2(Statistic Pooling), two fully-connected layers at the sentence level (e.g., Figure 2 (Linear1 and Linear2 in Figure 2 ), and one activation function layer (e.g.,

[0026] Softmax in

[0027] ); among them, the voiceprint feature to be recognized is the output of Linear2.

[0028] Optionally, in this embodiment, the cosine distance can be used to compare the similarity between the voiceprint feature to be recognized and the existing voiceprint features, and it is expressed by the formula as follows:

[0029]

[0030] where Similarity is the similarity, is the voiceprint feature to be recognized, is the existing voiceprint feature.

[0031] In the above design method, this application uses the trained voiceprint extraction network to extract the voiceprint feature to be recognized; among them, when training this voiceprint extraction network, a large amount of speaker-unlabeled voice data that is easy to obtain and has a low cost and a small amount of speaker-labeled voice data are used for training, and a combination of a supervised loss function and a semi-supervised loss function is used as the objective function for training. Compared with the prior art, which only uses a small amount of speaker-labeled voice data for training, its robustness is better, it can learn more speaker-related information in the large amount of speaker-unlabeled voice data, significantly improve the recognition performance of the voiceprint extraction network, and improve the accuracy of speaker recognition.

[0032] The training process of the above voiceprint extraction network will be described in detail below. Please refer to Figure 3 and Figure 4 , Figure 3 is a schematic flowchart of an implementation manner of the training method of the voiceprint extraction network of this application, Figure 4 is Figure 3 a schematic structural diagram of an implementation manner of the training method of the voiceprint extraction network in

[0033] S201: Construct a first sample set and a second sample set; wherein, the first sample set contains multiple first audio samples, and each first audio sample is set with a speaker label; the second sample set contains multiple second audio samples, and each second audio sample is not set with a speaker label.

[0034] Specifically, in this embodiment, the number of the above-mentioned first audio samples may be less than the number of the second audio samples, and the second audio samples can be obtained by means of the network or the like.

[0035] S202: Respectively perform feature extraction on multiple first audio samples to obtain multiple first audio features, and respectively perform feature extraction on multiple second audio samples to obtain multiple second audio features.

[0036] Specifically, the first audio features and the second audio features may be FB features (i.e., frequency domain features), and the extraction process is similar to the process of obtaining the audio features to be recognized in the above step S101, and will not be elaborated here.

[0037] S203: Input multiple first audio features and multiple second audio features into the first speaker verification network 10 (as shown in Figure 4 ) to obtain the first posterior probability corresponding to each audio feature, and input multiple first audio features and multiple second audio features into the second speaker verification network 12 (as shown in Figure 4 ) to obtain the second posterior probability corresponding to each audio feature.

[0038] Specifically, the semi-supervised training process in this application is composed of two parallel first speaker verification networks 10 and second speaker verification networks 12, and the first speaker verification network 10 and the second speaker verification network 12 have the same network structure, but the randomly initialized parameters are different. Optionally, both the first speaker verification network 10 and the second speaker verification network 12 are Xvector network models based on the structure of the time delay neural network (TDNN), and the network structures of the first speaker verification network 10 and the second speaker verification network 12 can be as shown in Figure 2 . Among them, the first posterior probability may be the network output after Softmax normalization in the first speaker verification network 10, and the second posterior probability may be the network output after Softmax normalization in the second speaker verification network 12.

[0039] S204: Obtain a supervised loss based on the first posterior probability and the second posterior probability of all first audio features, and obtain a semi-supervised loss based on the first posterior probability and the second posterior probability of all second audio features.

[0040] Specifically, the above semi-supervised loss is related to the second audio feature without a speaker label, and the above supervised loss is related to the first audio feature with a speaker label.

[0041] In one embodiment, the process of obtaining the semi-supervised loss based on the first posterior probability and the second posterior probability of all the second audio features in step S204 may be as follows: A. Obtain the corresponding first pseudo-speaker label based on the first posterior probability of each second audio sample, and obtain the corresponding second pseudo-speaker label based on the second posterior probability of each second audio sample. Specifically, the corresponding first pseudo-speaker label can be predicted through the first posterior probability, and the corresponding second pseudo-speaker label can be predicted through the second posterior probability. B. Obtain the first sub-loss based on the first posterior probability and the second pseudo-speaker label of the same second audio sample, and obtain the second sub-loss based on the second posterior probability and the first pseudo-speaker label of the same second audio sample. The purpose of this step is to use the pseudo-label output by one speaker verification network to supervise the posterior probability output by the other speaker verification network. C. Obtain the first sum value of the first sub-loss and the second sub-loss of all the second audio samples, and use the first ratio of the first sum value to the number of the second audio samples as the semi-supervised loss. It is expressed by the formula as follows:

[0042]

[0043] where, Lssl is the semi-supervised loss, score1 is the first posterior probability, Pseudolabel2 is the second pseudo-speaker label, score2 is the second posterior probability, Pseudolabel1 is the first pseudo-speaker label, D u is the number of the second audio samples; L CE is the cross-entropy loss function. It can be said that the above semi-supervised loss is the sum of the losses of the samples without speaker labels on two parallel speaker verification networks.

[0044] In addition, in the above embodiment, the first sub-loss and the second sub-loss are calculated based on the L CE cross-entropy loss function; of course, in other embodiments, other loss functions can also be used to calculate the first sub-loss and the second sub-loss. For example, replace the L CE cross-entropy loss function in the above formula with L AM (AM-softmax) loss function; where, the calculation formula of the L AM loss function is expressed as follows:

[0045]

[0046] where, N is the size of the current sample quantity, cosθ y,i is the audio feature of sample i and the ythi The cosine distance between the class center vectors of the speakers, m is the cosine distance threshold, s is the scale factor which serves to accelerate convergence, and c is the number of speakers. By adjusting the threshold m, the minimum cosine distance within the class is made less than the maximum cosine distance between classes.

[0047] In another embodiment, the step of obtaining the supervised loss based on the first posterior probability and the second posterior probability of all the first audio features in the above step S204 includes: A. Obtaining a third sub-loss based on the first posterior probability of each first audio sample and the corresponding speaker label, and obtaining a fourth sub-loss based on the second posterior probability of each first audio sample and the corresponding speaker label. B. Obtaining a second sum of the third sub-loss and the fourth sub-loss of all the first audio samples, and taking the second ratio of the second sum to the number of the first audio samples as the supervised loss. It is expressed by the formula as follows:

[0048]

[0049] where Ls is the supervised loss, score1 is the first posterior probability, label is the corresponding speaker label, score2 is the second posterior probability, D l is the number of the first audio samples; L AM is the AM-softmax loss function, and its specific calculation process can be as shown above and will not be elaborated here. It can be said that the above supervised loss is the sum of the losses of the samples with speaker labels on two parallel speaker verification networks. Of course, in other embodiments, other loss functions can also be used to calculate the third sub-loss and the fourth sub-loss, and the present application does not limit this.

[0050] S205: Obtain the total loss based on the supervised loss and the semi-supervised loss, and respectively adjust the parameters in the first speaker verification network and the second speaker verification network based on the total loss.

[0051] Specifically, the sum of the above supervised loss and semi-supervised loss can be used as the total loss. Or, in other embodiments, the first product can be obtained by multiplying the supervised loss by the first weight, the second product can be obtained by multiplying the semi-supervised loss by the second weight, and the sum of the first product and the second product can be used as the total loss.

[0052] S206: In response to reaching the stop training condition, output one of the first speaker verification network and the second speaker verification network as the trained speaker verification network.

[0053] Specifically, the above stop training condition can be that the total loss converges, or a preset number of training times is reached, etc. After stopping training, one of the trained first speaker verification network and the second speaker verification network can be intercepted as the output of the trained speaker verification network.

[0054] Optionally, the specific extraction method may be: obtaining a third sum value of the second sub-loss and the third sub-loss related to the first voiceprint extraction network 10 when stopping training, and obtaining a fourth sum value of the first sub-loss and the fourth sub-loss related to the second voiceprint extraction network 12 when stopping training; outputting the voiceprint extraction network corresponding to the smaller one of the third sum value and the fourth sum value.

[0055] In addition, after the pre-training in Figure 3 is completed, the trained voiceprint extraction network output can be further fine-tuned. For details, please refer to Figure 5 , Figure 5 For Figure 3 is a schematic flowchart of an implementation manner after step S206 in

[0056] S301: Construct a third sample set; where the third sample set contains a plurality of third audio samples, and each third audio sample is set with a speaker label, and at least one third audio sample corresponds to the same speaker label.

[0057] Specifically, generally speaking, since the speech data with speaker labels is limited, we can augment the speech data through pitch-shifting without changing the speed. Specifically, the pitch is determined by the fundamental frequency, that is, the higher the fundamental frequency, the higher the pitch, and the lower the fundamental frequency, the lower the pitch. Therefore, pitch-shifting without changing the speed of speech means changing the size of the speaker's fundamental frequency while keeping the speech rate and semantics unchanged, that is, keeping the short-time spectral envelope (the position and bandwidth of the formants) and the time process basically unchanged. The main idea of this operation is to adjust the pitch of all the speech of the same person to N times the original, and then consider it as a new speaker. Here, N takes 0.8 / 0.9 / 1.1 / 1.2, that is, the number of speakers increases by 4 times, and the number of training samples also increases by 4 times, enriching the training samples, expanding the voiceprint sample space, enhancing the content of its speaker information, and helping to train the speaker model.

[0058] Correspondingly, the specific process may be: A. Obtain a plurality of third audio samples, and each third audio sample is set with a speaker label; optionally, the third audio sample may be the first audio sample used in step S201. B. Adjust the pitch of at least some of the third audio samples to form new third audio samples, and the speaker label of the adjusted third audio sample is different from that of the third audio sample before adjustment. C. Construct a third sample set based on all the third audio samples before adjustment and all the third audio samples after adjustment.

[0059] S302: Select a subset of samples from the third sample set using the trained voiceprint extraction network; wherein, the subset of samples contains multiple third audio samples with similarity exceeding a threshold and having different speaker labels.

[0060] Specifically, the specific implementation process of the above step S302 can be as follows:

[0061] A. Obtain the class center matrix corresponding to each speaker label in the third sample set using the trained voiceprint extraction network. Specifically, multiple third audio samples in the third sample set can be input into the trained voiceprint extraction network; there is a virtual layer before the posterior probability is output in the voiceprint extraction network, and the virtual layer can output the center vector corresponding to the current third audio sample, and the center vectors with the same speaker label are aggregated to form the corresponding class center matrix.

[0062] B. Obtain the similarity matrix between the class center matrix of each speaker label in the third sample set and the class center matrices of the remaining speaker labels. Specifically, for each speaker label in the third sample set, the cosine similarity between the class center matrix of the current speaker label and the class center matrices of the remaining speaker labels can be obtained, and then the similarity matrix is constructed.

[0063] C. Randomly select at least one speaker label from the third sample set as the first target label, and obtain at least one of the remaining speaker labels with relatively large similarity values from the similarity matrix of each first target label as the second target label.

[0064] D. For each second target label, randomly select the third audio samples with the second target label and add them to the subset of samples.

[0065] Specifically, N speaker labels can be randomly selected from the third sample set as the first target labels; for each first target label, according to the relevant similarity matrix, M most similar speaker labels related to it are obtained as the second target labels; then one sample with the second target label is randomly selected and added to the subset of samples, and finally the subset of samples contains N * M samples.

[0066] S303: Adjust the trained voiceprint extraction network based on the loss corresponding to the subset of samples.

[0067] Specifically, obtain the L AM loss corresponding to the subset of samples, and adjust the trained voiceprint extraction network based on the L AM loss.

[0068] In addition, the training process of the above steps S301 - S303 can be carried out multiple times. And during the training cycle, the similarity matrix is updated at the beginning of each generation, and the top M similar examples are resampled.

[0069] In the above design method, after the previous pre - training is completed, the network is fine - tuned using the training data with speaker labels after pitch - modulation augmentation, and a supervised training fine - tuning method is adopted in the way of hard example mining. Hard example mining can select the most similar speakers in the data and sample the data of similar speakers as hard examples for key training. This method can enhance the model's ability to distinguish voiceprints with similar timbres, and enhance the robustness and voiceprint recognition performance of the model.

[0070] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of the voiceprint recognition device of the present application. The voiceprint recognition device specifically includes: a first extraction module 20, a second extraction module 22, and a determination module 24. Among them, the first extraction module 20 is used to extract features from the audio to be recognized to obtain the audio features to be recognized. The second extraction module 22 is connected to the first extraction module 20 and is used to input the audio features to be recognized into the trained voiceprint extraction network to obtain the voiceprint features to be recognized; wherein, when training the voiceprint extraction network, it is based on the speech data with speaker labels and the speech data without speaker labels, and the total loss adopted when training the voiceprint extraction network is related to the supervised loss and the semi - supervised loss. The determination module 24 is connected to the second extraction module 22 and is used to determine the speaker corresponding to the audio to be recognized based on the voiceprint features to be recognized.

[0071] Please continue to refer to Figure 6, the voiceprint recognition device provided by this application further includes a first training module 26. The first training module 26 is connected to the second extraction module 22 and is used to construct a first sample set and a second sample set. Among them, the first sample set contains multiple first audio samples, and each first audio sample is set with a speaker label; the second sample set contains multiple second audio samples, and each second audio sample is not set with a speaker label. Feature extraction is respectively performed on multiple first audio samples to obtain multiple first audio features, and feature extraction is respectively performed on multiple second audio samples to obtain multiple second audio features. The multiple first audio features and the multiple second audio features are input into the first voiceprint extraction network to obtain the first posterior probability corresponding to each audio feature, and the multiple first audio features and the multiple second audio features are input into the second voiceprint extraction network to obtain the second posterior probability corresponding to each audio feature. A supervised loss is obtained based on the first posterior probability and the second posterior probability of all first audio features, and a semi-supervised loss is obtained based on the first posterior probability and the second posterior probability of all second audio features. A total loss is obtained based on the supervised loss and the semi-supervised loss, and the parameters in the first voiceprint extraction network and the second voiceprint extraction network are respectively adjusted based on the total loss. In response to reaching the stop training condition, one of the first voiceprint extraction network and the second voiceprint extraction network is output as the trained voiceprint extraction network.

[0072] Among them, the step of obtaining the semi-supervised loss based on the first posterior probability and the second posterior probability of all second audio features includes: obtaining the corresponding first pseudo-speaker label based on the first posterior probability of each second audio sample, and obtaining the corresponding second pseudo-speaker label based on the second posterior probability of each second audio sample; obtaining the first sub-loss based on the first posterior probability and the second pseudo-speaker label of the same second audio sample, and obtaining the second sub-loss based on the second posterior probability and the first pseudo-speaker label of the same second audio sample; obtaining the first sum value of the first sub-loss and the second sub-loss of all second audio samples, and taking the first ratio of the first sum value to the number of second audio samples as the semi-supervised loss.

[0073] The step of obtaining the supervised loss based on the first posterior probability and the second posterior probability of all first audio features includes: obtaining the third sub-loss based on the first posterior probability and the corresponding speaker label of each first audio sample, and obtaining the fourth sub-loss based on the second posterior probability and the corresponding speaker label of each first audio sample; obtaining the second sum value of the third sub-loss and the fourth sub-loss of all first audio samples, and taking the second ratio of the second sum value to the number of first audio samples as the supervised loss.

[0074] Please continue to refer to Figure 6, the voiceprint recognition device provided by this application further includes a second training module 28, which is connected between the first training module 26 and the second extraction module 22, and is used to construct a third sample set after the first training module 26 finishes training; wherein, the third sample set contains multiple third audio samples, and each third audio sample is set with a speaker label, and at least one third audio sample corresponds to the same speaker label; use the trained voiceprint extraction network to screen out a subset from the third sample set; wherein, the subset contains multiple third audio samples with a similarity exceeding the threshold and having different speaker labels; adjust the trained voiceprint extraction network based on the loss corresponding to the subset.

[0075] Among them, the step of screening out a subset from the third sample set by using the trained voiceprint extraction network includes: obtaining a class center matrix corresponding to each speaker label in the third sample set by using the trained voiceprint extraction network; obtaining a similarity matrix between the class center matrix of each speaker label in the third sample set and the class center matrices of the remaining speaker labels; randomly selecting at least one speaker label from the third sample set as a first target label, and obtaining at least one remaining speaker label with a larger similarity value from the similarity matrix of each first target label as a second target label; for each second target label, randomly select a third audio sample with the second target label and add it to the subset.

[0076] Among them, the step of constructing the third sample set includes: obtaining multiple third audio samples, and each third audio sample is set with a speaker label; adjusting the pitch of at least some of the third audio samples to form new third audio samples, and the speaker label of the adjusted third audio sample is different from that of the third audio sample before adjustment; constructing a third sample set based on all the third audio samples before adjustment and all the third audio samples after adjustment.

[0077] Please refer to Figure 7 , Figure 7FIG. 0 is a schematic structural diagram of an embodiment of an electronic device according to the present application. The electronic device includes a memory 32 and a processor 30 that are coupled to each other. Program instructions are stored in the memory 32, and the processor 30 is configured to execute the program instructions to implement the method in any of the above embodiments. Specifically, the electronic device includes, but is not limited to, a desktop computer, a laptop computer, a tablet computer, a server, etc., which are not limited herein. In addition, the processor 30 may also be referred to as a CPU (Central Processing Unit). The processor 30 may be an integrated circuit chip with signal processing capabilities. The processor 30 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 30 may be implemented jointly by integrated circuit chips.

[0078] Please refer to Figure 8 , Figure 8 FIG. 7 is a schematic structural diagram of an embodiment of a storage device according to the present application. The storage device 40 stores program instructions 400 that can be run by a processor, and the program instructions 400 are used to implement the method in any of the above embodiments.

[0079] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0080] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0081] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0082] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0083] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A voiceprint recognition method, characterized in that, Including: Performing feature extraction on the audio to be recognized to obtain the audio feature to be recognized; Inputting the audio feature to be recognized into the trained voiceprint extraction network to obtain the voiceprint feature to be recognized; wherein, when training the voiceprint extraction network, it is based on the speech data with speaker labels and the speech data without speaker labels, and the total loss used when training the voiceprint extraction network is related to the supervised loss and the semi-supervised loss, and the supervised loss and the semi-supervised loss are obtained based on the following steps: constructing a first sample set and a second sample set; wherein, the first sample set contains multiple first audio samples, and each of the first audio samples is set with a speaker label; the second sample set contains multiple second audio samples, and each of the second audio samples is not set with a speaker label; performing feature extraction on the multiple first audio samples respectively to obtain multiple first audio features, and performing feature extraction on the multiple second audio samples respectively to obtain multiple second audio features; inputting the multiple first audio features and the multiple second audio features into the first voiceprint extraction network to obtain the first posterior probability corresponding to each audio feature, and inputting the multiple first audio features and the multiple second audio features into the second voiceprint extraction network to obtain the second posterior probability corresponding to each audio feature; obtaining the supervised loss based on the first posterior probability and the second posterior probability of all the first audio features, and obtaining the semi-supervised loss based on the first posterior probability and the second posterior probability of all the second audio features; the voiceprint extraction network is one of the first voiceprint extraction network and the second voiceprint extraction network; Determining the speaker corresponding to the audio to be recognized based on the voiceprint feature to be recognized.

2. The voiceprint recognition method according to claim 1, characterized in that, The process of training the voiceprint extraction network includes: Obtaining the total loss based on the supervised loss and the semi-supervised loss, and respectively adjusting the parameters in the first voiceprint extraction network and the second voiceprint extraction network based on the total loss; In response to reaching the stop training condition, outputting one of the first voiceprint extraction network and the second voiceprint extraction network as the trained voiceprint extraction network.

3. The voiceprint recognition method according to claim 2, characterized in that, The step of obtaining the semi-supervised loss based on the first posterior probability and the second posterior probability of all the second audio features includes: Obtaining the corresponding first pseudo-speaker label based on the first posterior probability of each second audio sample, and obtaining the corresponding second pseudo-speaker label based on the second posterior probability of each second audio sample; Obtaining the first sub-loss based on the first posterior probability and the second pseudo-speaker label of the same second audio sample, and obtaining the second sub-loss based on the second posterior probability and the first pseudo-speaker label of the same second audio sample; Obtaining the first sum value of the first sub-loss and the second sub-loss of all the second audio samples, and taking the first ratio of the first sum value to the number of the second audio samples as the semi-supervised loss.

4. The voiceprint recognition method according to claim 2, characterized in that, The step of obtaining the supervised loss based on the first posterior probability and the second posterior probability of all the first audio features includes: Obtain a third sub-loss based on the first posterior probability of each of the first audio samples and the corresponding speaker label, and obtain a fourth sub-loss based on the second posterior probability of each of the first audio samples and the corresponding speaker label; Obtain a second sum value of the third sub-loss and the fourth sub-loss of all the first audio samples, and use the second ratio of the second sum value to the number of the first audio samples as the supervised loss.

5. The voiceprint recognition method according to claim 2, characterized in that, After the step of outputting one of the first speaker verification network and the second speaker verification network as the trained speaker verification network in response to reaching the stop training condition, it includes: Construct a third sample set; wherein, the third sample set contains a plurality of third audio samples, and each of the third audio samples is set with a speaker label, and at least one of the third audio samples corresponds to the same speaker label; Use the trained speaker verification network to screen out a subset from the third sample set; wherein, the subset contains a plurality of third audio samples with a similarity exceeding a threshold and having different speaker labels; Adjust the trained speaker verification network based on the loss corresponding to the subset.

6. The voiceprint recognition method according to claim 5, characterized in that, The step of using the trained speaker verification network to screen out a subset from the third sample set includes: Use the trained speaker verification network to obtain a class center matrix corresponding to each speaker label in the third sample set; Obtain a similarity matrix between the class center matrix of each speaker label in the third sample set and the class center matrices of the remaining speaker labels; Randomly select at least one speaker label from the third sample set as a first target label, and obtain at least one remaining speaker label with a relatively large similarity value from the similarity matrix of each first target label as a second target label; For each of the second target labels, randomly select a third audio sample with the second target label and add it to the subset.

7. The voiceprint recognition method according to claim 5, characterized in that, The step of constructing the third sample set includes: Obtain a plurality of third audio samples, and each of the third audio samples is set with a speaker label; Adjust the pitch of at least some of the third audio samples to form new third audio samples, and the speaker label of the adjusted third audio sample is different from that of the third audio sample before adjustment; Construct the third sample set based on all the third audio samples before adjustment and all the third audio samples after adjustment.

8. A training method for a voiceprint extraction network, characterized in that, Include: Construct a first sample set and a second sample set; wherein, the first sample set contains a plurality of first audio samples, and each of the first audio samples is set with a speaker label; the second sample set contains a plurality of second audio samples, and each of the second audio samples is not set with a speaker label; Perform feature extraction on the plurality of first audio samples respectively to obtain a plurality of first audio features, and perform feature extraction on the plurality of second audio samples respectively to obtain a plurality of second audio features; Input multiple of the first audio features and multiple of the second audio features into a first voiceprint extraction network to obtain a first posterior probability corresponding to each audio feature, and input multiple of the first audio features and multiple of the second audio features into a second voiceprint extraction network to obtain a second posterior probability corresponding to each audio feature; Obtain a supervised loss based on the first posterior probabilities and the second posterior probabilities of all the first audio features, and obtain a semi-supervised loss based on the first posterior probabilities and the second posterior probabilities of all the second audio features; Obtain a total loss based on the supervised loss and the semi-supervised loss, and respectively adjust the parameters in the first voiceprint extraction network and the second voiceprint extraction network based on the total loss; In response to reaching a stop training condition, output one of the first voiceprint extraction network and the second voiceprint extraction network as the trained voiceprint extraction network.

9. A voiceprint recognition device, characterized in that, Comprising: A first extraction module, configured to perform feature extraction on an audio to be recognized to obtain audio features to be recognized; A second extraction module, connected to the first extraction module, configured to input the audio features to be recognized into the trained voiceprint extraction network to obtain voiceprint features to be recognized; wherein, when training the voiceprint extraction network, it is based on speech data with speaker labels and speech data without speaker labels, and the total loss used when training the voiceprint extraction network is related to the supervised loss and the semi-supervised loss, and the supervised loss and the semi-supervised loss are obtained based on the following steps: construct a first sample set and a second sample set; wherein, the first sample set contains multiple first audio samples, and each of the first audio samples is set with a speaker label; the second sample set contains multiple second audio samples, and each of the second audio samples is not set with a speaker label; perform feature extraction on multiple of the first audio samples respectively to obtain multiple first audio features, and perform feature extraction on multiple of the second audio samples respectively to obtain multiple second audio features; input multiple of the first audio features and multiple of the second audio features into a first voiceprint extraction network to obtain a first posterior probability corresponding to each audio feature, and input multiple of the first audio features and multiple of the second audio features into a second voiceprint extraction network to obtain a second posterior probability corresponding to each audio feature; obtain the supervised loss based on the first posterior probabilities and the second posterior probabilities of all the first audio features, and obtain the semi-supervised loss based on the first posterior probabilities and the second posterior probabilities of all the second audio features; the voiceprint extraction network is one of the first voiceprint extraction network and the second voiceprint extraction network; A determination module, connected to the second extraction module, configured to determine the speaker corresponding to the audio to be recognized based on the voiceprint features to be recognized.

10. An electronic device, characterized in that, Comprising a mutually coupled memory and a processor, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the voiceprint recognition method according to any one of claims 1 to 7 or the training method of the voiceprint extraction network according to claim 8.

11. A storage device, characterized in that, Stored with program instructions that can be run by a processor, the program instructions are used to implement the voiceprint recognition method described in any one of claims 1 to 7 or the training method of the voiceprint extraction network described in claim 8.

Citation Information

Patent Citations

  • Model generation method, voiceprint recognition method and corresponding device

    CN110838295A

  • Voiceprint recognition system, method and device and electronic equipment

    CN111462760A