Voiceprint Extraction Method, Device, Equipment and Readable Storage Medium
By convolution processing and environmental prediction of the speech spectrum fragments of speech data, high-dimensional feature vectors are generated, which solves the accuracy of the voiceprint extraction model under different recording environments, and achieves higher voiceprint information accuracy.
Patent Information
- Application Number
- CN202210616862.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-06-01
AI Technical Summary
In the prior art, the accuracy of voiceprint extraction models of voiceprint data under different recording environments has decreased, and the impact of environmental differences on voiceprint information cannot be effectively eliminated.
By obtaining the spectral fragments of speech data, performing convolution processing and environmental prediction, the depth feature map and environmental prediction posterior probability are obtained, and statistical pooling is performed in combination with the frame-level environmental attention weight coefficient to generate high-dimensional feature vectors, and finally a voiceprint representation vector with fused recording environment information is obtained.
The accuracy of the voiceprint representation vector of voiceprint data is improved, effectively eliminating the impact of recording environment differences on voiceprint information.
Smart Images

Figure CN115019808B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of voice data processing, and more specifically, to a voiceprint extraction method, device, equipment and readable storage medium. Background Art
[0002] Voiceprint recognition and authentication is one of the key technologies in the field of biometric authentication. By inputting two pieces of voice data into a voiceprint extraction model, the voiceprint extraction model extracts the voiceprint information of the two pieces of voice data, and then uses the similarity of the voiceprint information of the two pieces of voice data to obtain an authentication result.
[0003] Currently, most deep neural networks, such as Convolutional Neural Networks (CNNs), are trained based on training voices to obtain a voiceprint extraction model. For the voice data to be extracted for voiceprint information, input it into the voiceprint extraction model, and the voiceprint extraction model can output the voiceprint feature vector of the voice data. However, if the voice data to be extracted for voiceprint information is not recorded in the same environment as the training voice data, it will cause the accuracy of the voiceprint feature vector output by the voiceprint extraction model to decrease.
[0004] Therefore, how to provide a voiceprint extraction method to eliminate the influence of the voice data recording environment difference on the accuracy of the voiceprint information has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0005] In view of the above problems, this application proposes a voiceprint extraction method, device, equipment and readable storage medium. The specific solutions are as follows:
[0006] A voiceprint extraction method, the method includes:
[0007] Obtain the voice data to be extracted for voiceprint;
[0008] Determine the spectrogram segment corresponding to the voice data;
[0009] For each spectrogram segment, perform voiceprint extraction on the spectrogram segment to obtain the voiceprint feature vector of the spectrogram segment; the voiceprint feature vector is fused with the recording environment information of the voice data;
[0010] Based on the voiceprint feature vectors of each spectrogram segment, obtain the voiceprint feature vector of the voice data.
[0011] Optionally, performing voiceprint extraction on the spectrogram segment to obtain the voiceprint feature vector of the spectrogram segment includes:
[0012] Perform convolution processing on the spectrogram segment to obtain the depth feature map of the spectrogram segment;
[0013] Perform environmental prediction on the depth feature map of the spectrogram segment to obtain the posterior probability of environmental prediction;
[0014] Based on the depth feature map of the spectrogram segment and the posterior probability of environmental prediction, obtain the high-dimensional feature vector of the spectrogram segment;
[0015] Perform voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain the voiceprint characterization vector of the spectrogram segment.
[0016] Optionally, the obtaining the high-dimensional feature vector of the spectrogram segment based on the depth feature map of the spectrogram segment and the posterior probability of environmental prediction includes:
[0017] Perform a linear transformation on the posterior probability of environmental prediction to obtain a linear transformation parameter;
[0018] Based on the linear transformation parameter, calculate the frame-level environmental attention weight coefficient of the spectrogram segment;
[0019] Perform statistical pooling on the depth feature map of the spectrogram segment and the frame-level environmental attention weight coefficient of the spectrogram segment to obtain the high-dimensional feature vector of the spectrogram segment.
[0020] Optionally, the performing statistical pooling on the depth feature map of the spectrogram segment and the frame-level environmental attention weight coefficient of the spectrogram segment to obtain the high-dimensional feature vector of the spectrogram segment includes:
[0021] Perform statistical pooling on the depth feature map of the spectrogram segment and the frame-level environmental attention weight coefficient of the spectrogram segment to obtain the mean and standard variance corresponding to the spectrogram segment;
[0022] Concatenate the mean and standard variance corresponding to the spectrogram segment to obtain the high-dimensional feature vector of the spectrogram segment.
[0023] Optionally, the performing voiceprint extraction on the spectrogram segment to obtain the voiceprint characterization vector of the spectrogram segment includes:
[0024] Input the spectrogram segment into a voiceprint extraction model, the voiceprint extraction model performs convolution processing on the spectrogram segment to obtain the depth feature map of the spectrogram segment; perform environmental prediction on the depth feature map of the spectrogram segment to obtain the posterior probability of environmental prediction; based on the depth feature map of the spectrogram segment and the posterior probability of environmental prediction, obtain the high-dimensional feature vector of the spectrogram segment; perform voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain the voiceprint characterization vector of the spectrogram segment.
[0025] Optionally, the voiceprint extraction model is trained as follows:
[0026] Obtain a training data set and a pre-constructed voiceprint extraction model. The training data set includes multiple spectrogram segments, and each spectrogram segment is labeled with an environment label and a speaker label;
[0027] Determine multiple training data subsets based on the training data set. Each training data subset includes spectrogram segments of different speakers in different environments, spectrogram segments of different speakers in the same environment, and spectrogram segments of the same speaker in different environments;
[0028] Each time of training, input a training data subset into the voiceprint extraction model, train the voiceprint extraction model, and obtain the loss function of the voiceprint extraction model for this training;
[0029] When the loss function converges, determine that the training of the voiceprint extraction model is completed.
[0030] Optionally, the loss function of the voiceprint extraction model is composed of an environment prediction loss, a speaker prediction loss, a triplet loss in the same environment, and a triplet loss in different environments;
[0031] The environment prediction loss is used to characterize the error between the environment label predicted based on the spectrogram segment and the environment label annotated for the spectrogram segment. The speaker prediction loss is used to characterize the error between the speaker label predicted based on the spectrogram segment and the speaker label annotated for the spectrogram segment. The triplet loss in the same environment is used to characterize the error between the voiceprint representation vectors of three spectrogram segments in the same environment. The triplet loss in different environments is used to characterize the error between the voiceprint representation vectors of three spectrogram segments in different environments.
[0032] A voiceprint extraction device, the device includes:
[0033] An acquisition unit, configured to acquire voice data to be subjected to voiceprint extraction;
[0034] A spectrogram segment determination unit, configured to determine the spectrogram segment corresponding to the voice data;
[0035] A voiceprint representation vector determination unit for spectrogram segments, configured to perform voiceprint extraction on each spectrogram segment to obtain the voiceprint representation vector of the spectrogram segment; the voiceprint representation vector is fused with the recording environment information of the voice data;
[0036] A voiceprint representation vector determination unit for voice data, configured to obtain the voiceprint representation vector of the voice data based on the voiceprint representation vectors of each spectrogram segment.
[0037] Optionally, the voiceprint characterization vector determination unit of the spectrogram segment includes:
[0038] A convolution processing unit, configured to perform convolution processing on the spectrogram segment to obtain a depth feature map of the spectrogram segment;
[0039] An environment prediction unit, configured to perform environment prediction on the depth feature map of the spectrogram segment to obtain an environment prediction posterior probability;
[0040] A high-dimensional feature vector determination unit, configured to obtain a high-dimensional feature vector of the spectrogram segment based on the depth feature map of the spectrogram segment and the environment prediction posterior probability;
[0041] A voiceprint extraction unit, configured to perform voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain a voiceprint characterization vector of the spectrogram segment.
[0042] Optionally, the high-dimensional feature vector determination unit includes:
[0043] A linear transformation unit, configured to perform linear transformation on the environment prediction posterior probability to obtain a linear transformation parameter;
[0044] A calculation unit, configured to calculate a frame-level environment attention weight coefficient of the spectrogram segment based on the linear transformation parameter;
[0045] A statistical pooling processing unit, configured to perform statistical pooling processing on the depth feature map of the spectrogram segment and the frame-level environment attention weight coefficient of the spectrogram segment to obtain a high-dimensional feature vector of the spectrogram segment.
[0046] Optionally, the statistical pooling processing unit includes:
[0047] A mean and standard deviation determination unit, configured to perform statistical pooling processing on the depth feature map of the spectrogram segment and the frame-level environment attention weight coefficient of the spectrogram segment to obtain a mean and a standard deviation corresponding to the spectrogram segment;
[0048] A splicing unit, configured to splice the mean and the standard deviation corresponding to the spectrogram segment to obtain a high-dimensional feature vector of the spectrogram segment.
[0049] Optionally, the voiceprint characterization vector determination unit of the spectrogram segment is specifically configured to:
[0050] Input the spectrogram segment into a voiceprint extraction model. The voiceprint extraction model performs convolutional processing on the spectrogram segment to obtain a depth feature map of the spectrogram segment; perform environmental prediction on the depth feature map of the spectrogram segment to obtain an environmental prediction posterior probability; based on the depth feature map of the spectrogram segment and the environmental prediction posterior probability, obtain a high-dimensional feature vector of the spectrogram segment; perform voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain a voiceprint representation vector of the spectrogram segment.
[0051] Optionally, the training method of the voiceprint extraction model is as follows:
[0052] Obtain a training data set and a pre-constructed voiceprint extraction model. The training data set includes multiple spectrogram segments, and each spectrogram segment is labeled with an environmental label and a speaker label;
[0053] Based on the training data set, determine multiple training data subsets. Each training data subset includes spectrogram segments of different speakers in different environments, spectrogram segments of different speakers in the same environment, and spectrogram segments of the same speaker in different environments;
[0054] Each time of training, input a training data subset into the voiceprint extraction model, train the voiceprint extraction model, and obtain the loss function of the voiceprint extraction model for this training;
[0055] When the loss function converges, determine that the training of the voiceprint extraction model is completed.
[0056] Optionally, the loss function of the voiceprint extraction model is composed of an environmental prediction loss, a speaker prediction loss, a triplet loss in the same environment, and a triplet loss in different environments;
[0057] The environmental prediction loss is used to represent the error between the environmental label predicted based on the spectrogram segment and the environmental label annotated for the spectrogram segment. The speaker prediction loss is used to represent the error between the speaker label predicted based on the spectrogram segment and the speaker label annotated for the spectrogram segment. The triplet loss in the same environment is used to represent the error between the voiceprint representation vectors of three spectrogram segments in the same environment. The triplet loss in different environments is used to represent the error between the voiceprint representation vectors of three spectrogram segments in different environments.
[0058] A voiceprint extraction device includes a memory and a processor;
[0059] The memory is used to store programs;
[0060] The processor is used to execute the programs to implement each step of the voiceprint extraction method as described above.
[0061] A readable storage medium stores a computer program thereon. It is characterized in that when the computer program is executed by a processor, each step of the voiceprint extraction method described above is implemented.
[0062] By means of the above technical solution, the present application discloses a voiceprint extraction method, device, equipment and readable storage medium. After obtaining the voice data to be subjected to voiceprint extraction, first determine the spectrogram segment corresponding to the voice data, and then for each spectrogram segment, perform voiceprint extraction on the spectrogram segment to obtain a voiceprint feature vector of the spectrogram segment fused with the recording environment information of the voice data; perform weighted averaging on the voiceprint feature vectors of each spectrogram segment fused with the environment information to obtain a voiceprint feature vector of the voice data fused with the recording environment information of the voice data. In the above solution, the voiceprint feature vector of the voice data is fused with the recording environment information of the voice data, and its accuracy is higher. Therefore, adopting the above solution can eliminate the influence of the difference in the recording environment of the voice data on the accuracy of the voiceprint information. Description of the Drawings
[0063] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0064] Figure 1 It is a schematic flowchart of a voiceprint extraction method disclosed in an embodiment of the present application;
[0065] Figure 2 It is a schematic structural diagram of a voiceprint extraction model disclosed in an embodiment of the present application;
[0066] Figure 3 It is a schematic flowchart of a method for performing voiceprint extraction on a spectrogram segment to obtain a voiceprint feature vector of the spectrogram segment disclosed in an embodiment of the present application;
[0067] Figure 4 It is a schematic flowchart of a method for training a voiceprint extraction model disclosed in an embodiment of the present application;
[0068] Figure 5 It is a schematic structural diagram of a voiceprint extraction device disclosed in an embodiment of the present application;
[0069] Figure 6 It is a hardware structure block diagram of a voiceprint extraction device disclosed in an embodiment of the present application. Detailed Embodiments
[0070] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0071] Next, a voiceprint extraction method provided by the present application will be introduced through the following embodiments.
[0072] Referring to Figure 1 , Figure 1 which is a schematic flowchart of a voiceprint extraction method disclosed in an embodiment of the present application. The method may include:
[0073] Step S101: Obtain voice data to be subjected to voiceprint extraction.
[0074] In the present application, the voice data to be subjected to voiceprint extraction may be voice data recorded in any environment.
[0075] Step S102: Determine the spectrogram segment corresponding to the voice data.
[0076] In the present application, after windowing and Fourier transform are performed on the voice data, acoustic features such as Mel-frequency cepstral coefficients (MFCC) or filterbank features can be obtained. The acoustic features are used to form a spectrogram, which is segmented according to the window length l to obtain N spectrogram segments {Seg 1 , Seg 2 ,..., Seg N}, and the size of each segment is l×d. It should be noted that if the voice data is less than the window length l, the voice data is copied several times, and the voice data with the window length l is retained, and the excess is discarded.
[0077] It should be noted that in the present application, filterbank features can be used as acoustic features because, for voiceprint recognition methods based on deep neural networks (such as CNNs, etc.), filterbank features have the best effect.
[0078] In addition, it should be noted that in this application, if l is set too small, the original spectrogram will be fragmented, and the coherent spectrogram information will be cut into several small spectrogram segments, resulting in excessive loss of information between spectrogram segments and making it impossible to model the long-term correlation of speech. If l is set too large, it will affect the network training efficiency, and at the same time, the GPU (Graphics Processing Unit) training resources will be significantly increased. As an implementable way, l can be set to 1 / 2 of the average effective duration of each speech data in the training dataset.
[0079] Step S103: For each spectrogram segment, perform voiceprint extraction on the spectrogram segment to obtain a voiceprint characterization vector of the spectrogram segment; the voiceprint characterization vector is fused with the recording environment information of the speech data.
[0080] In this application, the voiceprint extraction of spectrogram segments can be implemented by using a deep neural network model to obtain the voiceprint characterization vector of the spectrogram segment. When this neural network model is trained, the recording environment information of each speech data in the training dataset needs to be considered. The specific implementation will be described in detail in the following embodiments and will not be elaborated here.
[0081] Step S104: Based on the voiceprint characterization vectors of each spectrogram segment, obtain the voiceprint characterization vector of the speech data.
[0082] As an implementable way, the voiceprint characterization vectors of each spectrogram segment can be averaged, or weighted averaged, to obtain the voiceprint characterization vector of the speech data.
[0083] This embodiment discloses a voiceprint extraction method. After obtaining the speech data to be subjected to voiceprint extraction, first determine the spectrogram segments corresponding to the speech data, and then for each spectrogram segment, perform voiceprint extraction on the spectrogram segment to obtain a voiceprint characterization vector of the spectrogram segment that is fused with the recording environment information of the speech data; perform weighted averaging on the voiceprint characterization vectors of each spectrogram segment that are fused with the environment information to obtain a voiceprint characterization vector of the speech data that is fused with the recording environment information of the speech data. In the above method, the voiceprint characterization vector of the speech data is fused with the recording environment information of the speech data, and its accuracy is higher. Therefore, the above method can eliminate the influence of the difference in the recording environment of the speech data on the accuracy of the voiceprint information.
[0084] It should be noted that in this application, a voiceprint extraction model can be constructed based on a deep neural network model, and then the voiceprint extraction model is trained, and then the above step S103 is executed based on the trained voiceprint extraction model.
[0085] Refer to Figure 2 , Figure 2A structural schematic diagram of a voiceprint extraction model disclosed in an embodiment of the present application. The voiceprint extraction model includes a TDNN (Time-Delay Neural Network) module, a two-layer fully connected module, a statistical pooling module, and a one-layer fully connected module. Then, performing voiceprint extraction on the spectrogram segment to obtain the voiceprint characterization vector of the spectrogram segment includes: inputting the spectrogram segment into the voiceprint extraction model, and the voiceprint extraction model performing convolution processing on the spectrogram segment to obtain the depth feature map of the spectrogram segment; performing environment prediction on the depth feature map of the spectrogram segment to obtain the environmental prediction posterior probability; obtaining the high-dimensional feature vector of the spectrogram segment based on the depth feature map of the spectrogram segment and the environmental prediction posterior probability; performing voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain the voiceprint characterization vector of the spectrogram segment.
[0086] Based on Figure 2 The voiceprint extraction model shown, in another embodiment of the present application, the specific implementation manner of performing voiceprint extraction on the spectrogram segment in step S103 to obtain the voiceprint characterization vector of the spectrogram segment is described. Refer to Figure 3 , Figure 3 A flowchart of a method for performing voiceprint extraction on a spectrogram segment to obtain the voiceprint characterization vector of the spectrogram segment disclosed in an embodiment of the present application. The method may include the following steps:
[0087] Step S201: Performing convolution processing on the spectrogram segment to obtain the depth feature map of the spectrogram segment.
[0088] In the present application, the spectrogram segment may be input into the voiceprint extraction model, and the TDNN module in the voiceprint extraction model performs multiple one-dimensional convolution processes on the voice segment to obtain the depth feature map of the spectrogram segment.
[0089] Step S202: Performing environment prediction on the depth feature map of the spectrogram segment to obtain the environmental prediction posterior probability.
[0090] In the present application, the two-layer fully connected module in the voiceprint extraction model may use the softmax criterion to perform environment prediction on the depth feature map of the spectrogram segment to obtain the environmental prediction posterior probability. It should be noted that the environmental prediction posterior probability is at the frame level and can be represented by p n to represent p n (n = 1, 2,..., l0), where l0 represents the number of frames.
[0091] Step S203: Obtaining the high-dimensional feature vector of the spectrogram segment based on the depth feature map of the spectrogram segment and the environmental prediction posterior probability.
[0092] In the present application, a high-dimensional feature vector of the spectral segment can be obtained by a statistical pooling module in the voiceprint extraction model based on the depth feature map of the spectral segment and the environmental prediction posterior probability.
[0093] As an implementable manner, the process of obtaining the high-dimensional feature vector of the spectral segment based on the depth feature map of the spectral segment and the environmental prediction posterior probability may include: performing a linear transformation on the environmental prediction posterior probability to obtain a linear transformation parameter; calculating a frame-level environmental attention weight coefficient of the spectral segment based on the linear transformation parameter; and performing a statistical pooling process on the depth feature map of the spectral segment and the frame-level environmental attention weight coefficient of the spectral segment to obtain the high-dimensional feature vector of the spectral segment.
[0094] Specifically, the performing a statistical pooling process on the depth feature map of the spectral segment and the frame-level environmental attention weight coefficient of the spectral segment to obtain the high-dimensional feature vector of the spectral segment includes: performing a statistical pooling process on the depth feature map of the spectral segment and the frame-level environmental attention weight coefficient of the spectral segment to obtain the mean value and standard variance corresponding to the spectral segment; and splicing the mean value and standard variance corresponding to the spectral segment to obtain the high-dimensional feature vector of the spectral segment.
[0095] For ease of understanding, assume that the depth feature map of the spectral segment is The environmental prediction posterior probability of the spectral segment is p n (n = 1, 2,..., l0), perform a linear transformation on p n (n = 1, 2,..., l0): e n = Wp n + B, where W represents a linear transformation projection matrix, B represents a bias, and is a parameter of the model;
[0096] Calculate the frame-level environmental attention weight coefficient of the spectral segment based on the following formula:
[0097]
[0098] After performing a statistical pooling (Static Pooling) process on the depth feature map of the spectral segment based on the frame-level environmental attention weight coefficient of the spectral segment, obtain the mean value μ i corresponding to the spectral segment and the standard variance σ i corresponding to the spectral segment:
[0099]
[0100] where diag represents taking diagonal elements.
[0101] Concatenate the mean μ i and the standard deviation σ i to form the high-dimensional feature vector h of the spectrogram segment i .
[0102] Step S204: Perform voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain the voiceprint characterization vector of the spectrogram segment.
[0103] In this application, the high-dimensional feature vector h i can be fed into a fully connected layer to perform voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain the voiceprint characterization vector of the spectrogram segment.
[0104] In another embodiment of this application, the training method of the voiceprint extraction model is described. Refer to Figure 4 , Figure 4 which is a schematic flow chart of a training method for a voiceprint extraction model disclosed in an embodiment of this application, and it may include the following steps:
[0105] Step 301: Obtain a training data set and a pre-constructed voiceprint extraction model. The training data set includes multiple spectrogram segments, and each spectrogram segment is labeled with an environment label and a speaker label.
[0106] It should be noted that the spectrogram segments in the training data set can be obtained in the manner described in step S102. For specific details, please refer to the relevant description of step S102, which will not be elaborated here.
[0107] Step 302: Determine multiple training data subsets based on the training data set.
[0108] It should be noted that each training data subset includes spectrogram segments of different speakers in different environments, spectrogram segments of different speakers in the same environment, and spectrogram segments of the same speaker in different environments.
[0109] Step 303: Each time of training, input a training data subset into the voiceprint extraction model, train the voiceprint extraction model, and obtain the loss function of the voiceprint extraction model for this training.
[0110] In this application, the loss function of the voiceprint extraction model is composed of an environment prediction loss, a speaker prediction loss, a triplet loss in the same environment, and a triplet loss in different environments;
[0111] Among them, the environmental prediction loss is used to represent the error between the environmental label predicted based on the speech spectrum segment and the environmental label annotated for the speech spectrum segment, the speaker prediction loss is used to represent the error between the speaker label predicted based on the speech spectrum segment and the speaker label annotated for the speech spectrum segment, the triplet loss in the same environment is used to represent the error between the voiceprint representation vectors of three speech spectrum segments in the same environment, and the triplet loss in different environments is used to represent the error between the voiceprint representation vectors of three speech spectrum segments in different environments.
[0112] As an implementable manner, the environmental prediction loss and the speaker prediction loss can be calculated using the cross-entropy (CE) criterion. The triplet loss in the same environment can be calculated using the following formula L s = D(w as w ps - w as w ns + m s ), where w as and w ps represent different voiceprint representation vectors of the same speaker in the same environment, w as and w ns represent the voiceprint representation vectors of different speakers in the same environment, m s represents the distance metric boundary in the same environment; the triplet loss in different environments can be calculated using the following formula L d = D(w as w ps - w as w nd + m d ), where w as and w ps represent different voiceprint representation vectors of the same speaker in the same environment, w as and w nd represent the voiceprint representation vectors of different speakers in different environments, m d represents the distance metric boundary in different environments; D() represents the cosine distance between vectors.
[0113] Step 304: When the loss function converges, determine that the voiceprint extraction model training is completed.
[0114] Next, a voiceprint extraction device disclosed in an embodiment of the present application will be described. The voiceprint extraction device described below can be correspondingly referred to the voiceprint extraction method described above.
[0115] Refer to Figure 5 , Figure 5 which is a schematic structural diagram of a voiceprint extraction device disclosed in an embodiment of the present application. AsFigure 5 As shown in Figure 5 , the voiceprint extraction device may include:
[0116] An acquisition unit 11, configured to acquire voice data to be subjected to voiceprint extraction;
[0117] A spectrogram segment determination unit 12, configured to determine a spectrogram segment corresponding to the voice data;
[0118] A voiceprint characterization vector determination unit 13 for the spectrogram segment, configured to perform voiceprint extraction on each spectrogram segment to obtain a voiceprint characterization vector of the spectrogram segment; the voiceprint characterization vector is fused with recording environment information of the voice data;
[0119] A voiceprint characterization vector determination unit 14 for the voice data, configured to obtain a voiceprint characterization vector of the voice data based on voiceprint characterization vectors of respective spectrogram segments.
[0120] As an implementable manner, the voiceprint characterization vector determination unit for the spectrogram segment includes:
[0121] A convolution processing unit, configured to perform convolution processing on the spectrogram segment to obtain a depth feature map of the spectrogram segment;
[0122] An environment prediction unit, configured to perform environment prediction on the depth feature map of the spectrogram segment to obtain an environment prediction posterior probability;
[0123] A high-dimensional feature vector determination unit, configured to obtain a high-dimensional feature vector of the spectrogram segment based on the depth feature map of the spectrogram segment and the environment prediction posterior probability;
[0124] A voiceprint extraction unit, configured to perform voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain a voiceprint characterization vector of the spectrogram segment.
[0125] As an implementable manner, the high-dimensional feature vector determination unit includes:
[0126] A linear transformation unit, configured to perform linear transformation on the environment prediction posterior probability to obtain a linear transformation parameter;
[0127] A calculation unit, configured to calculate a frame-level environment attention weight coefficient of the spectrogram segment based on the linear transformation parameter;
[0128] A statistical pooling processing unit, configured to perform statistical pooling processing on the depth feature map of the spectrogram segment and the frame-level environment attention weight coefficient of the spectrogram segment to obtain a high-dimensional feature vector of the spectrogram segment.
[0129] As an implementable manner, the statistical pooling processing unit includes:
[0130] A mean and standard deviation determination unit for performing statistical pooling on the depth feature map of the spectrogram segment and the frame-level environmental attention weight coefficient of the spectrogram segment to obtain the mean and standard deviation corresponding to the spectrogram segment;
[0131] A splicing unit for splicing the mean and standard deviation corresponding to the spectrogram segment to obtain a high-dimensional feature vector of the spectrogram segment.
[0132] As an implementable manner, the voiceprint characterization vector determination unit of the spectrogram segment is specifically used for:
[0133] Inputting the spectrogram segment into a voiceprint extraction model, the voiceprint extraction model performing convolution processing on the spectrogram segment to obtain the depth feature map of the spectrogram segment; performing environmental prediction on the depth feature map of the spectrogram segment to obtain the environmental prediction posterior probability; obtaining the high-dimensional feature vector of the spectrogram segment based on the depth feature map of the spectrogram segment and the environmental prediction posterior probability; and performing voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain the voiceprint characterization vector of the spectrogram segment.
[0134] As an implementable manner, the training method of the voiceprint extraction model is as follows:
[0135] Obtaining a training data set and a pre-constructed voiceprint extraction model, where the training data set includes multiple spectrogram segments, and each spectrogram segment is labeled with an environmental label and a speaker label;
[0136] Determining multiple training data subsets based on the training data set, where each training data subset includes spectrogram segments of different speakers in different environments, spectrogram segments of different speakers in the same environment, and spectrogram segments of the same speaker in different environments;
[0137] Each time of training, inputting a training data subset into the voiceprint extraction model, training the voiceprint extraction model, and obtaining the loss function of the voiceprint extraction model for this training;
[0138] When the loss function converges, it is determined that the training of the voiceprint extraction model is completed.
[0139] As an implementable manner, the loss function of the voiceprint extraction model is composed of an environmental prediction loss, a speaker prediction loss, a triplet loss in the same environment, and a triplet loss in different environments;
[0140] The environmental prediction loss is used to characterize the error between the environmental label predicted based on the spectrogram segment and the environmental label annotated for the spectrogram segment. The speaker prediction loss is used to characterize the error between the speaker label predicted based on the spectrogram segment and the speaker label annotated for the spectrogram segment. The triplet loss in the same environment is used to characterize the error between the voiceprint feature vectors of three spectrogram segments in the same environment. The triplet loss in different environments is used to characterize the error between the voiceprint feature vectors of three spectrogram segments in different environments.
[0141] Referring to Figure 6 , Figure 6 FIG. is a hardware structure block diagram of a voiceprint extraction device provided by an embodiment of the present application. Referring to Figure 6 , the hardware structure of voiceprint extraction may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0142] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete communication with each other through the communication bus 4;
[0143] The processor 1 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0144] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;
[0145] Wherein, the memory stores a program, and the processor can call the program stored in the memory. The program is used for:
[0146] Obtaining voice data to be subjected to voiceprint extraction;
[0147] Determining the spectrogram segment corresponding to the voice data;
[0148] For each spectrogram segment, performing voiceprint extraction on the spectrogram segment to obtain a voiceprint feature vector of the spectrogram segment; the voiceprint feature vector is fused with the recording environment information of the voice data;
[0149] Based on the voiceprint feature vectors of each spectrogram segment, obtaining a voiceprint feature vector of the voice data.
[0150] Optionally, the refined functions and extended functions of the program may refer to the above description.
[0151] An embodiment of the present application further provides a readable storage medium, which can store a program suitable for a processor to execute. The program is used for:
[0152] Obtain voice data to be subjected to voiceprint extraction;
[0153] Determine the spectrogram segment corresponding to the voice data;
[0154] For each spectrogram segment, perform voiceprint extraction on the spectrogram segment to obtain a voiceprint characterization vector of the spectrogram segment; the voiceprint characterization vector is fused with the recording environment information of the voice data;
[0155] Based on the voiceprint characterization vectors of each spectrogram segment, obtain the voiceprint characterization vector of the voice data.
[0156] Optionally, the refined functions and extended functions of the program can be referred to the above description.
[0157] Finally, it should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0158] Each embodiment in this specification is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same and similar parts between each embodiment can be referred to each other.
[0159] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voiceprint extraction method, characterized in that, The method includes: Obtaining speech data to be subjected to voiceprint extraction; Determining the spectrogram segment corresponding to the speech data; For each spectrogram segment, performing voiceprint extraction on the spectrogram segment to obtain a voiceprint feature vector of the spectrogram segment; the voiceprint feature vector is fused with the recording environment information of the speech data; Based on the voiceprint feature vectors of each spectrogram segment, obtaining a voiceprint feature vector of the speech data; Wherein, the performing voiceprint extraction on the spectrogram segment to obtain a voiceprint feature vector of the spectrogram segment includes: Performing convolution processing on the spectrogram segment to obtain a depth feature map of the spectrogram segment; Performing environment prediction on the depth feature map of the spectrogram segment to obtain an environment prediction posterior probability; Based on the depth feature map of the spectrogram segment and the environment prediction posterior probability, obtaining a high-dimensional feature vector of the spectrogram segment; Performing voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain a voiceprint feature vector of the spectrogram segment.
2. The method according to claim 1, characterized in that, The obtaining a high-dimensional feature vector of the spectrogram segment based on the depth feature map of the spectrogram segment and the environment prediction posterior probability includes: Performing a linear transformation on the environment prediction posterior probability to obtain a linear transformation parameter; Based on the linear transformation parameter, calculating a frame-level environment attention weight coefficient of the spectrogram segment; Performing statistical pooling processing on the depth feature map of the spectrogram segment and the frame-level environment attention weight coefficient of the spectrogram segment to obtain a high-dimensional feature vector of the spectrogram segment.
3. The method according to claim 2, wherein The performing statistical pooling processing on the depth feature map of the spectrogram segment and the frame-level environment attention weight coefficient of the spectrogram segment to obtain a high-dimensional feature vector of the spectrogram segment includes: Performing statistical pooling processing on the depth feature map of the spectrogram segment and the frame-level environment attention weight coefficient of the spectrogram segment to obtain the mean value and standard variance corresponding to the spectrogram segment; Concatenating the mean value and standard variance corresponding to the spectrogram segment to obtain a high-dimensional feature vector of the spectrogram segment.
4. The method according to claim 1, characterized in that, The performing voiceprint extraction on the spectrogram segment to obtain a voiceprint feature vector of the spectrogram segment includes: Inputting the spectrogram segment into a voiceprint extraction model, the voiceprint extraction model performing convolution processing on the spectrogram segment to obtain a depth feature map of the spectrogram segment; performing environment prediction on the depth feature map of the spectrogram segment to obtain an environment prediction posterior probability; based on the depth feature map of the spectrogram segment and the environment prediction posterior probability, obtaining a high-dimensional feature vector of the spectrogram segment; performing voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain a voiceprint feature vector of the spectrogram segment.
5. The method according to claim 4, characterized in that The training method of the voiceprint extraction model is as follows: Obtaining a training data set and a pre-constructed voiceprint extraction model, the training data set including a plurality of spectrogram segments, and each spectrogram segment being labeled with an environment label and a speaker label; Based on the training data set, determining a plurality of training data subsets, each training data subset including spectrogram segments of different speakers in different environments, spectrogram segments of different speakers in the same environment, and spectrogram segments of the same speaker in different environments; In each training, a subset of training data is input into the voiceprint extraction model, the voiceprint extraction model is trained, and the loss function of the voiceprint extraction model for this training is obtained. When the loss function converges, it is determined that the training of the voiceprint extraction model is completed.
6. The method according to claim 5, wherein The loss function of the voiceprint extraction model is composed of an environment prediction loss, a speaker prediction loss, a triplet loss in the same environment, and a triplet loss in different environments. The environment prediction loss is used to characterize the error between the environment label predicted based on the spectrogram segment and the environment label annotated for the spectrogram segment. The speaker prediction loss is used to characterize the error between the speaker label predicted based on the spectrogram segment and the speaker label annotated for the spectrogram segment. The triplet loss in the same environment is used to characterize the error between the voiceprint feature vectors of three spectrogram segments in the same environment. The triplet loss in different environments is used to characterize the error between the voiceprint feature vectors of three spectrogram segments in different environments.
7. A voiceprint extraction device, characterized in that, The device includes: An acquisition unit, configured to acquire voice data to be subjected to voiceprint extraction. A spectrogram segment determination unit, configured to determine the spectrogram segment corresponding to the voice data. A voiceprint feature vector determination unit for spectrogram segments, configured to perform voiceprint extraction on each spectrogram segment to obtain the voiceprint feature vector of the spectrogram segment; the voiceprint feature vector is fused with the recording environment information of the voice data. A voiceprint feature vector determination unit for voice data, configured to obtain the voiceprint feature vector of the voice data based on the voiceprint feature vectors of the respective spectrogram segments. Among them, the voiceprint feature vector determination unit for spectrogram segments is specifically configured to perform convolution processing on the spectrogram segment to obtain the depth feature map of the spectrogram segment; perform environment prediction on the depth feature map of the spectrogram segment to obtain the posterior probability of environment prediction; based on the depth feature map of the spectrogram segment and the posterior probability of environment prediction, obtain the high-dimensional feature vector of the spectrogram segment; perform voiceprint extraction on the high-dimensional feature vector of the spectrogram segment to obtain the voiceprint feature vector of the spectrogram segment.
8. A voiceprint extraction device, characterized in that, It includes a memory and a processor. The memory is used to store programs. The processor is configured to execute the program to implement each step of the voiceprint extraction method as described in any one of claims 1 to 6.
9. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the voiceprint extraction method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Method, device and equipment for training voice model and computer readable storage medium
CN110428842A
Language recognition method and device, model training method and device, and facility
CN110853618A
Attentive adversarial domain-invariant training
US20200335108A1