A method, apparatus, device and storage medium for voiceprint feature extraction
By using a pre-trained vocalprint extraction model, using the training spectral fragments with unblocked timing and disrupted timing, the problem of vocalprint feature extraction in the prior art is solved, and accurate and robust vocalprint feature extraction is achieved, improving recognition accuracy.
Patent Information
- Application Number
- CN202310362146.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-03
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-04-03
AI Technical Summary
Existing voiceprint recognition technology is difficult to extract accurate and robust voiceprint features without being disturbed by voice timing.
By using a pre-trained vocalprint extraction model, the model uses the training sample of the training spectral fragments that are not disrupted and disrupted in timing, and learns a vocalprint feature extraction method that is not affected by timing.
It realizes the extraction of accurate and robust voiceprint features without being disturbed by voice timing, which improves the accuracy and stability of voiceprint recognition.
Smart Images

Figure CN116312563B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voiceprint recognition, and particularly to a method, device, equipment and storage medium for extracting voiceprint features. Background Art
[0002] Voiceprint recognition technology is one of the key technologies in the field of biometric authentication. Voiceprint recognition technology, also known as speaker recognition technology, uses the voice of a speaker to authenticate the identity of the speaker. Using voice for identity authentication not only has the characteristics of not requiring memory and simple judgment, but also can be authenticated without the user's knowledge, with a high user acceptance rate, and is widely used in fields such as finance and smart home.
[0003] The key to voiceprint recognition technology lies in the extraction of voiceprint features. For some fields with high requirements for authentication accuracy, it is necessary to extract relatively accurate and robust voiceprint features. It can be understood that if relatively accurate and robust voiceprint features are to be obtained, the voiceprint extraction method should not be interfered by the speech time sequence. That is, for a piece of speech data, regardless of how the time sequence of the speech data changes, the voiceprint features extracted for this piece of speech data should basically remain the same. However, there is currently no voiceprint extraction method that is not interfered by the speech time sequence. Summary of the Invention
[0004] In view of this, the present invention provides a method, device, equipment and storage medium for extracting voiceprint features. The voiceprint feature extraction method is not interfered by the speech time sequence, and its technical solution is as follows:
[0005] A method for extracting voiceprint features includes:
[0006] Obtaining a plurality of spectrogram segments of target speech data;
[0007] Based on a pre-trained voiceprint extraction model, extracting voiceprint features from the plurality of spectrogram segments of the target speech data respectively to obtain voiceprint features corresponding to the plurality of spectrogram segments of the target speech data. Among them, the voiceprint extraction model uses a plurality of training spectrogram segments with unchanged time sequence and a plurality of training spectrogram segments with disrupted time sequence as training samples, and uses the true identity labels corresponding to the respective training spectrogram segments included in the training samples as sample labels, and is trained with the goal of making the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples tend to be consistent with the corresponding true identity labels;
[0008] Based on the voiceprint features corresponding to the plurality of spectrogram segments of the target speech data, determining the voiceprint features corresponding to the target speech data.
[0009] Optionally, the process of obtaining the training samples includes:
[0010] Obtain a number of spectrogram segments from a pre-constructed set of spectrogram segments, where the set of spectrogram segments includes a plurality of spectrogram segments with non-disordered time sequences;
[0011] For each spectrogram segment obtained from the set of spectrogram segments:
[0012] Randomly generate a time sequence scrambling probability corresponding to this spectrogram segment;
[0013] If the time sequence scrambling probability corresponding to this spectrogram segment is greater than a set probability threshold, then scramble the time sequence of this spectrogram segment to obtain a spectrogram segment with a scrambled time sequence as a training spectrogram segment; if the time sequence scrambling probability corresponding to this spectrogram segment is less than or equal to the set probability threshold, then use this spectrogram segment as a training spectrogram segment;
[0014] Compose a training sample from the obtained training spectrogram segments.
[0015] Optionally, the scrambling of the time sequence of this spectrogram segment to obtain a spectrogram segment with a scrambled time sequence includes:
[0016] Slice this spectrogram segment into a plurality of spectrogram sub-segments, where each spectrogram sub-segment is a spectrogram sub-segment of consecutive multiple frames of speech;
[0017] Randomly scramble and combine the plurality of spectrogram sub-segments into a new spectrogram segment to obtain a spectrogram segment with a scrambled time sequence.
[0018] Optionally, the training process of the voiceprint extraction model includes:
[0019] For each training spectrogram segment included in the training sample:
[0020] Extract voiceprint features from this training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint features corresponding to this training spectrogram segment;
[0021] Predict the probabilities of the identity labels corresponding to this training spectrogram segment being each set identity label based on the voiceprint features corresponding to this training spectrogram segment;
[0022] Determine the prediction loss corresponding to this training spectrogram segment based on the probabilities of the identity labels corresponding to this training spectrogram segment being each set identity label and the true identity label corresponding to this training spectrogram segment;
[0023] Update the parameters of the voiceprint extraction model based on the prediction losses respectively corresponding to the training spectrogram segments included in the training sample.
[0024] Optionally, the updating of the parameters of the voiceprint extraction model based on the prediction losses respectively corresponding to the training spectrogram segments included in the training sample includes:
[0025] Fuse the prediction losses corresponding to each training spectrogram segment included in the training sample to obtain a fused loss;
[0026] Update the parameters of the voiceprint extraction model according to the fused loss.
[0027] Optionally, the extracting the voiceprint feature of the training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint feature corresponding to the training spectrogram segment includes:
[0028] Extract shallow features and deep features from the training spectrogram segment based on the voiceprint extraction model;
[0029] Fuse the shallow features and the deep features based on the voiceprint extraction model to obtain a fused feature;
[0030] Extract features from the fused feature based on the voiceprint extraction model as the target feature of the training spectrogram segment;
[0031] Average the features of each frame in the target feature of the training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint feature corresponding to the spectrogram segment.
[0032] Optionally, the voiceprint extraction model includes: a first feature extraction module, a second feature extraction module, a feature fusion part, a third feature extraction part and a feature processing module, wherein the second feature extraction module includes a plurality of cascaded feature extraction sub-modules;
[0033] The extracting the voiceprint feature of the training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint feature corresponding to the training spectrogram segment includes:
[0034] Input the training spectrogram segment into the first feature extraction module for feature extraction;
[0035] Input the feature output by the first feature extraction module into the second feature extraction module for feature extraction, wherein the input of the first feature extraction sub-module in the second feature extraction module is the feature output by the first feature extraction module, and the input of other feature extraction sub-modules is the feature output by the previous feature extraction sub-module;
[0036] Input the features output by each feature extraction sub-module in the second feature extraction module into the feature fusion module for feature fusion;
[0037] Input the fused feature output by the feature fusion module into the third feature extraction module for feature extraction;
[0038] Input the features output by the third feature extraction module into the feature processing module for processing to obtain the voiceprint features corresponding to the training spectrogram segment output by the feature processing module, where the feature processing module calculates the mean value of the features of each frame in the input features.
[0039] A voiceprint feature extraction device, comprising: a spectrogram segment acquisition module, a voiceprint feature extraction module, and a voiceprint feature determination module;
[0040] The spectrogram segment acquisition module is used to acquire a plurality of spectrogram segments of target voice data;
[0041] The voiceprint feature extraction module is used to respectively extract voiceprint features from a plurality of spectrogram segments of the target voice data based on a pre-trained voiceprint extraction model to obtain the voiceprint features corresponding to the plurality of spectrogram segments of the target voice data, where the voiceprint extraction model uses a plurality of training spectrogram segments with non-shuffled time sequences and a plurality of training spectrogram segments with shuffled time sequences as training samples, and uses the true identity labels corresponding to the respective training spectrogram segments included in the training samples as sample labels, and is trained with the goal that the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples are consistent with the corresponding true identity labels;
[0042] The voiceprint feature determination module is used to determine the voiceprint features corresponding to the target voice data based on the voiceprint features corresponding to the plurality of spectrogram segments of the target voice data.
[0043] A voiceprint feature extraction device, comprising: a memory and a processor;
[0044] The memory is used to store programs;
[0045] The processor is used to execute the program to implement each step of the voiceprint feature extraction method described in any one of the above.
[0046] A readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each step of the voiceprint feature extraction method described in any one of the above is implemented.
[0047] The voiceprint feature extraction method provided by the present invention first obtains several spectrogram segments of target voice data, then extracts voiceprint features from the several spectrogram segments of the target voice data respectively based on a pre-trained voiceprint extraction model to obtain the voiceprint features corresponding to the several spectrogram segments of the target voice data, and finally determines the voiceprint feature corresponding to the target voice data based on the voiceprint features corresponding to the several spectrogram segments of the target voice data. Since the voiceprint extraction model uses several training spectrogram segments with unchanged time series and several training spectrogram segments with disrupted time series as training samples, and uses the true identity labels corresponding to the respective training spectrogram segments included in the training samples as sample labels, and is trained with the goal of making the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples tend to be consistent with the corresponding true identity labels, therefore, whether it is a spectrogram segment with unchanged time series or a spectrogram segment with disrupted time series, the trained voiceprint extraction model can extract voiceprint features that can accurately predict the identity labels, that is, when extracting voiceprint features based on the trained voiceprint extraction model, it is not easily affected by time series information. Thus, accurate and robust voiceprint features can be obtained based on the trained voiceprint extraction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0049] Figure 1 It is a schematic diagram of the hardware architecture related to the present invention;
[0050] Figure 2 It is a schematic flow chart of the voiceprint feature extraction method provided by the embodiment of the present invention;
[0051] Figure 3 It is a schematic flow chart of obtaining training samples provided by the embodiment of the present invention;
[0052] Figure 4 It is a schematic diagram of a spectrogram segment with unchanged time series and spectrogram segments obtained by cutting it according to different segmentation lengths, randomly disrupting and then recombining provided by the embodiment of the present invention;
[0053] Figure 5 It is a schematic flow chart of training a voiceprint extraction model provided by the embodiment of the present invention;
[0054] Figure 6 It is an example of the voiceprint extraction model provided by the embodiment of the present invention;
[0055] Figure 7 Schematic structural diagram of the voiceprint feature extraction device provided by an embodiment of the present invention;
[0056] Figure 8 Schematic structural diagram of the voiceprint feature extraction device provided by an embodiment of the present application. Detailed implementation manners
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0058] The inventor of this case found in the process of implementing this case that the current mainstream voiceprint feature extraction method is the method based on total variation factor analysis. The overall process of this method is to pre-train a total variation space covering various environments and channels by using a large amount of corpus. For a piece of voice data, the voice data is mapped into a voiceprint feature with a fixed and unified dimension by using the pre-trained total variation space.
[0059] However, in the case of short voice duration, the above voiceprint feature extraction method based on total variation factor analysis will result in unstable voiceprint features due to insufficient calculation of statistics, and further lead to low accuracy of subsequent voiceprint authentication. Moreover, the above solution does not consider the influence of voice timing on voiceprint feature extraction.
[0060] In view of the defects of the above voiceprint feature extraction method based on total variation factor analysis, and coupled with the achievements of deep learning methods in many research fields in recent years, there has emerged a voiceprint feature extraction solution based on deep learning. Specifically, a voiceprint extraction model is pre-trained by using the spectrogram of training voice data. In the actual application stage, the spectrogram of the target voice data is obtained, and the voiceprint feature is extracted from the spectrogram of the target voice data based on the pre-trained voiceprint extraction model.
[0061] Since the voiceprint extraction model usually adopts a convolutional neural network, and due to the inherent receptive field limitation in the convolutional neural network, during the feature extraction process, only the speech features of several consecutive frames within the receptive field can be focused on, and the speech features outside the receptive field cannot be attended to. This results in the trained convolutional neural network, i.e., the voiceprint extraction model, being dependent on the temporal information of the speech data. If the temporal order of the speech data is disrupted, the obtained voiceprint features will be different. That is, the voiceprint features extracted from an unshuffled speech and its corresponding shuffled speech using the above voiceprint feature extraction method are very different. In fact, the disruption of the temporal order only affects the coherence of the speech content and does not change the personalized voiceprint information contained in the speech. That is to say, the above feature extraction scheme is vulnerable to the interference of the speech temporal order.
[0062] In order to make the speech temporal order not affect the extraction of voiceprint features, so as to obtain accurate and robust voiceprint features, the inventors of this case conducted in-depth research. Through continuous research, a voiceprint feature extraction method that is not interfered by the speech temporal order was finally proposed. Before introducing the voiceprint feature extraction method provided by the present invention, the hardware architecture involved in the present invention will be described first.
[0063] In a possible implementation manner, as Figure 1 shown, the hardware architecture involved in the present invention may include: an electronic device 101 and a server 102.
[0064] Exemplarily, the electronic device 101 can be any electronic product that can perform human-computer interaction with the user in one or more ways such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device. For example, a personal computer, a laptop computer, a tablet computer, a mobile phone, a smart TV, etc.
[0065] It should be noted that Figure 1 is only an example, and there can be various types of electronic devices, not limited to Figure 1 the laptop computer in
[0066] Exemplarily, the server 102 can be a single server, a server cluster composed of multiple servers, or a cloud computing server center. The server 102 may include a processor, a memory, and a network interface, etc.
[0067] Exemplarily, the electronic device 101 can establish a connection and communicate with the server 102 through a wireless communication network; exemplarily, the electronic device 101 can establish a connection and communicate with the server 102 through a wired network.
[0068] The electronic device 101 can obtain target speech data, send the target speech data to the server 102, and the server 102 obtains the voiceprint features corresponding to the target speech data according to the voiceprint feature extraction method provided by the present invention.
[0069] In another possible implementation, the hardware architecture involved in the present invention may include: an electronic device. The electronic device is a device with strong data processing capabilities.
[0070] Exemplarily, the electronic device can be any electronic product that can perform human-computer interaction with a user in one or more ways such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device. For example, a personal computer, a laptop computer, a tablet computer, a mobile phone, a smart TV, etc.
[0071] The electronic device can obtain target voice data and obtain the voiceprint feature corresponding to the target voice data according to the voiceprint feature extraction method provided by the present invention.
[0072] Those skilled in the art should understand that the above-mentioned electronic devices and servers are only examples. Other existing or future possible electronic devices or servers that can be applied to the present invention should also be included within the protection scope of the present invention and are hereby incorporated herein by reference.
[0073] Next, the voiceprint feature extraction method provided by the present invention will be introduced through the following embodiments.
[0074] Please refer to Figure 1 , which shows a schematic flowchart of the voiceprint feature extraction method provided by an embodiment of the present invention. The method may include:
[0075] Step S201: Obtain a plurality of spectrogram segments of the target voice data.
[0076] In this embodiment, the process of obtaining the voice segment sequence corresponding to the target voice data may include:
[0077] Step S2011: Obtain the voice feature of the target voice data.
[0078] Among them, the voice feature of the target voice data may be a filter bank feature (i.e., filterbank). Specifically, the target voice data can be windowed and Fourier-transformed to obtain the filter bank feature. This embodiment does not limit the voice feature of the target voice data to the filter bank feature, and it can also be other features, such as Mel-frequency cepstral coefficients (MFCC).
[0079] Step S2012: Construct the voice feature corresponding to the target voice data into a spectrogram as the spectrogram corresponding to the target voice data.
[0080] Step S2013: Segment the spectrogram corresponding to the target voice data according to a preset segmentation length to obtain a plurality of spectrogram segments of the target voice data.
[0081] Exemplarily, if the segmentation length is 400, each spectrogram segment obtained by segmenting the spectrogram corresponding to the target speech data according to the segmentation length of 400 is a spectrogram segment of 400 frames of speech.
[0082] Step S202: Based on the pre-trained voiceprint extraction model, extract voiceprint features from several spectrogram segments of the target speech data respectively, and obtain the voiceprint features corresponding to several spectrogram segments of the target speech data respectively.
[0083] Specifically, input each spectrogram segment of the target speech data into the pre-trained voiceprint extraction model, and the voiceprint extraction model outputs the voiceprint features corresponding to each spectrogram segment of the target speech data respectively.
[0084] The voiceprint extraction model in this embodiment uses several training spectrogram segments with unshuffled time series and several training spectrogram segments with shuffled time series as training samples, and is trained with the true identity labels corresponding to each training spectrogram segment included in the training samples. The training objective of the voiceprint extraction model is to make the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples based on the voiceprint extraction model tend to be consistent with the corresponding true identity labels.
[0085] It should be noted that each training spectrogram segment included in the training samples can be a spectrogram segment obtained from the same speech data, or a spectrogram segment obtained from different speech data.
[0086] In addition, it should be noted that the training samples include training spectrogram segments with unshuffled time series because, in the actual application stage, voiceprint features are extracted from each spectrogram segment of the target speech data (the spectrogram segments obtained by segmenting the spectrogram constructed based on the spectrogram features of the target speech data, rather than the spectrogram segments with shuffled time series obtained by segmenting the spectrogram segments). The training samples include training spectrogram segments with shuffled time series in order to train the voiceprint extraction model to be unaffected by the time series.
[0087] Step S203: Based on the voiceprint features corresponding to several spectrogram segments of the target speech data, determine the voiceprint feature corresponding to the target speech data.
[0088] Specifically, fuse the voiceprint features corresponding to several spectrogram segments of the target speech data, and the fused features are used as the voiceprint features corresponding to the target speech data.
[0089] There are various implementation methods for fusing the voiceprint features corresponding to several spectrogram segments of the target speech data:
[0090] In a possible implementation, the average value of the voiceprint features corresponding to several spectrogram segments of the target voice data can be calculated, and the obtained average value is used as the voiceprint feature corresponding to the target voice data, that is:
[0091]
[0092] In another possible implementation, the average value can be calculated after weighting the voiceprint features corresponding to several spectrogram segments of the target voice data, and the obtained average value after weighting is used as the voiceprint feature corresponding to the target voice data, that is:
[0093]
[0094] In formulas (1) and (2), N represents the total number of spectrogram segments of the target voice data, w i represents the voiceprint feature corresponding to the i-th spectrogram segment of the target voice data, α i represents the weight corresponding to the voiceprint feature w i and w represents the voiceprint feature corresponding to the target voice data.
[0095] The voiceprint feature extraction method provided by the embodiments of the present invention first obtains several spectrogram segments of the target voice data, then extracts the voiceprint features of several spectrogram segments of the target voice data respectively based on a pre-trained voiceprint extraction model, obtains the voiceprint features corresponding to several spectrogram segments of the target voice data respectively, and finally determines the voiceprint feature corresponding to the target voice data based on the voiceprint features corresponding to several spectrogram segments of the target voice data. Since the voiceprint extraction model uses several training spectrogram segments with unchanged time sequence and several training spectrogram segments with disrupted time sequence as training samples, and uses the true identity labels corresponding to each training spectrogram segment included in the training samples as sample labels, and is trained with the goal of making the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples tend to be consistent with the corresponding true identity labels, therefore, whether it is a spectrogram segment with unchanged time sequence or a spectrogram segment with disrupted time sequence, the trained voiceprint extraction model can extract voiceprint features that can accurately predict the identity labels, that is, when extracting voiceprint features based on the trained voiceprint extraction model, it is not easily affected by the time sequence information. Thus, accurate and robust voiceprint features can be obtained based on the trained voiceprint extraction model.
[0096] The voiceprint extraction model in the voiceprint feature extraction method provided in the above embodiments is trained using training samples including several training spectrogram segments with unchanged time sequence and several training spectrogram segments with disrupted time sequence. In another embodiment of the present invention, the process of obtaining the training samples is introduced.
[0097] Please refer to Figure 3, showing a schematic flow chart for obtaining a training sample, may include:
[0098] Step S301: Obtain a number of spectrogram segments from a pre-constructed spectrogram segment set.
[0099] Among them, the spectrogram segment set includes multiple spectrogram segments. Each spectrogram segment in the spectrogram segment set is a spectrogram segment with the time sequence not shuffled, and each spectrogram segment has a true identity label.
[0100] Optionally, the number of spectrogram segments included in the training sample can be set. When obtaining spectrogram segments from the spectrogram segment set, obtain the set number of spectrogram segments from the spectrogram segment set. For example, if the set number is 64, then obtain 64 spectrogram segments from the spectrogram segment set.
[0101] The spectrogram segment set in this embodiment can be constructed in the following way: Obtain a voice data set. The voice data set includes multiple voice data, and each voice data in the voice data set has a true identity label; for each voice data in the voice data set, obtain the spectrogram of the voice data (first obtain the voice features of the voice data, such as filterbank features, and then construct the spectrogram of the voice data from the voice features of the voice data), and segment the spectrogram of the voice data according to a set segmentation length to obtain a number of spectrogram segments of the voice data, and use the true identity label of the voice data as the true identity label of each spectrogram segment of the voice data; form the spectrogram segment set from all the obtained spectrogram segments with true identity labels.
[0102] Assume that the dimension of the voice features corresponding to the voice data is d, and the preset segmentation length is l. The size of each spectrogram segment obtained by segmenting the spectrogram of the voice data according to the preset segmentation length l is l×d. It should be noted that if the length of the voice data is less than l, before obtaining the spectrogram of the voice data, the voice data can be copied several times to make its length greater than or equal to l. If there is redundant data after copying, the redundant data is discarded.
[0103] It should be noted that the number of spectrogram segments obtained from the pre-constructed spectrogram segment set may be spectrogram segments of the same voice data, or may be voice segments of different voice data.
[0104] Step S302: For each spectrogram segment obtained from the spectrogram segment set, randomly generate the time sequence shuffling probability corresponding to the spectrogram segment.
[0105] It should be noted that the time sequence shuffling probability corresponding to the spectrogram segment is the probability of shuffling the time sequence of the spectrogram segment.
[0106] Exemplarily, a number can be randomly generated in the range of [0,1] as the time sequence shuffling probability corresponding to the spectrogram segment.
[0107] Step S303: Determine whether the temporal scrambling probability corresponding to this spectrogram segment is greater than a set probability threshold. If so, execute Step S304-a; if not, execute Step S304-b.
[0108] In this embodiment, by comparing the temporal scrambling probability corresponding to this spectrogram segment with the set probability threshold, it is decided whether to scramble the time sequence of this spectrogram segment.
[0109] Step S304-a: Scramble the time sequence of this spectrogram segment to obtain a spectrogram segment with a scrambled time sequence, which is used as a training spectrogram segment.
[0110] If the temporal scrambling probability corresponding to this spectrogram segment is greater than the set probability threshold, then scramble the time sequence of this spectrogram segment, and use the spectrogram segment with the scrambled time sequence as a training spectrogram segment.
[0111] Exemplarily, the set probability threshold is 0.5, and the temporal scrambling probability corresponding to this spectrogram segment is 0.8, which is greater than 0.5. Then, scramble the time sequence of this spectrogram segment, and use the spectrogram segment with the scrambled time sequence as a training spectrogram segment.
[0112] Optionally, the process of scrambling the time sequence of this spectrogram segment to obtain a spectrogram segment with a scrambled time sequence may include: dividing this spectrogram segment into multiple spectrogram sub-segments, where each spectrogram sub-segment is a spectrogram sub-segment of consecutive multiple frames of speech; randomly scrambling the multiple spectrogram sub-segments obtained by the division and then combining them into a new spectrogram segment to obtain a spectrogram segment with a scrambled time sequence.
[0113] Considering that frame-level scrambling of the spectrogram segment will seriously damage the spectrogram segment, in this embodiment, segment-level scrambling is performed on the spectrogram segment, that is, segment-level segmentation is performed on the spectrogram segment (each spectrogram sub-segment obtained by the segmentation is a spectrogram sub-segment of consecutive multiple frames of speech), and then the multiple spectrogram sub-segments obtained by the segmentation are randomly scrambled and combined to obtain a spectrogram segment with segment-level scrambling.
[0114] When segmenting a spectrogram segment, the spectrogram segment can be segmented according to the set segmentation length s. It should be noted that the segmentation length s needs to be set appropriately. Different segmentation lengths will affect the integrity and coherence of the spectrogram segment. If the segmentation length s is small, the spectrogram segment will be severely damaged, and the local continuous frame-level correlation will be lost completely, approaching the noise level. If the segmentation length s is large, it is not enough to change the dependence of the voiceprint extraction model on the speech time series. Optionally, the segmentation length s can be 100, that is, each spectrogram sub-segment obtained by segmenting according to the segmentation length of 100 is a spectrogram sub-segment of 100-frame speech. It should be noted that the segmentation length s being 100 is only an example. The segmentation length s can also be other values, such as 90, 110, etc., as long as it can change the dependence of the voiceprint extraction model on the time series and retain the local continuous frame-level correlation.
[0115] Please refer to Figure 4 , which shows a schematic diagram of a spectrogram segment without time series shuffling and spectrogram segments obtained by segmenting it according to different segmentation lengths, randomly shuffling and then recombining them. Figure (a) is the spectrogram segment without time series shuffling, Figure (b) is the spectrogram after frame-level shuffling and recombination (that is, the spectrogram segment obtained by segmenting, shuffling, and recombining the spectrogram segment in Figure (a) frame by frame), and Figures (c) to (h) are spectrogram segments obtained by segmenting, shuffling, and recombining the spectrogram segment in Figure (a) according to the segmentation lengths of 10, 20, 50, 80, 100, and 200 in sequence.
[0116] Step S304-b: Use this spectrogram segment as the training spectrogram segment.
[0117] If the time series shuffling probability corresponding to this spectrogram segment is less than or equal to the set probability threshold, directly use this spectrogram segment as the training spectrogram segment.
[0118] Exemplarily, the set probability threshold is 0.5, and the time series shuffling probability corresponding to this spectrogram segment is 0.3, which is less than 0.5. Then, do not shuffle the time series of this spectrogram segment and directly use this spectrogram segment as the training spectrogram segment.
[0119] Step S305: Compose a training sample from all the obtained training spectrogram segments.
[0120] The training sample obtained through the above method contains both spectrogram segments with unshuffled time series and spectrogram segments with shuffled time series.
[0121] The above steps S301 to S305 can be executed multiple times, and in this way, multiple training samples can be obtained. It should be noted that each time steps S301 to S305 are executed, steps S302 to S305 can be executed multiple times. Since the training spectrogram segments to be shuffled are determined by random probability, the training samples obtained by executing steps S302 to S305 multiple times are different.
[0122] In addition, it should be noted that this embodiment does not limit the use of the above method to obtain training samples, and other methods can also be used to obtain training samples. For example, several spectrogram segments can be first obtained from a pre-constructed spectrogram segment set, assumed to be M, and then a ratio is randomly generated as the proportion of the spectrogram segments to be shuffled in time series. Based on this ratio and M, the number of spectrogram segments to be shuffled in time series is determined, assumed to be Q. Q of the M obtained spectrogram segments are shuffled in time series, and the Q spectrogram segments shuffled in time series and the M - Q spectrogram segments not shuffled in time series are combined to form a training sample.
[0123] In another embodiment of the present invention, the training process of the voiceprint extraction model is introduced.
[0124] Please refer to Figure 5 , which shows a schematic flowchart of training the voiceprint extraction model, and may include:
[0125] Step S501: For each training spectrogram segment included in the training sample, based on the voiceprint extraction model, extract the voiceprint feature of this training spectrogram segment to obtain the voiceprint feature corresponding to this training spectrogram segment.
[0126] Specifically, the process of extracting the voiceprint feature of this training spectrogram segment based on the voiceprint extraction model may include: extracting shallow features and deep features of this training spectrogram segment based on the voiceprint extraction model; fusing the shallow features and deep features extracted from this training spectrogram segment based on the voiceprint extraction model to obtain the fused features; extracting features from the fused features based on the voiceprint extraction model as the target features of this training spectrogram segment; and averaging the features of each frame in the target features of this training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint feature corresponding to this spectrogram segment.
[0127] Please refer to Figure 6 , which shows an example of the voiceprint extraction model. Figure 6The shown voiceprint extraction model includes a first feature extraction module 601, a second feature extraction module 602, a feature fusion part 603, a third feature extraction module 604, and a feature processing module 605. Among them, the structures of the first feature extraction module 601 and the third feature extraction module 604 can be the same. The second feature extraction module 602 includes multiple cascaded feature extraction sub-modules (such as 3 cascaded feature extraction sub-modules), and the structures of the respective feature extraction sub-modules included in the second feature extraction module 602 can be the same. Considering that a convolutional neural network can perform joint analysis on the time domain and the frequency domain and deeply excavate the voiceprint information in the spectrogram, the first feature extraction module 601 and the third feature extraction module 604 can adopt a traditional convolutional network, and the respective feature extraction sub-modules included in the second feature extraction module 602 can adopt a residual network, that is, the second feature extraction module 602 can include multiple groups of residual networks. Optionally, the residual network can be SE-ResNet or SE-Res2Block.
[0128] The following combines Figure 6 the shown voiceprint extraction model to further introduce the process of extracting voiceprint features from the training spectrogram segment based on the voiceprint extraction model.
[0129] Based on Figure 6 the shown voiceprint extraction model, the process of extracting voiceprint features from the training spectrogram segment can include:
[0130] Step a1: Input the training spectrogram segment into the first feature extraction module 601 of the voiceprint extraction model for feature extraction.
[0131] The first feature extraction module 601 extracts features from the input training spectrogram segment and outputs the extracted features.
[0132] Step a2: Input the features output by the first feature extraction module 601 into the second feature extraction module 602 of the voiceprint extraction model for feature extraction.
[0133] Among them, the input of the first feature extraction sub-module in the second feature extraction module 602 is the features output by the first feature extraction module 601. It extracts features from the input features and outputs the extracted features. The input of each of the other feature extraction sub-modules is the features output by the previous feature extraction sub-module, and each of the other feature extraction sub-modules also extracts features from the input features and outputs the extracted features.
[0134] Step a3: Input the features output by the respective feature extraction sub-modules in the second feature extraction module 602 into the feature fusion module 603 of the voiceprint extraction model for feature fusion.
[0135] The features output by each feature extraction sub-module in the second feature extraction module 602 include shallow features and deep features (the features output by the relatively front feature extraction sub-modules are shallow features, and the features output by the relatively back feature extraction sub-modules are deep features). The feature fusion module 603 fuses the shallow features and deep features and outputs the fused features.
[0136] Step a4: Input the fused features output by the feature fusion module 603 into the third feature extraction module 604 of the voiceprint extraction model for feature extraction.
[0137] The third feature extraction module 604 extracts features from the input fused features and outputs the extracted features.
[0138] Step a5: Input the features output by the third feature extraction module 604 into the feature processing module 605 of the voiceprint extraction model for processing to obtain the voiceprint features corresponding to the training spectrogram segment output by the feature processing module 605.
[0139] Among them, the feature processing module 605 calculates the mean value of the features of each frame in the input features and outputs it.
[0140] Step S502: Based on the voiceprint features corresponding to the training spectrogram segment, predict the probabilities that the identity labels corresponding to the training spectrogram segment are each set identity label.
[0141] Among them, each set identity label may include the true identity labels of each spectrogram segment in the spectrogram segment set.
[0142] Step S503: Based on the probabilities that the identity label corresponding to the training spectrogram segment is each set identity label and the true identity label corresponding to the training spectrogram segment, determine the prediction loss corresponding to the training spectrogram segment.
[0143] Optionally, the cross-entropy loss can be determined based on the probabilities that the identity label corresponding to the training spectrogram segment is each set identity label and the true identity label corresponding to the training spectrogram segment, and used as the prediction loss corresponding to the training spectrogram segment.
[0144] Through the above process, the prediction losses corresponding to each training spectrogram segment included in the training sample can be obtained.
[0145] Step S504: Based on the prediction losses corresponding to each training spectrogram segment included in the training sample, update the parameters of the voiceprint extraction model.
[0146] Specifically, the process of updating the parameters of the voiceprint extraction model based on the prediction losses corresponding to the respective training spectrogram segments included in the training samples may include: fusing the prediction losses corresponding to the respective training spectrogram segments included in the training samples to obtain a fused loss; and updating the parameters of the voiceprint extraction model in the reverse direction according to the fused loss. Among them, the manner of fusing the prediction losses corresponding to the respective training spectrogram segments included in the training samples may be to sum the prediction losses corresponding to the respective training spectrogram segments included in the training samples. Of course, this embodiment is not limited thereto. For example, the prediction losses corresponding to the respective training spectrogram segments included in the training samples may also be weighted and summed.
[0147] Exemplarily, the training sample includes 64 training spectrogram segments. Through steps S301 to S304, the prediction losses corresponding to the 64 training spectrogram segments can be obtained respectively. After obtaining the prediction losses corresponding to the 64 training spectrogram segments respectively, the prediction losses corresponding to the 64 training spectrogram segments can be fused, and the voiceprint extraction model can be adjusted according to the fused loss.
[0148] The voiceprint extraction model can be trained by using multiple different training samples (the training samples are obtained according to the training sample acquisition method provided in the above embodiment) according to the process of steps S501 to S504 until the training end condition is met (for example, the model converges or reaches a preset number of training times, etc.).
[0149] After the training is completed, a voiceprint extraction model that is not easily affected by temporal information, has a strong voiceprint feature expression ability, and can extract accurate and robust voiceprint features can be obtained.
[0150] After obtaining the trained voiceprint extraction model, the voiceprint features corresponding to the target voice data can be obtained based on this model. That is, first, several spectrogram segments of the target voice data are obtained, and then the several spectrogram segments of the target voice data are input into the trained voiceprint extraction model. The voiceprint extraction model extracts the voiceprint features of the several spectrogram segments of the target voice data respectively to obtain the voiceprint features corresponding to the several spectrogram segments of the target voice data. Finally, the voiceprint features corresponding to the target voice data are determined according to the voiceprint features corresponding to the several spectrogram segments of the target voice data. It should be noted that the process of extracting the voiceprint features of the several spectrogram segments of the target voice data based on the voiceprint extraction model is similar to the process of extracting the voiceprint features of the respective training spectrogram segments included in the training samples based on the voiceprint extraction model. For details, reference can be made to the relevant part, and this embodiment will not be elaborated here.
[0151] The embodiment of the present invention also provides a voiceprint feature extraction device. The voiceprint feature extraction device provided by the embodiment of the present invention will be described below. The voiceprint feature extraction device described below can be correspondingly referred to the voiceprint feature extraction method described above.
[0152] Please refer to Figure 7 which shows a schematic structural diagram of the voiceprint feature extraction device provided by the embodiment of the present invention, and may include: a spectrogram segment acquisition module 701, a voiceprint feature extraction module 702, and a voiceprint feature determination module 703.
[0153] The spectrogram segment acquisition module 701 is used to acquire a plurality of spectrogram segments of the target voice data;
[0154] The voiceprint feature extraction module 702 is used to respectively extract voiceprint features from a plurality of spectrogram segments of the target voice data based on a pre-trained voiceprint extraction model, and obtain the voiceprint features corresponding to the plurality of spectrogram segments of the target voice data.
[0155] Wherein, the voiceprint extraction model uses a plurality of training spectrogram segments with non-disrupted time sequences and a plurality of training spectrogram segments with disrupted time sequences as training samples, and uses the true identity labels corresponding to the respective training spectrogram segments included in the training samples as sample labels, and is trained with the goal of making the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples tend to be consistent with the corresponding true identity labels.
[0156] The voiceprint feature determination module 703 is used to determine the voiceprint feature corresponding to the target voice data based on the voiceprint features corresponding to the plurality of spectrogram segments of the target voice data.
[0157] Optionally, the voiceprint feature extraction device provided by the embodiment of the present invention may further include: a training sample acquisition module for acquiring training samples.
[0158] The training sample acquisition module is used for:
[0159] Acquire a plurality of spectrogram segments from a pre-constructed spectrogram segment set, wherein the spectrogram segment set includes a plurality of spectrogram segments with non-disrupted time sequences;
[0160] For each spectrogram segment acquired from the spectrogram segment set:
[0161] Randomly generate a time sequence disruption probability corresponding to this spectrogram segment;
[0162] If the time sequence disruption probability corresponding to this spectrogram segment is greater than a set probability threshold, then disrupt the time sequence of this spectrogram segment to obtain a spectrogram segment with disrupted time sequence as a training spectrogram segment; if the time sequence disruption probability corresponding to this spectrogram segment is less than or equal to the set probability threshold, then use this spectrogram segment as a training spectrogram segment;
[0163] Form a training sample from the obtained training spectrogram segments.
[0164] Optionally, when the training sample acquisition module shuffles the time sequence of the spectrogram segment to obtain a spectrogram segment with a shuffled time sequence, it is specifically used for:
[0165] Segment the spectrogram segment into multiple spectrogram sub-segments, where each spectrogram sub-segment is a spectrogram sub-segment of consecutive multiple frames of speech;
[0166] Randomly shuffle and combine the multiple spectrogram sub-segments into a new spectrogram segment to obtain a spectrogram segment with a shuffled time sequence.
[0167] Optionally, the voiceprint feature extraction device provided by the embodiments of the present invention may further include: a model training module for training a voiceprint feature extraction model.
[0168] The model training module is used for:
[0169] For each training spectrogram segment included in the training sample:
[0170] Extract voiceprint features from the training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint features corresponding to the training spectrogram segment;
[0171] Predict the probabilities that the identity label corresponding to the training spectrogram segment is each set identity label based on the voiceprint features corresponding to the training spectrogram segment;
[0172] Determine the prediction loss corresponding to the training spectrogram segment based on the probabilities that the identity label corresponding to the training spectrogram segment is each set identity label and the true identity label corresponding to the training spectrogram segment;
[0173] Update the parameters of the voiceprint extraction model based on the prediction losses respectively corresponding to the training spectrogram segments included in the training sample.
[0174] Optionally, when the model training module updates the parameters of the voiceprint extraction model based on the prediction losses respectively corresponding to the training spectrogram segments included in the training sample, it is specifically used for:
[0175] Fuse the prediction losses respectively corresponding to the training spectrogram segments included in the training sample to obtain a fused loss;
[0176] Update the parameters of the voiceprint extraction model according to the fused loss.
[0177] Optionally, when the model training module extracts voiceprint features from the training spectrogram segment to obtain the voiceprint features corresponding to the training spectrogram segment, it is specifically used for:
[0178] Extract shallow features and deep features from the training spectrogram segment based on the voiceprint extraction model;
[0179] Based on the voiceprint extraction model, fuse the shallow features and the deep features to obtain the fused features;
[0180] Based on the voiceprint extraction model, extract features from the fused features as the target features of this training spectrogram segment;
[0181] Based on the voiceprint extraction model, calculate the mean value of the features of each frame in the target features of this training spectrogram segment to obtain the voiceprint feature corresponding to this spectrogram segment.
[0182] Optionally, the voiceprint extraction model includes: a first feature extraction module, a second feature extraction module, a feature fusion part, a third feature extraction part, and a feature processing module, wherein the second feature extraction module includes a plurality of cascaded feature extraction sub-modules;
[0183] When the model training module extracts the voiceprint feature from this training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint feature corresponding to this training spectrogram segment, it is specifically used for:
[0184] Input this training spectrogram segment into the first feature extraction module for feature extraction;
[0185] Input the features output by the first feature extraction module into the second feature extraction module for feature extraction, wherein the input of the first feature extraction sub-module in the second feature extraction module is the features output by the first feature extraction module, and the input of other feature extraction sub-modules is the features output by the previous feature extraction sub-module;
[0186] Input the features output by each feature extraction sub-module in the second feature extraction module into the feature fusion module for feature fusion;
[0187] Input the fused features output by the feature fusion module into the third feature extraction module for feature extraction;
[0188] Input the features output by the third feature extraction module into the feature processing module for processing to obtain the voiceprint feature corresponding to this training spectrogram segment output by the feature processing module, wherein the feature processing module calculates the mean value of the features of each frame in the input features.
[0189] The voiceprint feature extraction device provided by the embodiment of the present invention first obtains several spectrogram segments of target voice data, then extracts voiceprint features from the several spectrogram segments of the target voice data respectively based on a pre-trained voiceprint extraction model to obtain the voiceprint features corresponding to the several spectrogram segments of the target voice data respectively, and finally determines the voiceprint feature corresponding to the target voice data based on the voiceprint features corresponding to the several spectrogram segments of the target voice data respectively. Since the voiceprint extraction model uses several training spectrogram segments with unchanged time series and several training spectrogram segments with shuffled time series as training samples, and uses the true identity labels corresponding to the respective training spectrogram segments included in the training samples as sample labels, and is trained with the goal of making the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples tend to be consistent with the corresponding true identity labels, therefore, whether it is a spectrogram segment with unchanged time series or a spectrogram segment with shuffled time series, the trained voiceprint extraction model can extract voiceprint features that can accurately predict the identity labels, that is, when extracting voiceprint features based on the trained voiceprint extraction model, it is not easily affected by time series information. Thus, accurate and robust voiceprint features can be obtained based on the trained voiceprint extraction model.
[0190] The embodiment of the present invention also provides a voiceprint feature extraction device. Please refer to Figure 8 , which shows a schematic structural diagram of the voiceprint feature extraction device. The voiceprint feature extraction device may include: at least one processor 801, at least one communication interface 802, at least one memory 803, and at least one communication bus 804;
[0191] In the embodiment of the present application, the number of the processor 801, the communication interface 802, the memory 803, and the communication bus 804 is at least one, and the processor 801, the communication interface 802, and the memory 803 complete mutual communication through the communication bus 804;
[0192] The processor 801 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiment of the present invention, etc.;
[0193] The memory 803 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory;
[0194] Among them, the memory stores a program, and the processor can call the program stored in the memory. The program is used for:
[0195] Obtain several spectrogram segments of target voice data;
[0196] Extract voiceprint features from several spectrogram segments of the target voice data respectively based on a pre-trained voiceprint extraction model, and obtain the voiceprint features corresponding to several spectrogram segments of the target voice data. Among them, the voiceprint extraction model uses several training spectrogram segments with unchanged time sequences and several training spectrogram segments with disrupted time sequences as training samples, and uses the true identity labels corresponding to each training spectrogram segment included in the training samples as sample labels, and is trained with the goal of making the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples tend to be consistent with the corresponding true identity labels;
[0197] Determine the voiceprint feature corresponding to the target voice data based on the voiceprint features corresponding to several spectrogram segments of the target voice data.
[0198] Optionally, the refinement function and expansion function of the program can be referred to the above description.
[0199] The embodiment of the present invention also provides a readable storage medium, which can store a program suitable for a processor to execute. The program is used for:
[0200] Obtain several spectrogram segments of the target voice data;
[0201] Extract voiceprint features from several spectrogram segments of the target voice data respectively based on a pre-trained voiceprint extraction model, and obtain the voiceprint features corresponding to several spectrogram segments of the target voice data. Among them, the voiceprint extraction model uses several training spectrogram segments with unchanged time sequences and several training spectrogram segments with disrupted time sequences as training samples, and uses the true identity labels corresponding to each training spectrogram segment included in the training samples as sample labels, and is trained with the goal of making the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples tend to be consistent with the corresponding true identity labels;
[0202] Determine the voiceprint feature corresponding to the target voice data based on the voiceprint features corresponding to several spectrogram segments of the target voice data.
[0203] Optionally, the refinement function and expansion function of the program can be referred to the above description.
[0204] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0205] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0206] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voiceprint feature extraction method, characterized in that, Including: Obtaining a plurality of spectrogram segments of target voice data; Based on a pre-trained voiceprint extraction model, extracting voiceprint features from the plurality of spectrogram segments of the target voice data respectively, to obtain voiceprint features corresponding to the plurality of spectrogram segments of the target voice data respectively. Among them, the voiceprint extraction model uses a plurality of training spectrogram segments with unchanged time sequences and a plurality of training spectrogram segments with shuffled time sequences as training samples, and uses the true identity labels corresponding to the respective training spectrogram segments included in the training samples as sample labels, so as to train the voiceprint extraction model with the goal that the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples are consistent with the corresponding true identity labels; Based on the voiceprint features corresponding to the plurality of spectrogram segments of the target voice data respectively, determining the voiceprint feature corresponding to the target voice data; The obtaining process of the training samples includes: Obtaining a plurality of spectrogram segments from a pre-constructed spectrogram segment set, where the spectrogram segment set includes a plurality of spectrogram segments with unchanged time sequences; For each spectrogram segment obtained from the spectrogram segment set: Randomly generating a time sequence shuffling probability corresponding to the spectrogram segment; If the time sequence shuffling probability corresponding to the spectrogram segment is greater than a set probability threshold, then shuffling the time sequence of the spectrogram segment at the segment level to obtain a spectrogram segment with a shuffled time sequence, which is used as a training spectrogram segment; if the time sequence shuffling probability corresponding to the spectrogram segment is less than or equal to the set probability threshold, then using the spectrogram segment as a training spectrogram segment; Composing the obtained training spectrogram segments into a training sample.
2. The voiceprint feature extraction method according to claim 1, characterized in that The shuffling the time sequence of the spectrogram segment to obtain a spectrogram segment with a shuffled time sequence includes: Cutting the spectrogram segment into a plurality of spectrogram sub-segments, where each spectrogram sub-segment is a spectrogram sub-segment of consecutive multiple frames of speech; Randomly shuffling the plurality of spectrogram sub-segments and then combining them into a new spectrogram segment to obtain a spectrogram segment with a shuffled time sequence.
3. The voiceprint feature extraction method according to any one of claims 1 to 2, characterized in that The training process of the voiceprint extraction model includes: For each training spectrogram segment included in the training sample: Based on the voiceprint extraction model, extracting voiceprint features from the training spectrogram segment to obtain the voiceprint features corresponding to the training spectrogram segment; Based on the voiceprint features corresponding to the training spectrogram segment, predicting the probabilities that the identity label corresponding to the training spectrogram segment is each set identity label; Based on the probabilities that the identity label corresponding to the training spectrogram segment is each set identity label and the true identity label corresponding to the training spectrogram segment, determining the prediction loss corresponding to the training spectrogram segment; Based on the prediction losses corresponding to the respective training spectrogram segments included in the training sample, updating the parameters of the voiceprint extraction model.
4. The voiceprint feature extraction method according to claim 3, characterized in that The updating the parameters of the voiceprint extraction model based on the prediction losses corresponding to the respective training spectrogram segments included in the training sample includes: Fusing the prediction losses corresponding to the respective training spectrogram segments included in the training sample to obtain a fused loss; According to the fused loss, updating the parameters of the voiceprint extraction model.
5. The voiceprint feature extraction method according to claim 3, characterized in that The extracting voiceprint features from the training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint features corresponding to the training spectrogram segment includes: Extract shallow features and deep features from the training spectrogram segment based on the voiceprint extraction model; Fuse the shallow features and the deep features based on the voiceprint extraction model to obtain fused features; Extract features from the fused features based on the voiceprint extraction model as the target features of the training spectrogram segment; Calculate the mean value of the features of each frame in the target features of the training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint feature corresponding to the spectrogram segment.
6. The voiceprint feature extraction method according to claim 3, wherein The voiceprint extraction model includes: a first feature extraction module, a second feature extraction module, a feature fusion module, a third feature extraction module, and a feature processing module, wherein the second feature extraction module includes a plurality of cascaded feature extraction sub-modules; The extracting the voiceprint feature from the training spectrogram segment based on the voiceprint extraction model to obtain the voiceprint feature corresponding to the training spectrogram segment includes: Input the training spectrogram segment into the first feature extraction module for feature extraction; Input the features output by the first feature extraction module into the second feature extraction module for feature extraction, wherein the input of the first feature extraction sub-module in the second feature extraction module is the features output by the first feature extraction module, and the input of other feature extraction sub-modules is the features output by the previous feature extraction sub-module; Input the features output by each feature extraction sub-module in the second feature extraction module into the feature fusion module for feature fusion; Input the fused features output by the feature fusion module into the third feature extraction module for feature extraction; Input the features output by the third feature extraction module into the feature processing module for processing to obtain the voiceprint feature corresponding to the training spectrogram segment output by the feature processing module, wherein the feature processing module calculates the mean value of the features of each frame in the input features.
7. A voiceprint feature extraction device, characterized in that, Includes: A spectrogram segment acquisition module, a voiceprint feature extraction module, and a voiceprint feature determination module; The spectrogram segment acquisition module is used to acquire a plurality of spectrogram segments of the target voice data; The voiceprint feature extraction module is used to extract voiceprint features from a plurality of spectrogram segments of the target voice data respectively based on a pre-trained voiceprint extraction model to obtain voiceprint features corresponding to the plurality of spectrogram segments of the target voice data, wherein the voiceprint extraction model uses a plurality of training spectrogram segments with non-shuffled time series and a plurality of training spectrogram segments with shuffled time series as training samples, and uses the true identity labels corresponding to the respective training spectrogram segments included in the training samples as sample labels, and is trained with the goal of making the identity labels predicted by the voiceprint features extracted from each training spectrogram segment included in the training samples tend to be consistent with the corresponding true identity labels; The voiceprint feature determination module is used to determine the voiceprint feature corresponding to the target voice data based on the voiceprint features corresponding to the plurality of spectrogram segments of the target voice data; The acquisition process of the training samples includes: Obtain a plurality of spectrogram segments from a pre-constructed spectrogram segment set, wherein the spectrogram segment set includes a plurality of spectrogram segments with non-shuffled time series; For each spectrogram segment obtained from the spectrogram segment set: Randomly generate the temporal scrambling probability corresponding to the spectrogram segment; If the temporal scrambling probability corresponding to the spectrogram segment is greater than the set probability threshold, perform segment-level scrambling on the time sequence of the spectrogram segment to obtain a spectrogram segment with scrambled time sequence as the training spectrogram segment; if the temporal scrambling probability corresponding to the spectrogram segment is less than or equal to the set probability threshold, use the spectrogram segment as the training spectrogram segment; Form a training sample from the obtained training spectrogram segments.
8. A voiceprint feature extraction device, characterized in that Comprising: A memory and a processor; The memory is used for storing programs; The processor is used for executing the program to implement each step of the voiceprint feature extraction method described in any one of claims 1 to 6.
9. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, each step of the voiceprint feature extraction method described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Voiceprint information extraction model generation method and device and voiceprint information extraction method and device
CN109584887A
Voiceprint vector extraction method and device, equipment and storage medium
CN113140222A
Speech recognition verification processing method and device
CN115022087A
Transform-based sound scene classification method
CN115798515A