A speaker-specific speech separation method based on dual-path self-attention mechanism
By introducing a dual-path self-attention mechanism and signal-to-noise ratio estimation module in a specific human speech separation method, the problem of poor separation effect in the prior art is solved, and a more efficient and stable speech separation effect is achieved.
Patent Information
- Application Number
- CN202210088494.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-01-25
AI Technical Summary
Existing human-specific speech separation algorithms are not effective when processing mixed corpus containing noise and multi-person speech interference, especially when training and testing conditions do not match exactly, the signal-to-noise ratio mismatch can affect the separation performance.
A specific person-specific speech separation method based on a dual-path self-attention mechanism is adopted. By obtaining registered corpus and mixed corpus, identity characteristics and speech characteristics are extracted, and then the signal-to-noise ratio estimation module and a dual-path self-attention mechanism separator are used to gradually restore the clean speech signal of the target speaker.
The separation performance in different signal-to-noise ratio scenarios is improved, the network's attention to the signal-to-noise ratio is enhanced, and the stability and accuracy of the separation effect are improved.
Smart Images

Figure CN114495973B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech separation, and in particular to a method for separating speech of a specific person based on a dual-path self-attention mechanism. Background Art
[0002] Specific speaker separation technology refers to extracting the target speaker's voice from a mixed corpus containing noise and multi-person voice interference given the target speaker's reference voice. At present, there are two major schools of specific speaker separation algorithms based on deep learning methods: time-frequency domain and time-domain. The time-frequency domain method ignores the importance of phase information in the process of speech signal reconstruction and uses the existing extracted amplitude spectrum, which limits the feature learning of the original speech signal by the separation network to a certain extent; the signal sequence obtained by the encoder of the specific speaker separation algorithm based on the time domain is often significantly longer than the sequence length obtained after the traditional short-time Fourier transform (STFT), making the network modeling and learning process more difficult. In addition, whether it is the time-frequency domain or the time domain method, the separation effect will be reduced due to the incomplete matching of training and testing conditions. Among them, the difference in the signal-to-noise ratio level (in the separation task, it refers to the ratio of the energy level of the target speaker's voice to the other interfering voices, hereinafter referred to as SNR) between the test data and the training data is a major factor affecting the separation effect. Summary of the invention
[0003] In order to solve the above technical problems, the purpose of the present invention is to provide a method for separating the speech of a specific person based on a dual-path self-attention mechanism, which can quickly and accurately extract the voice of the target speaker from a mixed corpus containing noise and multi-person voice interference.
[0004] The first technical solution adopted by the present invention is: a method for separating speech of a specific person based on a dual-path self-attention mechanism, comprising the following steps:
[0005] Obtain registered corpus and mixed corpus;
[0006] Mel-spectrogram extraction is performed on the registered corpus and input into the pre-trained speaker encoder to obtain identity features;
[0007] The mixed corpus is processed based on the pre-trained speech encoder to obtain speech features;
[0008] Fusing the identity feature and the voice feature to obtain a fused feature;
[0009] The fusion features are processed based on the pre-trained signal-to-noise ratio estimation module to obtain the signal-to-noise ratio estimation value;
[0010] The fused features and the signal-to-noise ratio estimation are sequentially passed through the pre-trained speech separator and speech decoder to obtain the clean speech signal of the target speaker.
[0011] Furthermore, the step of extracting mel-spectrograms from the registered corpus and inputting the mel-spectrograms into a pre-trained speaker encoder to obtain identity features specifically includes:
[0012] The registered corpus is subjected to frame processing, pre-emphasis processing and windowing processing in sequence to obtain a windowed signal;
[0013] Perform short-time Fourier transform on the windowed signal to obtain a linear spectrum;
[0014] Convert the linear spectrum into a Mel nonlinear spectrum to obtain a Mel spectrum;
[0015] The mel-spectrogram is processed based on the pre-trained speaker encoder to obtain the identity features.
[0016] Furthermore, the pre-trained speaker encoder includes a front-end feature extraction network, an encoding layer, and a fully connected layer.
[0017] Furthermore, the step of processing the mel spectrum based on the pre-trained speaker encoder to obtain the identity feature specifically includes:
[0018] Based on the front-end feature extraction network, the features of the speaker verification task are learned from the mel spectrum to obtain the front-end features;
[0019] Convert the front-end features into encoding vectors based on the encoding layer;
[0020] The encoded vector is processed based on the fully connected layer to obtain the speaker identity feature of fixed dimension.
[0021] Furthermore, the step of processing the mixed corpus based on the pre-trained speech encoder to obtain speech features specifically includes:
[0022] Convert the mixed corpus into speech signal sequence based on the pre-trained speech encoder;
[0023] The speech signal sequence is cut and reorganized into three-dimensional feature blocks to obtain speech features.
[0024] Furthermore, the pre-trained signal-to-noise ratio estimation module includes a hole convolution layer, an LSTM layer and a fully connected layer. The pre-trained signal-to-noise ratio estimation module processes the fusion features to obtain a signal-to-noise ratio estimation value, which specifically includes:
[0025] Extract features from fused features based on dilated convolution;
[0026] Mining the temporal information between features based on the LSTM layer;
[0027] Get the signal-to-noise ratio estimate for each frame based on the fully connected layer;
[0028] Take the average in the time dimension to get the signal-to-noise ratio estimate of the speech segment.
[0029] Furthermore, the step of sequentially passing the fusion features and the signal-to-noise ratio estimation value through a pre-trained speech separator and a speech decoder to obtain a clean speech signal of the target speaker specifically includes:
[0030] Concatenate the fusion features and the signal-to-noise ratio estimation value in the feature dimension to obtain the concatenated features;
[0031] The pre-trained dual-path self-attention mechanism speech separator separates the spliced features to obtain the separated three-dimensional feature modules;
[0032] Based on the speech decoder, the separated three-dimensional feature modules are reorganized and spliced to restore the clean speech signal of the target speaker.
[0033] Furthermore, the training steps of the speaker encoder specifically include:
[0034] Construct the first training set;
[0035] The data of the first training set are sampled, framed, pre-emphasized, windowed, Fourier transformed, and Mel-filtered to obtain a training Mel spectrum;
[0036] Using Cross Entropy Loss as the loss function, the speaker encoder is trained according to the training Mel spectrum and the true labels in the first training set to obtain a pre-trained speaker encoder.
[0037] The beneficial effects of the method of the present invention are as follows: the present invention adds a signal-to-noise ratio estimation module to the network and extracts the signal-to-noise ratio of the target speech and the interfering speech, and uses the estimated signal-to-noise ratio as one of the inputs of the separation network, so that the separation network pays attention to the signal-to-noise ratio level of the current speech segment to improve the separation performance of the network in different signal-to-noise ratio scenarios. In addition, through a separation algorithm of a dual-path self-attention mechanism, the advantages of network autonomous learning are fully utilized to improve the separation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of the steps of a method for separating speech of a specific person based on a dual-path self-attention mechanism of the present invention;
[0039] Figure 2 It is a structural block diagram of a specific embodiment method of the present invention;
[0040] Figure 3 It is a schematic diagram of the mel spectrum extraction process of a specific embodiment of the present invention. DETAILED DESCRIPTION
[0041] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only provided for the convenience of explanation and description, and the order between the steps is not limited in any way. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0042] Reference Figure 1 and Figure 2 The present invention provides a method for separating speech of a specific person based on a dual-path self-attention mechanism, the method comprising the following steps:
[0043] S1. Obtain registered corpus and mixed corpus;
[0044] S2, extract the mel spectrum from the registered corpus and input it into the pre-trained speaker encoder to obtain identity features;
[0045] Specifically, the speaker encoder is used to extract a coding vector that can characterize the identity characteristics of the target speaker from the registered audio of the target speaker. The present invention uses a text-independent speaker verification network to extract the identity characteristic coding vector of the target speaker.
[0046] S2.1, performing frame processing, pre-emphasis processing and windowing processing on the registered corpus in sequence to obtain a windowed signal;
[0047] Specifically, speech signals have short-term stability and can be regarded as quasi-static within 10 to 30 ms. Therefore, speech signals are usually framed in the time domain, and for smooth transition between frames, there is usually an overlap between frames, which is called frame shift. Here, the speech signals are all 16k by default, with a frame length of 25 ms and a frame shift of 10 ms. Due to the characteristics of human vocal structure, the high-frequency part of human speech signals will be suppressed, and the high-frequency energy attenuates greatly during the propagation of sound. In order to compensate for the deficiency of high-frequency signals, the high-frequency part of each frame signal is pre-emphasized. In order to reduce the Gibbs effect, the pre-emphasized signal is windowed using a Hamming window.
[0048] S2.2, performing short-time Fourier transform on the windowed signal to obtain a linear spectrum;
[0049] Specifically, the windowed signal is subjected to a short-time Fourier transform to convert the signal from the time domain to the frequency domain to obtain an amplitude spectrum. Here, the number of Fourier transform points is set to 512.
[0050] S2.3, converting the linear spectrum into a Mel nonlinear spectrum to obtain a Mel spectrum;
[0051] Specifically, the human ear's perception of sound frequency changes logarithmically. The linear spectrum after short-time Fourier transform is evenly spaced in the frequency domain, and the Mel spectrum conforms to the auditory characteristics of the human ear. The energy spectrum is multiplied by a set of triangular bandpass filters evenly distributed on the Mel spectrum, that is, the center frequency interval of the bandpass filter decreases as the filter index decreases and widens as the filter index increases. The logarithmic energy output of each filter is obtained, and the linear spectrum can be converted to the Mel nonlinear spectrum to obtain the required Mel spectrum. The Mel spectrum extraction process refers to Figure 3 .
[0052] S2.4. Process the mel spectrum based on the pre-trained speaker encoder to obtain identity features.
[0053] S2.4.1. Learning the features of the speaker verification task from the Mel-spectrograms based on the front-end feature extraction network to obtain the front-end features;
[0054] Specifically, the front-end feature extraction network is used to learn features suitable for speaker verification tasks from the Mel spectrum of the speech signal. ResNet-18 is used as a module for front-end feature learning.
[0055] S2.4.2, converting the front-end features into coding vectors based on the coding layer;
[0056] Specifically, the encoding layer is used to convert the features containing temporal relationships output by the front-end feature extractor into a fixed-length encoding vector that is independent of temporal relationships. At the same time, it realizes the dimensionality conversion between the feature extraction layer and the classifier and reduces the overfitting problem of the deep network.
[0057] S2.4.3. Process the encoded vector based on the fully connected layer to obtain the speaker identity feature of fixed dimension.
[0058] In addition, it also includes a classifier, which is used to classify the speaker identity features. Each node of its output represents each person in the training data. This part only exists during the training process. After the training is completed, this part is removed, and the output of the fully connected layer of the previous part is used as the speaker identity feature to input the subsequent separation network.
[0059] S3, processing the mixed corpus based on the pre-trained speech encoder to obtain speech features;
[0060] S3.1, converting the mixed corpus into a speech signal sequence based on a pre-trained speech encoder;
[0061] S3.2. Cut and reorganize the speech signal sequence into three-dimensional feature blocks to obtain speech features.
[0062] Specifically, the speech encoder is used to encode the mixed speech, simulate STFT, realize the conversion of the speech signal domain, and obtain a feature sequence with a length of L and a feature dimension of H. This process is implemented using a one-dimensional convolutional network. The speech signal sequence after convolution is usually too long, so the sequence is cut and reorganized here to obtain a three-dimensional feature block. That is, the feature sequence with a length of L and a feature dimension of H is cut into M short sequences, each of which has a length of N. There is an overlap of length P between adjacent short sequences, and K=2P, and finally a three-dimensional module with a dimension of N×M×H is obtained.
[0063] S4, fusing the identity feature and the voice feature to obtain a fused feature;
[0064] Specifically, it is used to fuse the speech features obtained by the speech encoder and the identity features obtained by the speaker encoder. The speech features and the identity features are concatenated in the feature dimension, and then the fused features are obtained through a fully connected layer. The fused features are simultaneously input into the speech separator and the signal-to-noise ratio estimation module.
[0065] S5, processing the fusion features based on the pre-trained signal-to-noise ratio estimation module to obtain a signal-to-noise ratio estimation value;
[0066] S5.1. Feature extraction of fused features based on dilated convolution;
[0067] S5.2, mining the temporal information between features based on the LSTM layer;
[0068] S5.3, obtaining the signal-to-noise ratio estimation value of each frame based on the fully connected layer;
[0069] S5.4. Take the average in the time dimension to obtain the signal-to-noise ratio estimate of the speech segment.
[0070] Specifically, the signal-to-noise ratio estimation module is used to estimate the signal-to-noise ratio in the corpus. The signal-to-noise ratio estimation module in the present invention is implemented by three layers of two-dimensional hole convolution, one layer of LSTM, and one layer of full connection. The hole convolution is used for feature extraction, the LSTM is used to explore the timing information between features, and the fully connected layer is used to obtain the signal-to-noise ratio estimation value of each frame. Finally, the mean is taken in the time dimension to smooth the signal-to-noise ratio of the entire segment. In the training phase, this module performs multi-task training together with the rest of the parts; in the testing phase, this module, together with the previous speaker encoder, speech encoder and feature fusion module, extracts the signal-to-noise ratio estimation value (SNR) in the current speech segment from the mixed corpus, replacing the GroundTruth SNR in the training phase as one of the inputs of the speech separator.
[0071] S6. The fused features and the signal-to-noise ratio estimation values are sequentially passed through a pre-trained speech separator and a speech decoder to obtain a clean speech signal of the target speaker.
[0072] S6.1, concatenating the fused features and the signal-to-noise ratio estimation value in the feature dimension to obtain a concatenated feature;
[0073] S6.2, based on the pre-trained dual-path self-attention mechanism speech separator, the spliced features are separated to obtain the separated three-dimensional feature module;
[0074] Specifically, the speech separator is used to separate the three-dimensional features corresponding to the target speaker, which is implemented by a dual-path self-attention mechanism module stacked with B layers. The fused features are first concatenated with the SNR of the current speech segment in the feature dimension, that is, the SNR is repeated N×M times first, and then concatenated in the fused feature dimension. Each dual-path self-attention mechanism module consists of self-attention mechanisms on two paths, intra-block and inter-block, and the parameters of each module are set the same. The self-attention mechanism allows the network to discover the correlation between its own sequences. First, the query matrix Q, key matrix K and value matrix V are generated according to the input feature vector. Then Q and K are matrix multiplied to obtain the correlation score matrix R between a certain moment and other moments in the sequence. R is normalized to between 0 and 1 by the Softmax function, and then matrix multiplied with V to obtain the value of the self-attention output. In this paper, the self-attention mechanism is implemented by a multi-head attention mechanism, which consists of h parallel self-attention modules. The results of each self-attention module are concatenated to obtain the final output. Intra-block self-attention regards N as the time dimension, which is equivalent to calculating the local attention of each short sequence, while inter-block self-attention regards M as the time dimension, which is equivalent to calculating the global attention of the entire long sequence. This combination of local and global methods can effectively model the long sequence information. In the training phase, the SNR of the speech segment uses the Ground Truth SNR, that is, the SNR value when generating the mixed corpus for training; in the testing phase, the SNR uses the SNR value calculated by the signal-to-noise ratio estimation module.
[0075] S6.2. Based on the speech decoder, the three-dimensional feature modules are reorganized and spliced to restore the clean speech signal of the target speaker.
[0076] Specifically, the speech decoder is used to restore the three-dimensional feature module obtained by the speech separator to obtain the clean audio corresponding to the target speaker, which is implemented by a one-dimensional deconvolution network. The three-dimensional module is reassembled and spliced according to the reverse process of cutting and splicing in the speech encoder to obtain a feature sequence of the same length as the output of the speech encoder, and then sent to the one-dimensional deconvolution network to restore the clean speech signal of the target speaker.
[0077] As a further preferred embodiment of the method, the training step of the speaker encoder specifically includes:
[0078] Construct the first training set DatasetA;
[0079] The data of the first training set DatasetA is sampled, framed, pre-emphasized, windowed, Fourier transformed, and Mel-filtered to obtain a training Mel spectrum;
[0080] Specifically, 2. The frame length is set to 25ms, the frame shift is set to 10ms, the pre-emphasis coefficient is set to 0.97, the window function uses the Hamming window, the number of Fourier transform points is set to 512, and the Mel filter coefficient is set to 64;
[0081] Using Cross Entropy Loss as the loss function, the speaker encoder is trained according to the training Mel spectrum and the true labels in the first training set to obtain a pre-trained speaker encoder.
[0082] As a further preferred embodiment of the method, the training step of the speech separator specifically comprises:
[0083] Construct the second training set DatasetB;
[0084] Two speakers A and B are randomly selected from the second training set DatasetB each time. A is used as the target speaker. An audio clip is selected from the corpus corresponding to A as the registration corpus to extract the identity feature vector E from the trained speaker encoder. Then an audio clip is randomly selected from the remaining corpus of A as the target restored speech signal S1. B is used as the interference speaker. An audio clip is randomly selected from B's corpus as the interference speech S2, which is mixed with S1 at a certain signal-to-noise ratio (-5 to 10 dB) as the training data for separation (all speech clips used for training are cropped to 3s). The signal-to-noise ratio data for mixing is also input into the model.
[0085] The speech separator is trained using SI-SNR (scale-invariant signal-to-noise ratio) as the loss function.
[0086] In addition, the training steps for other modules are similar.
[0087] SI-SNR and L1 Loss are used as the loss functions for the separation module and the signal-to-noise ratio estimation module respectively. The sum of the results is used as the loss function of the entire model to jointly optimize the two modules. Adam is used as the optimizer, and the initial learning rate is set to 0.001. When the loss of the validation set does not drop for more than 10 epochs, the learning rate is adjusted. Iterate 50 times to complete the training and save the model.
[0088] A speaker-specific speech separation system based on a dual-path self-attention mechanism, comprising:
[0089] Data acquisition module, used to acquire registered corpus and mixed corpus;
[0090] The identity feature extraction module is used to extract the mel-spectrogram from the registered corpus and input it into the pre-trained speaker encoder to obtain the identity feature;
[0091] The speech feature extraction module processes the mixed corpus based on the pre-trained speech encoder to obtain speech features;
[0092] A fusion module is used to fuse the identity feature and the voice feature to obtain a fusion feature;
[0093] A signal-to-noise ratio estimation module processes the fusion features based on the pre-trained signal-to-noise ratio estimation module to obtain a signal-to-noise ratio estimation value;
[0094] The separation module is used to sequentially pass the fused features and the signal-to-noise ratio estimation values through the pre-trained speech separator and speech decoder to obtain the clean speech signal of the target speaker.
[0095] As a further preferred embodiment of the system, it also includes:
[0096] The training module is used to train the speaker encoder, the speech encoder, the signal-to-noise ratio estimation module, the speech separator and the speech decoder.
[0097] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0098] A person-specific speech separation device based on dual-path self-attention mechanism:
[0099] at least one processor;
[0100] at least one memory for storing at least one program;
[0101] When the at least one program is executed by the at least one processor, the at least one processor implements the method for separating specific person's speech based on the dual-path self-attention mechanism as described above.
[0102] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0103] A storage medium storing processor-executable instructions, characterized in that the processor-executable instructions are used to implement a method for separating speech of a specific person based on a dual-path self-attention mechanism as described above when executed by the processor.
[0104] The contents of the above method embodiments are all applicable to the present storage medium embodiments. The functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0105] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A method for separating speech of a specific person based on a dual-path self-attention mechanism, characterized in that: The following steps are involved: Obtain registered corpus and mixed corpus; Mel-spectrogram extraction is performed on the registered corpus and input into the pre-trained speaker encoder to obtain identity features; The mixed corpus is processed based on the pre-trained speech encoder to obtain speech features; Fusing the identity feature and the voice feature to obtain a fused feature; The fusion features are processed based on the pre-trained signal-to-noise ratio estimation module to obtain the signal-to-noise ratio estimation value; The fused features and the signal-to-noise ratio estimation are sequentially passed through the pre-trained speech separator and speech decoder to obtain the clean speech signal of the target speaker.
2. According to claim 1, a method for separating speech of a specific person based on a dual-path self-attention mechanism is characterized in that: The step of extracting mel-spectrograms from the registered corpus and inputting the mel-spectrograms into a pre-trained speaker encoder to obtain identity features specifically includes: The registered corpus is subjected to frame processing, pre-emphasis processing and windowing processing in sequence to obtain a windowed signal; Perform short-time Fourier transform on the windowed signal to obtain a linear spectrum; Convert the linear spectrum into a Mel nonlinear spectrum to obtain a Mel spectrum; The mel-spectrogram is processed based on the pre-trained speaker encoder to obtain the identity features.
3. According to claim 2, a method for separating speech of a specific person based on a dual-path self-attention mechanism is characterized in that: The pre-trained speaker encoder includes a front-end feature extraction network, an encoding layer, and a fully connected layer.
4. According to claim 3, a method for separating speech of a specific person based on a dual-path self-attention mechanism is characterized in that: The step of processing the mel spectrum based on the pre-trained speaker encoder to obtain the identity feature specifically includes: Based on the front-end feature extraction network, the features of the speaker verification task are learned from the mel spectrum to obtain the front-end features; Convert the front-end features into encoding vectors based on the encoding layer; The encoded vector is processed based on the fully connected layer to obtain the speaker identity feature of fixed dimension.
5. According to claim 4, a method for separating speech of a specific person based on a dual-path self-attention mechanism is characterized in that: The step of processing the mixed corpus based on the pre-trained speech encoder to obtain speech features specifically includes: Convert the mixed corpus into speech signal sequence based on the pre-trained speech encoder; The speech signal sequence is cut and reorganized into three-dimensional feature blocks to obtain speech features.
6. According to claim 5, a method for separating speech of a specific person based on a dual-path self-attention mechanism is characterized in that: The pre-trained signal-to-noise ratio estimation module includes a hole convolution layer, an LSTM layer and a fully connected layer. The pre-trained signal-to-noise ratio estimation module processes the fusion features to obtain the signal-to-noise ratio estimation value, which specifically includes: Extract features from fused features based on dilated convolution; Mining the temporal information between features based on the LSTM layer; Get the signal-to-noise ratio estimate for each frame based on the fully connected layer; The average is taken in the time dimension to obtain the signal-to-noise ratio estimation value of the speech segment corresponding to the fusion feature.
7. According to claim 6, a method for separating speech of a specific person based on a dual-path self-attention mechanism is characterized in that: The step of sequentially passing the fusion features and the signal-to-noise ratio estimation value through a pre-trained speech separator and a speech decoder to obtain a clean speech signal of the target speaker specifically includes: Concatenate the fusion features and the signal-to-noise ratio estimation value in the feature dimension to obtain the concatenated features; The pre-trained dual-path self-attention mechanism speech separator separates the spliced features to obtain the separated three-dimensional feature modules; Based on the speech decoder, the separated three-dimensional feature modules are reorganized and spliced to restore the clean speech signal of the target speaker.
8. According to claim 7, a method for separating speech of a specific person based on a dual-path self-attention mechanism is characterized in that: The training steps of the speaker encoder specifically include: Construct the first training set; The data of the first training set are sampled, framed, pre-emphasized, windowed, Fourier transformed, and Mel-filtered to obtain a training Mel spectrum; Using Cross Entropy Loss as the loss function, the speaker encoder is trained according to the training Mel spectrum and the true labels in the first training set to obtain a pre-trained speaker encoder.
Citation Information
Patent Citations
Single-channel voice separation method and device and electronic equipment
CN111429938A
Voiceprint recognition method based on deep learning algorithm
CN112712814A