Voiceprint recognition method and system, terminal and medium
Through logarithmic Meier spectrum feature extraction and embedded feature extraction module based on pre-trained models, combining two-dimensional convolution, channel-by-channel convolution, frequency self-attention and time delay mechanisms, the existing voiceprint recognition technology has large calculation volume, high equipment requirements and low recognition accuracy, and efficient and accurate voiceprint recognition is achieved.
Patent Information
- Application Number
- CN202510177241.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-24
AI Technical Summary
The existing voiceprint recognition technology has problems such as excessive data calculation, high equipment capability requirements and low recognition accuracy.
Embedded feature vectors are generated by logarithmic Mel spectrum feature extraction on the speech data to be recognized by the current speaker that is acquired, and based on the embedded feature extraction module of the pre-trained speaker classification model. This module includes two-dimensional convolution processing, channel-by-channel convolution processing, frequency self-attention mechanism and time delay mechanism, which are used to extract multi-channel multi-scale time-frequency features and perform weight calibration.
It significantly improves the feature learning expressiveness and accuracy of the voiceprint recognition model, reduces the computational complexity and parameter volume, and reduces the computational cost, making it suitable for deployment in resource-constrained hardware environments.
Smart Images

Figure CN120199255A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and particularly to a voiceprint recognition method, system, terminal and medium. Background Art
[0002] Speech recognition is a technology that converts human speech into text or commands and is widely used in various terminal devices. Among them, voiceprint recognition refers to a technology that extracts the voiceprint information of a speaker through a section of speech, so as to confirm or distinguish the identity of the speaker. In recent years, with the rapid development of artificial intelligence technology, voiceprint recognition technology based on deep learning models has been widely used.
[0003] Currently, most of the deep learning models in the existing voiceprint recognition technology solutions based on deep learning models adopt the residual network model architecture. Specifically, a speaker embedding layer model is constructed using the residual network model architecture, and during the residual calculation process, through convolution and deconvolution operations, the output features of the residual module are made to be consistent with the input features in the time scale, so as to ensure that important information is not lost in the time domain of the features. Or a series of models of the Visual Geometry Group (VGG) are used as deep learning models to construct a speaker embedding layer model to extract the embedding vector of the speech. During the model training stage, by combining the kernel latent Dirichlet allocation and mutual information methods, the feature differences between positive and negative samples are maximized, so as to be used to optimize the parameter update of the model.
[0004] However, when performing voiceprint recognition tasks using the above types of deep learning models, the model lacks the ability to differentially process the importance of different time dimensions and frequency dimensions in the feature map, and cannot effectively cope with the complex changes of speech features in the time domain and frequency domain, and it is difficult to equally capture frequency domain information and time domain information, low-frequency information and high-frequency information, resulting in limited speaker recognition accuracy; moreover, the number of parameters and computational requirements of the model are large, and it is not suitable for deployment on embedded device terminals, such as resource-constrained hardware environments, thus limiting the application scope of the model.
[0005] Therefore, the existing voiceprint recognition technology still has technical problems such as excessive data calculation volume, high requirements for device capabilities, and low recognition accuracy. Summary of the Invention
[0006] In view of the above-mentioned disadvantages of the prior art, the purpose of this application is to provide a voiceprint recognition method, system, terminal and medium, which are used to solve the technical problems that the existing voiceprint recognition technology still has excessive data calculation volume, high requirements for device capabilities, and low recognition accuracy.
[0007] To achieve the above and other related objectives, a first aspect of the present application provides a voiceprint recognition method, which includes: extracting log Mel spectrogram features from the voice data to be recognized of the current speaker to generate Mel filter bank feature data of the voice data to be recognized; based on the embedding feature extraction module of a pre-trained speaker classification model, extracting embedding features from the Mel filter bank feature data to generate an embedding feature vector of the voice data to be recognized; calculating the similarity between the embedding feature vector and the template embedding feature vectors of one or more registered users respectively, and screening the target embedding feature vectors that match the embedding feature vector to verify that the current speaker is a registered user and identify the target registered user; wherein, based on the embedding feature extraction module, the way of extracting the embedding feature vector from the Mel filter bank feature data includes: through two-dimensional convolution processing of the Mel filter bank feature data, extracting local time-frequency feature information in the Mel filter bank feature data to generate corresponding multi-channel time-frequency feature data; dividing the multi-channel time-frequency feature data into multiple groups according to the feature channels, and performing per-channel convolution processing on each group of channel feature data respectively to extract the time-frequency features of the multi-channel time-frequency feature data at different scales to generate corresponding multi-channel multi-scale time-frequency feature data; by adopting a self-attention mechanism in the frequency dimension of the multi-channel multi-scale time-frequency feature data, adaptively calibrating the weight distribution of each time-frequency segment of the multi-channel multi-scale time-frequency feature data to generate corresponding multi-channel multi-scale enhanced time-frequency feature data; by adopting a time delay mechanism to capture the temporal dependence relationship in the multi-channel multi-scale enhanced time-frequency feature data, and based on the weight distribution of each time-frequency segment of the multi-channel multi-scale time-frequency feature data, linearly weighting the features in different frequency ranges to generate corresponding multi-channel global time-frequency feature data; performing an aggregation operation on all the frame-level features in the multi-channel global time-frequency feature data to compress the information in the time dimension into a fixed-length sentence-level feature vector; performing a linear transformation on the sentence-level feature vector and mapping the sentence-level feature vector into an embedding space for representing the voiceprint information of the speaker to generate a corresponding embedding feature vector.
[0008] In some embodiments of the first aspect of the present application, the embedding feature extraction module includes: a local time-frequency feature extraction unit, including: a first separable two-dimensional convolutional layer, a first batch normalization layer, a first activation function, a two-dimensional residual network, a second batch normalization layer, a second activation function, a second separable two-dimensional convolutional layer, a third batch normalization layer, a third activation function, and a frequency self-attention network connected in sequence; an embedding feature vector generation unit, including: a time delay neural network, a first global average pooling layer, a fourth batch normalization layer, a first linear layer, a fifth batch normalization layer, and a second linear layer connected in sequence.
[0009] In some embodiments of the first aspect of the present application, the two-dimensional residual network includes: a first two-dimensional convolutional layer, a second two-dimensional convolutional layer, and a third two-dimensional convolutional layer; the two-dimensional residual network is used to divide the multi-channel time-frequency feature data into four groups according to the feature channels, and perform per-channel convolution processing on each group of channel feature data respectively; the specific method includes: dividing the multi-channel time-frequency feature data into four groups according to the feature channels to obtain a first group of channel feature data, a second group of channel feature data, a third group of channel feature data, and a fourth group of channel feature data; outputting the first group of channel feature data as first-scale feature data; performing convolution processing on the second group of channel feature data through the first two-dimensional convolutional layer to output second-scale feature data; merging the third group of channel feature data with the second-scale feature data and then performing convolution processing through the second two-dimensional convolutional layer to output third-scale feature data; merging the fourth group of channel feature data with the third-scale feature data and then performing convolution processing through the third two-dimensional convolutional layer to output fourth-scale feature data; performing data fusion on the first-scale feature data, the second-scale feature data, the third-scale feature data, and the fourth-scale feature data to obtain multi-channel multi-scale time-frequency feature data.
[0010] In some embodiments of the first aspect of the present application, the frequency self-attention network includes: a second global average pooling layer, a first one-dimensional convolutional layer, and a fourth activation function connected in sequence; the frequency self-attention network is used to adaptively calibrate the weights of the multi-channel multi-scale time-frequency feature data in each time-frequency segment by using an attention mechanism to generate corresponding multi-channel multi-scale enhanced time-frequency feature data; the specific method includes: performing average pooling operation on the multi-channel multi-scale time-frequency feature data in the feature channel dimension through the second global average pooling layer to output corresponding single-channel multi-scale time-frequency feature data; performing self-attention calculation on the single-channel multi-scale time-frequency feature data through the first one-dimensional convolutional layer and the fourth activation function to output the weight distribution of the multi-channel multi-scale time-frequency feature data in each time-frequency segment; multiplying the multi-channel multi-scale time-frequency feature data by the weight distribution in each time-frequency segment to output corresponding multi-channel multi-scale enhanced time-frequency feature data.
[0011] In some embodiments of the first aspect of the present application, the time-delay neural network includes: a separable one-dimensional convolutional layer, a sixth batch normalization layer, a fifth activation function, a first one-dimensional residual block, a second one-dimensional residual block, a third one-dimensional residual block, a second one-dimensional convolutional layer, a third global average pooling layer, and a seventh batch normalization layer connected in sequence; the first one-dimensional residual block and the second one-dimensional residual block are also connected to the second one-dimensional convolutional layer; wherein, the time-delay neural network captures the temporal dependence relationship in the multi-channel multi-scale enhanced time-frequency feature data by adopting a time-delay mechanism, and extracts the global time-frequency features in the multi-channel multi-scale enhanced time-frequency feature data through the first one-dimensional residual block, the second one-dimensional residual block, and the third one-dimensional residual block; through the second one-dimensional convolutional layer, according to the weight distribution of each time-frequency segment of the multi-channel multi-scale time-frequency feature data, the features in different frequency ranges are linearly weighted to generate corresponding multi-channel global time-frequency feature data.
[0012] In some embodiments of the first aspect of the present application, the method for training the speaker classification model includes: randomly cropping a plurality of acquired audio files into a specified length, padding the audio files with insufficient length to the specified length to generate a plurality of audio segment data, adding random noise to each audio segment data to generate a plurality of speech training samples; marking the true speaker of each speech training sample, and respectively extracting the logarithmic Mel spectrogram features of each speech training sample to generate the Mel filter bank feature training data of each speech training sample; inputting the Mel filter bank feature training data into the speaker classification model, respectively extracting the embedding features of each Mel filter bank feature training data through the embedding feature extraction module of the speaker classification model to generate the embedding feature vectors of each speech training sample, and predicting the speaker of each speech training sample through the classification module of the speaker classification model; calculating the difference between the predicted speakers and the corresponding true speakers by using a classification loss function, calculating the gradient of the classification loss function with respect to the speaker classification model, and using an optimizer to update the corresponding model parameters until a converged speaker classification model is obtained.
[0013] In some embodiments of the first aspect of the present application, the method for calculating the similarity between the embedded feature vector and the template embedded feature vectors of one or more registered users respectively, screening the target embedded feature vector that matches the embedded feature vector to verify that the current speaker is a registered user, and identifying the target registered user includes: successively calculating the cosine similarity between the embedded feature vector and the template embedded feature vectors of each registered user, and obtaining the template embedded feature vector with the largest cosine similarity value as the primary target embedded feature vector; wherein, the template embedded feature vectors of each registered user are stored in a pre-constructed voice database; judging whether the cosine similarity value between the embedded feature vector and the primary target embedded feature vector is greater than a preset cosine similarity threshold; if the cosine similarity value is greater than or equal to the cosine similarity threshold, determining the primary target embedded feature vector as the target embedded feature vector, determining that the current speaker and the registered user corresponding to the target embedded feature vector are the same speaker, and identifying the registered user as the target registered user; if the cosine similarity value is less than the cosine similarity threshold, determining that the current speaker is an unregistered user.
[0014] To achieve the above object and other related objects, a second aspect of the present application provides a voiceprint recognition system, and the voiceprint recognition system includes: a first feature extraction module, configured to perform logarithmic Mel spectrogram feature extraction on the to-be-recognized voice data of the current speaker to generate Mel filter bank feature data of the to-be-recognized voice data; a second feature extraction module, connected to the first feature extraction module, configured to perform embedded feature extraction on the Mel filter bank feature data based on the embedded feature extraction module of the pre-trained speaker classification model to generate an embedded feature vector of the to-be-recognized voice data; a voiceprint recognition module, connected to the second feature extraction module, configured to calculate the similarity between the embedded feature vector and the template embedded feature vectors of one or more registered users respectively, and screen the target embedded feature vectors that match the embedded feature vector to verify that the current speaker is a registered user and identify the target registered user; wherein, based on the embedded feature extraction module, the manner of performing embedded feature extraction on the Mel filter bank feature data to generate the embedded feature vector of the to-be-recognized voice data includes: performing two-dimensional convolution processing on the Mel filter bank feature data to extract local time-frequency feature information in the Mel filter bank feature data to generate corresponding multi-channel time-frequency feature data; dividing the multi-channel time-frequency feature data into multiple groups according to the feature channels, and performing per-channel convolution processing on each group of channel feature data respectively to extract the time-frequency features of the multi-channel time-frequency feature data at different scales to generate corresponding multi-channel multi-scale time-frequency feature data; by adopting a self-attention mechanism in the frequency dimension of the multi-channel multi-scale time-frequency feature data, adaptively calibrating the weight distribution of each time-frequency segment of the multi-channel multi-scale time-frequency feature data to generate corresponding multi-channel multi-scale enhanced time-frequency feature data; by adopting a time delay mechanism to capture the temporal dependence relationship in the multi-channel multi-scale enhanced time-frequency feature data, and linearly weighting the features in different frequency ranges based on the weight distribution of each time-frequency segment of the multi-channel multi-scale time-frequency feature data to generate corresponding multi-channel global time-frequency feature data; performing an aggregation operation on all the frame-level features in the multi-channel global time-frequency feature data to compress the information in the time dimension into a fixed-length sentence-level feature vector; performing a linear transformation on the sentence-level feature vector and mapping the sentence-level feature vector into an embedding space for representing the voiceprint information of the speaker to generate a corresponding embedded feature vector.
[0015] To achieve the above object and other related objects, a third aspect of the present application provides a voiceprint recognition terminal, and the voiceprint recognition terminal includes: a processor and a memory; the memory is used for storing a computer program; the processor is used for executing the computer program stored in the memory so that the terminal executes any one of the voiceprint recognition methods provided in the above embodiments.
[0016] To achieve the above and other related objectives, a fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the voiceprint recognition methods provided in the above embodiments.
[0017] As described above, the present application provides a voiceprint recognition method, system, terminal, and medium. By extracting the logarithmic Mel spectrum features from the voice data to be recognized of the current speaker and further extracting the embedding features based on the speaker classification model, an embedding feature vector is generated for calculating the similarity with the template embedding feature vectors of one or more registered users to verify whether the current speaker is a registered user. Therefore, the present application has the following beneficial effects: By introducing the frequency self-attention mechanism, the speaker classification model can adaptively adjust the attention to different frequency features, accurately identify the key frequency information, and significantly improve the performance and accuracy of the model in feature learning; By introducing the multi-dimensional hybrid convolution strategy, the model has the dual capabilities of capturing frequency details and modeling temporal dependencies; By using the per-channel separable convolution strategy, the computational complexity and the number of parameters of the model are reduced, and the computational cost is reduced. Thus, the present application solves the technical problems of excessive data calculation amount, high requirements for device capabilities, and low recognition accuracy existing in the existing voiceprint recognition technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It shows a schematic flowchart of the voiceprint recognition method in an embodiment of the present application.
[0019] Figure 2 It shows a schematic structural diagram of the speaker classification model in an embodiment of the present application.
[0020] Figure 3 It shows a schematic structural diagram of the two-dimensional residual network in an embodiment of the present application.
[0021] Figure 4 It shows a schematic principle diagram of the frequency self-attention network in an embodiment of the present application.
[0022] Figure 5 It shows a schematic structural diagram of the time-delay neural network in an embodiment of the present application.
[0023] Figure 6 It shows a schematic structural diagram of the voiceprint recognition system in an embodiment of the present application.
[0024] Figure 7 It shows a schematic structural diagram of the voiceprint recognition terminal in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The following describes the implementation manners of the present application through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0026] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical items or similar items with basically the same functions and roles. For example, the first batch of normalization layers and the second batch of normalization layers are only used to distinguish different batch normalization layers, and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily limit being different.
[0027] To solve the problems in the above background art, the present invention provides a voiceprint recognition method, system, terminal and medium, aiming to solve the technical problems that the existing voiceprint recognition technology still has excessive data calculation amount, high requirements for device capabilities and low recognition accuracy.
[0028] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be further described in detail through the following embodiments in combination with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the invention.
[0029] As Figure 1 shown, a flowchart of a voiceprint recognition method in an embodiment of the present invention is shown. The voiceprint recognition method in this embodiment mainly includes the following steps.
[0030] Step S1: Extract logarithmic Mel-spectrum features from the to-be-recognized voice data of the current speaker to generate Mel-filter bank feature data of the to-be-recognized voice data.
[0031] In one embodiment, first extracting logarithmic Mel-spectrum features from the to-be-recognized voice data of the current speaker can extract Mel-filter bank features related to the auditory characteristics of the human ear from the to-be-recognized voice data, so as to be used for subsequently generating an embedding feature vector with more comprehensive, accurate and critical voiceprint information of the current speaker, which is of great significance for speech recognition. Specifically, the method for extracting logarithmic Mel-spectrum features from the to-be-recognized voice data includes the following steps.
[0032] ① Perform a pre-emphasis operation on the to-be-recognized voice data to generate corresponding high-frequency spectrum audio data.
[0033] The spectral characteristics of speech data usually have higher energy in the low-frequency components and lower energy in the high-frequency components. Moreover, during signal transmission, high-frequency components are more likely to be affected by noise and attenuate. Performing pre-emphasis operation on the speech data to be recognized can, on the one hand, increase the energy of the high-frequency components, making the entire spectrum flatter; on the other hand, it can enhance the energy of the high-frequency components, thereby improving the signal-to-noise ratio of the speech data to be recognized and better reflecting the voiceprint characteristics of the speech data to be recognized.
[0034] The pre-emphasis operation is usually implemented by a second-order high-pass filter, and the transfer function of the filter is shown in formula (1).
[0035] y(t) = x(t) - αx(t - 1); (1)
[0036] Where y(t) is the current sample of the high-frequency spectrum audio data generated after pre-emphasis; x(t) and x(t - 1) are the current sample and the previous sample of the input speech data to be recognized respectively; α is the filter coefficient. In a preferred embodiment, the value of α ranges from 0.9 to 1.
[0037] ② Perform a framing operation on the high-frequency spectrum audio data to generate a number of frame audio data.
[0038] In most cases, speech data is non-stationary and will lose the signal frequency profile over time. Therefore, it is meaningless to perform a fast Fourier transform on the entire speech data. However, speech data is a short-time stationary signal, so a fast Fourier transform can be performed on short-time frames, and a good approximation of the signal frequency profile can be obtained by connecting adjacent frames. After performing a framing operation on the high-frequency spectrum audio data, a number of frame audio data are obtained. In an embodiment, the frame length of each frame audio data is set to be about 20 - 40 ms. And to avoid excessive changes in the frame audio data of two adjacent frames, there is an overlapping area, that is, a frame shift, between the frame audio data of two adjacent frames. The frame shift is generally set to be about 10 ms.
[0039] ③ Perform a windowing operation on each frame audio data respectively to generate a number of corresponding frame windowed audio data.
[0040] After framing the high-frequency spectrum audio data, multiply each obtained frame audio data by a window function to increase the continuity at both ends of the frame, offset the assumption of infinite data in the subsequent fast Fourier transform, and reduce spectral leakage. In an embodiment, a Hamming window function can be used to perform the windowing operation. The mathematical expressions of the Hamming window function are shown in formula (2) and formula (3).
[0041] s(t) = W(n,a) × z(t); (2)
[0042] W(n,a) = 1 - a - a×cos(2πn) / (N - 1); (3)
[0043] Where 0 ≤ n ≤ N - 1, N is the window length, i.e., the number of samples; z(t) is the current sample of the frame audio data; s(t) is the current sample of the windowed frame audio data; a is the coefficient of the Hamming window function.
[0044] ④ Perform a fast Fourier transform operation on each windowed frame audio data to generate the linear spectral data of each frame of audio data.
[0045] Since it is difficult to observe the signal characteristics of speech data in the time domain, a fast Fourier transform is usually required to convert the finite-length sequence data from the time domain to the frequency domain, and different speech characteristics are extracted by observing the energy distribution in the frequency domain dimension.
[0046] Specifically, the calculation formula for performing a fast Fourier transform on each windowed frame audio data is shown in formula (4).
[0047]
[0048] And calculate the power spectrum P of the S(k) frequency-domain signal data according to formula (5).
[0049]
[0050] Where 0 ≤ k ≤ N - 1, N is the number of calculation points of the fast Fourier transform, which is set to 256 or 512 in a preferred embodiment; s(n) is the windowed frame audio data, representing the time-domain signal; S(k) is the converted linear spectral data, representing the frequency-domain signal; P is the power spectrum of the linear spectral data.
[0051] ⑤ Perform a Mel filtering operation on each linear spectral data using multiple Mel filters to generate the Mel spectral data of each frame of audio data.
[0052] The Mel filter simulates the human auditory perception system, i.e., the human ear. The human ear has different sensitivities to signals of different frequencies and only pays attention to certain specific frequency components. Therefore, the Mel filter is mainly used to allow only signal data of certain frequencies to pass through while filtering out those frequency signals that are not wanted. However, since the distribution of Mel filters on the frequency coordinate axis is not uniform, there are many filters in the low-frequency region, with a relatively dense distribution; in the high-frequency region, the number of filters is small and the distribution is sparse. Therefore, multiple Mel filters can be used to form a Mel filter bank, so as to simulate the non-linear perception of sound by the human ear, convert the spectrum of linear frequency to the spectrum of Mel frequency, and generate the corresponding Mel spectral data. The specific calculation steps are as follows.
[0053] First, determine the Mel frequency range to be covered. For speech data, the Mel frequency range can be from a low frequency of 0 to a high frequency.
[0054] The calculation formula for Mel frequency is shown in formula (6).
[0055]
[0056] Among them, f is the physical frequency; f mel is the Mel frequency.
[0057] Secondly, uniformly select M + 2 points within the Mel frequency range as Mel frequency points, m(0), m(1), …, m(M + 1). Among them, M is the number of filters. The Mel frequency points determine the center frequencies of the filters.
[0058] Then, convert these Mel frequency points back to physical frequencies using the inverse Mel frequency conversion formula (such as formula (7)).
[0059]
[0060] Finally, construct triangular filters to complete the Mel filtering operation. The calculation formula is shown in formula (8).
[0061]
[0062] The Mel filter bank includes M triangular filters. Among them, H m (k) is the frequency response of the triangular filter; f(m) is the center frequency of the m-th filter; the response at the center frequency of the triangular filter is 1 and linearly decreases to 0 until the responses at the center frequencies of two adjacent filters are 0; the interval between each f(m) widens as the value of m increases.
[0063] ⑥ Perform a logarithm operation on each Mel spectrum data respectively to generate the logarithmic Mel spectrum data of each frame of audio data. After information aggregation, generate the Mel filter bank feature data of the speech data to be recognized.
[0064] To improve the robustness of the Mel spectrum data to a certain degree of noise, it is necessary to compress each Mel spectrum data through logarithmic energy calculation. The specific calculation formula is shown in formula (9).
[0065] FBanks = 20 × log 10 (H m (k) × P); (9)
[0066] Among them, H m(k) is the frequency response of the triangular filter; P is the power spectrum of the linear spectral data; FBanks is the Mel filter bank feature data. The Mel filter bank feature data is a single-channel two-dimensional feature data of 1×N×F, where its height N is the number of Mel filters used when performing the Mel filtering operation, and its width F is the number of frames when performing the framing operation.
[0067] Step S2: Based on the embedding feature extraction module of the pre-trained speaker classification model, perform embedding feature extraction on the Mel filter bank feature data to generate an embedding feature vector of the speech data to be recognized.
[0068] In one embodiment, as Figure 2 shown, the speaker classification model includes an embedding feature extraction module and a classification module. Among them, the embedding feature extraction module is used to perform embedding feature extraction on the Mel filter bank feature data to generate an embedding feature vector of the speech data to be recognized; the classification module is used to perform a speaker category prediction task according to the extracted embedding feature vector during the model training phase, and calculate the loss of the model until the optimization of the model, so as to obtain a converged speaker classification model.
[0069] As Figure 2 shown, the embedding feature extraction module includes: a local time-frequency feature extraction unit and an embedding feature vector generation unit. The local time-frequency feature extraction unit includes: a first separable two-dimensional convolutional layer, a first batch normalization layer, a first activation function, a two-dimensional residual network, a second batch normalization layer, a second activation function, a second separable two-dimensional convolutional layer, a third batch normalization layer, a third activation function, and a frequency self-attention network connected in sequence.
[0070] In a specific embodiment, the first separable two-dimensional convolutional layer and the second separable convolutional layer decompose the traditional two-dimensional convolutional operation into two independent steps: depthwise separable convolution and pointwise convolution. The depthwise separable convolution performs a convolution operation on each input channel using a separate 3×3 convolutional layer; the pointwise convolution uses a 1×1 convolutional layer to adjust the number of output channels to the required number. Through the separable two-dimensional convolutional layer with a smaller receptive field, local time-domain information and frequency-domain information in the input Mel filter bank feature data can be extracted, so that by using repeated but less frequency-offset local frequency-domain features, the speaker classification model can effectively model finer frequency details with fewer feature channels. Moreover, the per-channel separable convolution strategy of depthwise separable convolution and pointwise convolution not only maintains or approaches the performance of traditional two-dimensional convolution, but also effectively reduces the complexity of traditional two-dimensional convolution calculation, greatly reducing the computational cost, enabling the speaker classification model to run more efficiently, and being suitable for deployment on resource-limited embedded devices, achieving a balance between performance and efficiency, and supporting real-time computing to meet the requirements of real-time applications.
[0071] As shown Figure 3 in the figure, the two-dimensional residual network includes: a plurality of two-dimensional convolutional layers, which are used to divide the input multi-channel two-dimensional feature data into multiple groups according to the feature channels, and perform per-channel convolutional processing on each group of channel feature data respectively, so as to endow the speaker classification model with the ability to extract time-frequency features at different scales. Compared with the traditional convolution method, this grouped convolution method not only retains the multi-scale feature expression, but also greatly reduces the computational complexity and the number of parameters. Therefore, the two-dimensional residual network can not only capture the relationship between detailed features and global information more flexibly, enhance the model expressiveness, but also significantly improve the efficiency and performance of the speaker classification model and optimize the computational cost in the case of limited computing resources.
[0072] In one embodiment, the two-dimensional residual network includes: a first two-dimensional convolutional layer, a second two-dimensional convolutional layer, and a third two-dimensional convolutional layer, which can divide the input multi-channel two-dimensional feature data into four groups, and perform per-channel convolutional processing on each group of channel feature data respectively. The specific method includes: dividing the input multi-channel two-dimensional feature data into four groups to obtain the first group of channel feature data, the second group of channel feature data, the third group of channel feature data, and the fourth group of channel feature data; outputting the first group of channel feature data as the first-scale feature data; performing convolutional processing on the second group of channel feature data through the first two-dimensional convolutional layer to output the second-scale feature data; merging the third group of channel feature data with the second-scale feature data and then performing convolutional processing through the second two-dimensional convolutional layer to output the third-scale feature data; merging the fourth group of channel feature data with the third-scale feature data and then performing convolutional processing through the third two-dimensional convolutional layer to output the fourth-scale feature data; fusing the first-scale feature data, the second-scale feature data, the third-scale feature data, and the fourth-scale feature data to obtain multi-channel multi-scale feature data.
[0073] Preferably, the first two-dimensional convolutional layer, the second two-dimensional convolutional layer, and the third two-dimensional convolutional layer adopt 3×3 convolutional layers.
[0074] As shown Figure 4 in the figure, the frequency self-attention network includes: a second global average pooling layer, a first one-dimensional convolutional layer, and a fourth activation function. In one embodiment, the fourth activation function adopts a sigmoid activation function.
[0075] It should be understood that the self-attention mechanism allows the model to calculate attention weights based on the relationships between input elements, and then perform weighted summation on the input elements according to these weights to achieve information redistribution and focusing. The frequency self-attention network applies the self-attention mechanism to the frequency dimension of local time periods and can be used to process audio data with frequency information, enabling the speaker classification model to effectively model features in different frequency ranges, adaptively focus on and capture relatively important frequency components in local time-frequency features, thereby enhancing the focusing ability of the speaker classification model on key audio information.
[0076] Specifically, as Figure 4 shown, given an input multi-channel feature X with a shape of C×M×N, where C is the number of feature channels, M is the frequency dimension, and N is the time dimension; the frequency self-attention network performs average pooling operation on the input feature X in the feature channel C dimension through the second global average pooling layer, and outputs a single-channel feature F of 1×M×N. The calculation formula is as shown in formula (10); self-attention calculation is performed on the single-channel feature F through the first one-dimensional convolutional layer and the fourth activation function. The calculation formula is as shown in formula (11), and the weight distribution Q of the multi-channel feature X in the frequency dimension is output, with a size of 1×M×N; the original input multi-channel feature X is multiplied by the obtained frequency weight distribution. The calculation formula is as shown in formula (12), and the multi-channel feature Y with frequency not re-calibrated is obtained for output. The size of the multi-channel feature Y is C×M×N.
[0077]
[0078] Q = σ(W2δ(W1F + b1)+b2); (11)
[0079] Y = Q·X; (12)
[0080] where C is the number of feature channels, M is the frequency dimension, and N is the time dimension; x ijk is the input multi-channel feature X; f ijk is the single-channel feature F; W1, W2, b1, b2 are learnable parameters, such as the weights and biases of the first one-dimensional convolutional layer; σ is the sigmoid activation function, and δ is the ReLU activation function.
[0081] It should be noted that in a preferred embodiment, the first one-dimensional convolutional layer uses a 1×1 convolutional kernel; and the first activation function, the second activation function, and the third activation function of the local time-frequency feature extraction unit can use the ReLU activation function.
[0082] As Figure 2As shown, the embedded feature vector generation unit includes: a time-delay neural network, a first global average pooling layer, a fourth batch normalization layer, a first linear layer, a fifth batch normalization layer, and a second linear layer connected in sequence.
[0083] In one embodiment, as Figure 5 shown, the time-delay neural network includes: a separable one-dimensional convolutional layer, a sixth batch normalization layer, a fifth activation function, a first one-dimensional residual block, a second one-dimensional residual block, a third one-dimensional residual block, a second one-dimensional convolutional layer, a third global average pooling layer, and a seventh batch normalization layer connected in sequence; and the first one-dimensional residual block and the second one-dimensional residual block are also connected to the second one-dimensional convolutional layer. Preferably, the fifth activation function uses the ReLU activation function.
[0084] Specifically, the time-delay neural network is used to capture the temporal dependence relationship in the input multi-channel sequence data by adopting a time-delay mechanism, and extract the global time-frequency features in the multi-channel sequence data through the first one-dimensional residual block, the second one-dimensional residual block, and the third one-dimensional residual block, so that the second one-dimensional convolutional layer connected to each one-dimensional residual block can linearly weight the features in different frequency ranges of the global time-frequency features according to the weight distribution of the input data in the frequency dimension, and generate corresponding multi-channel global time-frequency feature data.
[0085] In this embodiment, by introducing a time-delay mechanism and a weight sharing strategy, the time-delay neural network can efficiently capture and model the temporal dependence relationship in the input sequence data, not only improving the accuracy of the speaker classification model when processing dynamic data, but also optimizing the computational efficiency. On the one hand, each one-dimensional residual block in the time-delay neural network divides the feature channels into different groups, allowing the speaker classification model to model the temporal relationship of the sequence data at different scales, further enhancing the modeling ability of the speaker classification model at multiple scales, and through the cooperation of the second one-dimensional convolutional layer and each one-dimensional residual block, not only expanding the receptive field of the features, enabling the speaker classification model to simultaneously focus on local and global temporal information, but also being able to extract important features from a richer frequency range, with strong feature extraction ability; on the other hand, the second one-dimensional convolution hard-codes the absolute frequency position of the features through the order of the filter coefficients, that is, linearly weights its frequency features according to the weight distribution of the data in the frequency dimension, so as to be able to capture the unique frequency pattern in the speaker's voice, which helps to improve the accuracy of the voiceprint feature recognition in the voiceprint recognition task.
[0086] In one embodiment, the first one-dimensional residual block, the second one-dimensional residual block, and the third one-dimensional residual block include an input layer, one or more one-dimensional convolutional layers, a batch normalization layer, and an activation function connected in sequence, and a shortcut connection is established. After adjusting the input to the same dimension as the output of the activation function using a 1×1 convolution or pooling operation, the results are added and output. This structure helps to solve the problem of gradient disappearance in deep learning networks, enabling the network to be trained deeper. Preferably, the first one-dimensional residual block uses two 3×3 convolutional layers; the second one-dimensional residual block uses three 3×3 convolutional layers; the third one-dimensional residual block uses four 3×3 convolutional layers. Also, the second one-dimensional convolutional layer uses a 1×1 convolutional layer.
[0087] In one embodiment, the first global average pooling layer is used to perform an aggregation operation on all frame-level features in the input feature data and compress the information in the time dimension into a sentence-level feature vector of a fixed length, thereby providing a global representation of the entire speech data to be recognized. It should be understood that the frame-level features are extracted from each frame of the speech data to be recognized and contain the local voiceprint features of each time step. Therefore, in the voiceprint recognition task, the speaker classification model also needs to integrate these local voiceprint features globally to be able to more comprehensively represent the voiceprint features of the entire speech data to be recognized.
[0088] The first linear layer and the second linear layer map the sentence-level feature vector output by the first global average pooling layer into an embedding space for representing the speaker's voiceprint information, generating an embedded feature vector, thereby integrating the feature information of the sentence-level feature vector and reducing its feature dimension. In a specific embodiment, the dimension of the generated embedded feature vector is D = 192, that is, the embedded feature vector V = (v1, v2, …, v 192 )。
[0089] In this embodiment, through the linear transformation of two linear layers, the speaker classification model can compress rich time-frequency features into a compact and highly informative embedded feature vector. This embedded feature vector contains sufficient speaker identity information, which helps improve the accuracy and efficiency of voiceprint feature recognition in subsequent voiceprint recognition tasks. On the one hand, this embedded feature vector can capture the unique voiceprint features of the speaker and effectively distinguish between different speakers; on the other hand, this embedded feature vector can ensure that the speaker's voiceprint features are fully expressed in the embedding space while reducing data redundancy, making the voiceprint features more compact and efficient, suitable for fast and accurate voiceprint recognition.
[0090] In summary, the embedding feature extraction module of the speaker classification model in this application introduces a frequency self-attention mechanism, enabling the speaker classification model to autonomously and dynamically adjust the attention to different frequency features, especially to adaptively identify key frequency information according to the changes in features at different time periods; thereby enhancing the speaker classification model's perception ability of frequency features, enabling it to more accurately capture and process the importance differences of different frequency components in the time dimension, and significantly improving the performance and accuracy of the speaker classification model in feature learning. This flexibility enables the speaker classification model to perform more robustly and efficiently when facing complex time-frequency features.
[0091] Moreover, the embedding feature extraction module of the speaker classification model in this application adopts a hybrid convolution strategy that combines one-dimensional convolution and two-dimensional convolution, enabling the speaker classification model to simultaneously possess the dual capabilities of capturing frequency details and modeling temporal dependencies. One-dimensional convolution focuses on processing the temporal information in the time series, helping the model capture the dynamic changes in the sequence data; while two-dimensional convolution focuses on the joint feature extraction of frequency and time, and can process frequency details more precisely. This hybrid convolution strategy not only improves the speaker classification model's understanding of complex time-frequency features but also enhances its performance in processing temporal data, enabling the speaker classification model to efficiently capture and model the global and local dependencies of time-frequency features at multiple scales, thereby further improving the overall performance.
[0092] It should be noted that the specific structure of the speaker classification model in this application is not limited by this application. Users can appropriately adjust the network structure according to their needs and select different operation layers. For example, in the frequency attention network, the second global average pooling layer does not have to use global average pooling, but can use max pooling or weighted pooling operations of convolution; residual connections, dense connections, or improved residual connection and other deep convolution model structures can also be introduced into each convolution layer.
[0093] In one embodiment, the training method of the speaker classification model includes the following steps.
[0094] ① Randomly crop multiple obtained audio files into a specified length, and pad the audio files with insufficient length to the specified length to generate multiple audio segment data, and add random noise to each audio segment data to generate multiple speech training samples;
[0095] ② Mark the true speaker of each speech training sample, and separately extract the logarithmic Mel spectrogram features of each speech training sample to generate the Mel filter bank feature training data of each speech training sample.
[0096] ③ Input the training data of each Mel filter bank feature into the speaker classification model. Through the embedding feature extraction module, perform embedding feature extraction on the training data of each Mel filter bank feature respectively to generate the embedding feature vectors of each speech training sample, and through the classification module, predict the speakers of each speech training sample.
[0097] In one embodiment, the classification module includes a single-layer linear layer, which can perform class prediction operations according to the extracted embedding feature vectors for calculating the model loss and guiding model optimization. Through optimization in the training stage, the classification module can adapt to different voiceprint features and environmental conditions, improving the accuracy and robustness of the embedding feature extraction module.
[0098] ④ Use the classification loss function to calculate the difference between the predicted speakers and the corresponding true speakers, calculate the gradient of the classification loss function with respect to the speaker classification model, and use the optimizer to update the corresponding model parameters until a converged speaker classification model is obtained.
[0099] In one embodiment, the classification loss function uses the AAM-Softmax loss function, and the calculation formula is shown in formula (13).
[0100]
[0101] Among them, L AAM-Softmax is the loss value of the AAM-Softmax loss function; θ y is the cosine value of the angle between the embedding feature vector and the weight vector of the y-th class, j is the other vector classes except the true weight vector of the y-th class; m is a hyperparameter, or called the angular margin; s is a scaling factor used to control the numerical stability and gradient change.
[0102] It should be noted that the widely used standard Softmax loss function has some defects in some complex classification tasks: Softmax only relies on the dot product of the feature vector and the classification weight for classification. Even if the angles between the feature vector and the weight vectors of multiple categories are close, it may still produce a high category similarity, resulting in weak category discrimination; the classification basis of Softmax is the magnitude of the dot product of the feature vector and the weight vector, mainly depending on the magnitude and direction of the feature vector, but there is no explicit constraint on the difference in direction, which will lead to the fact that the direction information of the feature vector is not fully utilized, affecting the discriminative ability of the model; Softmax does not normalize the feature vector, so both the magnitude and direction of the feature affect the classification decision. Features with larger magnitudes may affect the classification result, and even if these magnitude differences have nothing to do with the category itself, it will lead to noise and bias in feature discrimination; in Softmax, the model pays more attention to the linear relationship between the feature vector and the weight vector, and does not fully utilize the angle information to enhance the discriminative ability of the features. For complex tasks such as speaker recognition, the direction (angle) of the feature is often more important than the magnitude; when dealing with samples with fuzzy or indistinguishable boundaries, Softmax may not provide sufficient discrimination, resulting in poor performance of the model on these complex samples.
[0103] Compared with the Softmax loss function, the AAM-Softmax loss function enhances the discriminative ability of the model in the feature space by introducing an angular margin and normalization processing, especially performing well in dealing with highly similar categories, difficult samples, and tasks that require strong generalization ability.
[0104] In one embodiment, during the parameter optimization process of the speaker classification model, the Adam optimizer with an adaptive learning rate optimization algorithm is used. It can use the first-order moment estimation (mean) and second-order moment estimation (variance) of the gradient to dynamically adjust the learning rate of each parameter, thereby accelerating the convergence of the speaker classification model. Preferably, in order to further improve the training effect of the speaker classification model, the adjustment strategy of the learning rate adopts a method of periodic change. Through this periodic learning rate adjustment strategy, the speaker classification model can specifically increase or decay the learning rate at different stages of training, thereby promoting rapid convergence in the initial stage and effectively avoiding falling into local optimal solutions in the later stage. Therefore, the speaker classification model has strong generalization ability and stability.
[0105] In one embodiment, the mel filter bank feature data of the speech data to be recognized is input into the converged speaker classification model, and the method for extracting the embedded feature vector by performing embedded feature extraction on the mel filter bank feature data includes the following steps.
[0106] ①By performing two-dimensional convolution processing on the Mel filter bank feature data, local time-frequency feature information in the Mel filter bank feature data is extracted to generate corresponding multi-channel time-frequency feature data.
[0107] In one embodiment, two-dimensional convolution processing is performed on the Mel filter bank feature data through a first separable two-dimensional convolution layer.
[0108] ②The multi-channel time-frequency feature data is divided into multiple groups according to the feature channels, and per-channel convolution processing is respectively performed on each group of channel feature data to extract the time-frequency features of the multi-channel time-frequency feature data at different scales, generating corresponding multi-channel multi-scale time-frequency feature data.
[0109] In one embodiment, the multi-channel time-frequency feature data is divided into four groups according to the feature channels through a two-dimensional residual network, and per-channel convolution processing is respectively performed on each group of channel feature data. The specific method includes: dividing the multi-channel time-frequency feature data into four groups according to the feature channels to obtain the first group of channel feature data, the second group of channel feature data, the third group of channel feature data, and the fourth group of channel feature data; outputting the first group of channel feature data as the first-scale feature data; performing convolution processing on the second group of channel feature data through a first two-dimensional convolution layer to output the second-scale feature data; after merging the third group of channel feature data with the second-scale feature data, performing convolution processing through a second two-dimensional convolution layer to output the third-scale feature data; after merging the fourth group of channel feature data with the third-scale feature data, performing convolution processing through a third two-dimensional convolution layer to output the fourth-scale feature data; performing data fusion on the first-scale feature data, the second-scale feature data, the third-scale feature data, and the fourth-scale feature data to obtain multi-channel multi-scale time-frequency feature data.
[0110] ③By adopting a self-attention mechanism in the frequency dimension of the multi-channel multi-scale time-frequency feature data, the weight distribution of each time-frequency segment of the multi-channel multi-scale time-frequency feature data is adaptively calibrated to generate corresponding multi-channel multi-scale enhanced time-frequency feature data.
[0111] In one embodiment, the weight of each time-frequency segment of the multi-channel multi-scale time-frequency feature data is adaptively calibrated through a frequency self-attention network using the attention mechanism to generate corresponding multi-channel multi-scale enhanced time-frequency feature data. The specific method includes: performing average pooling operation on the multi-channel multi-scale time-frequency feature data in the feature channel dimension through a second global average pooling layer to output corresponding single-channel multi-scale time-frequency feature data; performing self-attention calculation on the single-channel multi-scale time-frequency feature data through a first one-dimensional convolution layer and a fourth activation function to output the weight distribution of each time-frequency segment of the multi-channel multi-scale time-frequency feature data; multiplying the multi-channel multi-scale time-frequency feature data by the weight distribution of each time-frequency segment to output corresponding multi-channel multi-scale enhanced time-frequency feature data.
[0112] ④ By adopting a time delay mechanism to capture the temporal dependence relationship in the multi-channel multi-scale enhanced time-frequency feature data, and based on the weight distribution of the multi-channel multi-scale time-frequency feature data for each time-frequency segment, linearly weight the features in different frequency ranges to generate corresponding multi-channel global time-frequency feature data.
[0113] In one embodiment, a time delay neural network is used to adopt a time delay mechanism to capture the temporal dependence relationship in the multi-channel multi-scale enhanced time-frequency feature data, and the global time-frequency features in the multi-channel multi-scale enhanced time-frequency feature data are extracted through the first one-dimensional residual block, the second one-dimensional residual block, and the third one-dimensional residual block; through the second one-dimensional convolutional layer, according to the weight distribution of the multi-channel multi-scale time-frequency feature data for each time-frequency segment, linearly weight the features in different frequency ranges to generate corresponding multi-channel global time-frequency feature data.
[0114] ⑤ Perform an aggregation operation on all frame-level features in the multi-channel global time-frequency feature data, and compress the information in the time dimension into a sentence-level feature vector with a fixed length.
[0115] In one embodiment, a first global average pooling layer is used to perform an aggregation operation on all frame-level features in the multi-channel global time-frequency feature data, and compress the information in the time dimension into a sentence-level feature vector with a fixed length.
[0116] ⑥ Perform a linear transformation on the sentence-level feature vector, and map the sentence-level feature vector into an embedding space for representing the speaker's voiceprint information to generate a corresponding embedding feature vector.
[0117] In one embodiment, through the first linear layer and the second linear layer, perform a linear transformation on the sentence-level feature vector, and map the sentence-level feature vector into an embedding space for representing the speaker's voiceprint information, thereby generating a corresponding embedding feature vector.
[0118] Step S3: Calculate the similarity between the embedding feature vector and the template embedding feature vectors of one or more registered users respectively, and screen the target embedding feature vectors that match the embedding feature vector to verify that the current speaker is a registered user and identify the target registered user.
[0119] In one embodiment, step S3 specifically includes: successively calculating the cosine similarity between the embedded feature vector and the template embedded feature vectors of each registered user, and obtaining the template embedded feature vector with the largest cosine similarity value as the primary target embedded feature vector; determining whether the cosine similarity value between the embedded feature vector and the primary target embedded feature vector is greater than or equal to a preset cosine similarity threshold; if the cosine similarity value is greater than or equal to the cosine similarity threshold, determining the primary target embedded feature vector as the target embedded feature vector, and determining that the current speaker and the registered user corresponding to the target embedded feature vector are the same speaker, and identifying the registered user as the target registered user; if the cosine similarity value is less than the cosine similarity threshold, determining that the current speaker is an unregistered user.
[0120] It should be noted that the cosine similarity calculates the similarity degree of the corresponding voice data in terms of speaker features by measuring the angular difference between two vectors in a high-dimensional space. The higher the similarity, the more likely they are from the same speaker. However, the similarity calculation of two embedded feature vectors is not limited to cosine similarity. Other strategies such as Euclidean distance, Manhattan distance, and dynamic time warping can also be used, or the compared embedded feature vectors can be equally divided into multiple parts, and the similarity of each part can be calculated and the average score can be taken. Specifically, the present application does not limit this. Moreover, the present application does not limit the size of the preset cosine similarity threshold, and users can set and adjust it according to their needs.
[0121] The template embedded feature vectors of each registered user are stored in a pre-constructed voice database. In one embodiment, the construction method of the voice data includes: extracting the logarithmic Mel spectrogram features of the voice data of the registered user to generate the Mel filter bank feature data of the registered user; inputting the Mel filter bank feature data into a speaker classification model, and generating the template embedded feature vector of the registered user through an embedded feature extraction module; storing the registered user and the corresponding template embedded feature vector in the voice database.
[0122] As Figure 6 shown, a schematic structural diagram of a voiceprint recognition system 600 in an embodiment of the present invention is shown. The voiceprint recognition system 600 is used to implement the voiceprint recognition method provided in the above embodiments, and can solve the technical problems of excessive data calculation amount, high requirements for device capabilities, and low recognition accuracy existing in the existing voiceprint recognition technology. The voiceprint recognition system 600 includes: a first feature extraction module 601, a second feature extraction module 602, and a voiceprint recognition module 603.
[0123] Specifically, the first feature extraction module 601 is configured to extract the logarithmic Mel spectrogram features of the to-be-recognized voice data of the current speaker to generate the Mel filter bank feature data of the to-be-recognized voice data.
[0124] The second feature extraction module 602, connected to the first feature extraction module 601, is used to extract the embedding features of the Mel filter bank feature data based on the embedding feature extraction module of the pre-trained speaker classification model, and generate the embedding feature vector of the speech data to be recognized.
[0125] Among them, the specific method includes: performing two-dimensional convolution processing on the Mel filter bank feature data to extract the local time-frequency feature information in the Mel filter bank feature data and generate the corresponding multi-channel time-frequency feature data; dividing the multi-channel time-frequency feature data into multiple groups according to the feature channels, and performing per-channel convolution processing on each group of channel feature data respectively to extract the time-frequency features of the multi-channel time-frequency feature data at different scales and generate the corresponding multi-channel multi-scale time-frequency feature data; adopting a self-attention mechanism in the frequency dimension of the multi-channel multi-scale time-frequency feature data to adaptively calibrate the weight distribution of each time-frequency segment of the multi-channel multi-scale time-frequency feature data and generate the corresponding multi-channel multi-scale enhanced time-frequency feature data; adopting a time delay mechanism to capture the temporal dependence relationship in the multi-channel multi-scale enhanced time-frequency feature data, and linearly weighting the features in different frequency ranges based on the weight distribution of each time-frequency segment of the multi-channel multi-scale time-frequency feature data to generate the corresponding multi-channel global time-frequency feature data; performing an aggregation operation on all frame-level features in the multi-channel global time-frequency feature data to compress the information in the time dimension into a fixed-length sentence-level feature vector; performing a linear transformation on the sentence-level feature vector and mapping the sentence-level feature vector into an embedding space for representing the speaker's voiceprint information to generate the corresponding embedding feature vector.
[0126] The voiceprint recognition module 603, connected to the second feature extraction module 602, is used to calculate the similarity between the embedding feature vector and the template embedding feature vectors of one or more registered users respectively, and screen the target embedding feature vector that matches the embedding feature vector to verify that the current speaker is a registered user and identify the target registered user.
[0127] It should be understood that the specific processes of each module executing the above corresponding steps have been described in detail in the above method embodiments. For the sake of brevity, they will not be repeated here.
[0128] It should also be understood that the division of the modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, there may be other division methods. In addition, each functional module in the various embodiments of the present application may be integrated in one processor, or may exist separately physically, or two or more modules may be integrated in one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0129] The voiceprint recognition method provided by the embodiments of the present application can be implemented on the terminal side or the server side. In terms of the hardware structure of the voiceprint recognition terminal, please refer to Figure 7 , which is an optional schematic diagram of the hardware structure of the voiceprint recognition terminal 700 provided by the embodiments of the present invention. The voiceprint recognition terminal 700 can be a mobile phone, a computer device, a tablet device, a personal digital processing device, a factory background processing device, etc. The voiceprint recognition terminal 700 includes: at least one processor 701, a memory 702, at least one network interface 704, and a user interface 706. Moreover, each component in the voiceprint recognition terminal 700 is coupled together through a bus system 705. It can be understood that the bus system 705 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 705 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 7 all kinds of buses are labeled as the bus system. The user interface 706 can include a display, a keyboard, a mouse, a trackball, a click gun, a key, a button, a touchpad, or a touch screen, etc.
[0130] It can be understood that the memory 702 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The present application does not specifically limit this. The memory 702 in the embodiments of the present invention is used to store various types of data to support the operation of the voiceprint recognition terminal 700. Examples of these data include: any executable program for operating on the voiceprint recognition terminal 700, such as an operating system 7021 and an application program 7022; the operating system 7021 contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program 7022 can include various application programs, such as a MediaPlayer, a Browser, etc. Implementing the voiceprint recognition method provided by the embodiments of the present invention can be included in the application program 7022.
[0131] The voiceprint recognition method disclosed in the above embodiments of the present invention can be applied to the processor 701 or implemented by the processor 701. The processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 701 or by instructions in the form of software. The above-mentioned processor 701 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 701 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor 701 can be a microprocessor or any conventional processor, etc.
[0132] In an exemplary embodiment, the voiceprint recognition terminal 700 may be one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs) for performing the foregoing voiceprint recognition method.
[0133] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to a computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0134] As described above, the foregoing is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0135] In summary, the present application provides a voiceprint recognition method, system, terminal, and medium. By extracting log Mel spectral features from the voice data to be recognized of the current speaker and further extracting embedding features based on the speaker classification model, an embedding feature vector is generated for calculating the similarity with the template embedding feature vectors of one or more registered users to verify whether the current speaker is a registered user. Therefore, the present application has the following beneficial effects: by introducing a frequency self-attention mechanism, the speaker classification model can adaptively adjust the attention to different frequency features, accurately identify key frequency information, and significantly improve the expressiveness and accuracy of the model in feature learning; by introducing a multi-dimensional hybrid convolution strategy, the model simultaneously has the dual capabilities of capturing frequency details and modeling temporal dependencies; by means of a per-channel separable convolution strategy, the computational complexity and the number of parameters of the model are reduced, and the computational cost is reduced; thus, the present application solves the technical problems of excessive data calculation volume, high requirements for device capabilities, and low recognition accuracy existing in the existing voiceprint recognition technology.
[0136] Therefore, the present application effectively overcomes various disadvantages in the prior art and has high industrial utilization value.
[0137] The above embodiments are only illustrative of the principles and effects of the present application and are not used to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes completed by those of ordinary skill in the art in the technical field without departing from the spirit and technical ideas disclosed in the present application should still be covered by the claims of the present application.
Claims
1. A voiceprint recognition method, characterized in that: include: Performing logarithmic Mel spectrum feature extraction on the acquired speech data to be recognized of the current speaker to generate Mel filter bank feature data of the speech data to be recognized; An embedded feature extraction module based on a pre-trained speaker classification model performs embedded feature extraction on the Mel filter bank feature data to generate an embedded feature vector of the speech data to be recognized; Calculate similarity between the embedded feature vector and template embedded feature vectors of one or more registered users, select a target embedded feature vector that matches the embedded feature vector, so as to verify that the current speaker is a registered user, and identify the target registered user; Among them, based on the embedded feature extraction module, the method of performing embedded feature extraction on the Mel filter group feature data to generate the embedded feature vector of the speech data to be recognized includes: performing two-dimensional convolution processing on the Mel filter group feature data to extract local time-frequency feature information in the Mel filter group feature data to generate corresponding multi-channel time-frequency feature data; dividing the multi-channel time-frequency feature data into multiple groups according to feature channels, and performing channel-by-channel convolution processing on each group of channel feature data to extract the time-frequency features of the multi-channel time-frequency feature data at different scales to generate corresponding multi-channel multi-scale time-frequency feature data; using a self-attention mechanism in the frequency dimension of the multi-channel multi-scale time-frequency feature data to adaptively calibrate the multi-channel multi-scale The weight distribution of the time-frequency feature data in each time-frequency segment is used to generate corresponding multi-channel multi-scale enhanced time-frequency feature data; the temporal dependency in the multi-channel multi-scale enhanced time-frequency feature data is captured by adopting a time delay mechanism, and based on the weight distribution of the multi-channel multi-scale time-frequency feature data in each time-frequency segment, the features in different frequency ranges are linearly weighted to generate corresponding multi-channel global time-frequency feature data; all frame-level features in the multi-channel global time-frequency feature data are aggregated to compress the information in the time dimension into a sentence-level feature vector of fixed length; the sentence-level feature vector is linearly transformed, and the sentence-level feature vector is mapped to an embedding space for representing the speaker's voiceprint information to generate a corresponding embedded feature vector.
2. The voiceprint recognition method according to claim 1, characterized in that: The embedding feature extraction module comprises: A local time-frequency feature extraction unit, comprising: a first separable two-dimensional convolutional layer, a first normalization layer, a first activation function, a two-dimensional residual network, a second normalization layer, a second activation function, a second separable two-dimensional convolutional layer, a third normalization layer, a third activation function, and a frequency self-attention network connected in sequence; The embedding feature vector generation unit includes: a time-delay neural network, a first global average pooling layer, a fourth batch normalization layer, a first linear layer, a fifth batch normalization layer and a second linear layer connected in sequence.
3. The voiceprint recognition method according to claim 2, characterized in that: The two-dimensional residual network includes: a first two-dimensional convolutional layer, a second two-dimensional convolutional layer and a third two-dimensional convolutional layer; Among them, the two-dimensional residual network is used to divide the multi-channel time-frequency feature data into four groups according to the feature channels, and perform channel-by-channel convolution processing on each group of channel feature data; the specific method includes: dividing the multi-channel time-frequency feature data into four groups according to the feature channels, obtaining a first group of channel feature data, a second group of channel feature data, a third group of channel feature data and a fourth group of channel feature data; outputting the first group of channel feature data as first scale feature data; performing convolution processing on the second group of channel feature data through the first two-dimensional convolution layer, and outputting second scale feature data; merging the third group of channel feature data with the second scale feature data, performing convolution processing through the second two-dimensional convolution layer, and outputting third scale feature data; merging the fourth group of channel feature data with the third scale feature data, performing convolution processing through the third two-dimensional convolution layer, and outputting fourth scale feature data; performing data fusion on the first scale feature data, the second scale feature data, the third scale feature data and the fourth scale feature data to obtain multi-channel multi-scale time-frequency feature data.
4. The voiceprint recognition method according to claim 2, characterized in that: The frequency self-attention network includes: a second global average pooling layer, a first one-dimensional convolutional layer, and a fourth activation function connected in sequence; Among them, the frequency self-attention network is used to adopt the attention mechanism to adaptively calibrate the weights of the multi-channel multi-scale time-frequency feature data in each time and frequency segment, and generate corresponding multi-channel multi-scale enhanced time-frequency feature data; the specific method includes: performing an average pooling operation on the multi-channel multi-scale time-frequency feature data in the feature channel dimension through the second global average pooling layer, and outputting the corresponding single-channel multi-scale time-frequency feature data; performing self-attention calculation on the single-channel multi-scale time-frequency feature data through the first one-dimensional convolutional layer and the fourth activation function, and outputting the weight distribution of the multi-channel multi-scale time-frequency feature data in each time and frequency segment; multiplying the multi-channel multi-scale time-frequency feature data by the weight distribution in each time and frequency segment, and outputting the corresponding multi-channel multi-scale enhanced time-frequency feature data.
5. The voiceprint recognition method according to claim 2, characterized in that: The time-delay neural network includes: a separable one-dimensional convolutional layer, a sixth batch normalization layer, a fifth activation function, a first one-dimensional residual block, a second one-dimensional residual block, a third one-dimensional residual block, a second one-dimensional convolutional layer, a third global average pooling layer and a seventh batch normalization layer connected in sequence; the first one-dimensional residual block and the second one-dimensional residual block are further connected to the second one-dimensional convolutional layer; Among them, the time-delay neural network captures the temporal dependencies in the multi-channel multi-scale enhanced time-frequency feature data by adopting a time delay mechanism, and extracts the global time-frequency features in the multi-channel multi-scale enhanced time-frequency feature data through the first one-dimensional residual block, the second one-dimensional residual block and the third one-dimensional residual block; through the second one-dimensional convolutional layer, according to the weight distribution of the multi-channel multi-scale time-frequency feature data in each time and frequency segment, the features of different frequency ranges are linearly weighted to generate corresponding multi-channel global time-frequency feature data.
6. The voiceprint recognition method according to claim 1, characterized in that: Methods for training the speaker classification model include: The obtained multiple audio files are randomly cut into specified lengths, and the audio files with insufficient lengths are padded to the specified length to generate multiple audio segment data, and random noise is added to each audio segment data to generate multiple speech training samples; Mark the real speaker of each speech training sample, and perform logarithmic Mel spectrum feature extraction on each speech training sample to generate Mel filter bank feature training data for each speech training sample; Inputting each Mel filter group feature training data into the speaker classification model, performing embedding feature extraction on each Mel filter group feature training data through the embedding feature extraction module of the speaker classification model, generating an embedding feature vector of each speech training sample, and predicting the speaker of each speech training sample through the classification module of the speaker classification model; The difference between each predicted speaker and the corresponding real speaker is calculated using a classification loss function, and the gradient of the classification loss function with respect to the speaker classification model is calculated, and the corresponding model parameters are updated using an optimizer until a converged speaker classification model is obtained.
7. The voiceprint recognition method according to claim 1, characterized in that: The method of calculating similarity between the embedded feature vector and the template embedded feature vectors of one or more registered users, selecting a target embedded feature vector matching the embedded feature vector to verify that the current speaker is a registered user, and identifying the target registered user includes: The embedding feature vector is respectively subjected to cosine similarity calculation with the template embedding feature vector of each registered user, and the template embedding feature vector with the largest cosine similarity value is obtained as the primary target embedding feature vector; wherein the template embedding feature vector of each registered user is stored in a pre-constructed voice database; Determining whether a cosine similarity value between the embedded feature vector and the primary target embedded feature vector is greater than a preset cosine similarity threshold; If the cosine similarity value is greater than or equal to the cosine similarity threshold, determining the primary target embedding feature vector as the target embedding feature vector, determining that the current speaker and the registered user corresponding to the target embedding feature vector are the same speaker, and identifying the registered user as the target registered user; If the cosine similarity value is less than the cosine similarity threshold, it is determined that the current speaker is an unregistered user.
8. A voiceprint recognition system, characterized in that: include: A first feature extraction module is used to perform logarithmic Mel spectrum feature extraction on the acquired speech data to be recognized of the current speaker, and generate Mel filter bank feature data of the speech data to be recognized; A second feature extraction module, connected to the first feature extraction module, is used to perform embedded feature extraction on the Mel filter bank feature data based on an embedded feature extraction module of a pre-trained speaker classification model to generate an embedded feature vector of the speech data to be recognized; a voiceprint recognition module, connected to the second feature extraction module, configured to perform similarity calculations on the embedded feature vector and template embedded feature vectors of one or more registered users, select a target embedded feature vector that matches the embedded feature vector, so as to verify that the current speaker is a registered user, and identify the target registered user; Among them, based on the embedded feature extraction module, the method of performing embedded feature extraction on the Mel filter group feature data to generate the embedded feature vector of the speech data to be recognized includes: performing two-dimensional convolution processing on the Mel filter group feature data to extract local time-frequency feature information in the Mel filter group feature data to generate corresponding multi-channel time-frequency feature data; dividing the multi-channel time-frequency feature data into multiple groups according to feature channels, and performing channel-by-channel convolution processing on each group of channel feature data to extract the time-frequency features of the multi-channel time-frequency feature data at different scales to generate corresponding multi-channel multi-scale time-frequency feature data; using a self-attention mechanism in the frequency dimension of the multi-channel multi-scale time-frequency feature data to adaptively calibrate the multi-channel multi-scale The weight distribution of the time-frequency feature data in each time-frequency segment is used to generate corresponding multi-channel multi-scale enhanced time-frequency feature data; the temporal dependency in the multi-channel multi-scale enhanced time-frequency feature data is captured by adopting a time delay mechanism, and based on the weight distribution of the multi-channel multi-scale time-frequency feature data in each time-frequency segment, the features in different frequency ranges are linearly weighted to generate corresponding multi-channel global time-frequency feature data; all frame-level features in the multi-channel global time-frequency feature data are aggregated to compress the information in the time dimension into a sentence-level feature vector of fixed length; the sentence-level feature vector is linearly transformed, and the sentence-level feature vector is mapped to an embedding space for representing the speaker's voiceprint information to generate a corresponding embedded feature vector.
9. A voiceprint recognition terminal, characterized in that: include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory, so that the terminal executes the voiceprint recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voiceprint recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Voiceprint recognition method and system based on wavelet transform convolution, terminal and medium
CN120600032A
Wavelet transform convolution-based voiceprint recognition method and system, terminal and medium
CN120600032B
Voiceprint registration method and device based on deep voice embedding
CN121662052A