A Speaker Recognition Method and System for Real Scenarios
By using multi-resolution convolution kernel and grouping convolution technology in short-term speech recognition, multi-scale features are extracted in the time and frequency domains and masked, the problems of noise interference and insufficient information are solved, and the accuracy and stability of speaker recognition are improved.
Patent Information
- Application Number
- CN202411847687.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-16
AI Technical Summary
There are problems of noise interference and insufficient information in short-term speech recognition, resulting in a decrease in recognition accuracy.
Using a speaker recognition method for real scenes, features are extracted in the time domain through a multi-resolution convolution kernel, and feature extraction and fusion are extracted and fused in the frequency domain, and finally mask processing is performed to output the recognition results.
It enhances the extraction of the target speaker's characteristics, reduces the impact of noise on the recognition results, improves the stability and reliability of the recognition, and can effectively capture the subtle differences in the speaker.
Smart Images

Figure CN119314493B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speaker recognition, and specifically to a speaker recognition method for real scenarios. Background Art
[0002] The human auditory system is the most convenient and efficient recognition system except for vision. Today, with the rapid development of Internet and Internet of Things technologies, the application scenarios of personal identity authentication technologies are becoming more and more diversified. Currently, in the field of identity authentication, mainly single or multi-modal fusion personal identity authentication technologies such as fingerprint and face recognition are adopted. However, identity authentication technologies such as face recognition and fingerprint recognition not only rely on expensive professional equipment, but also can only achieve high accuracy in relatively specific environments. And voice, as a communication method that has existed since the birth of humans, itself has universality. And the individual differences and growth environment differences of different speakers lead to the differences and uniqueness of the voices of different speakers. Therefore, the voiceprint recognition technology based on voice can realize the verification of personal identity.
[0003] In real scenarios, compared with identity authentication technologies such as face recognition and fingerprint recognition, the speaker recognition technology has the following advantages: (1) High security, voice is a dynamic feature, which is more difficult to forge compared with static identity information, and voice can be combined with dynamic text passwords for application to reduce the risk brought by transcribed speech; (2) Strong privacy, compared with visual information such as pictures and videos, users are more inclined to upload their own voices; (3) Low cost, compared with acquisition devices such as fingerprint collectors and cameras, the cost of microphones is lower. In addition, the speaker recognition technology also helps with the voice recording in conference rooms, improves the accuracy of conference records, and provides users with a better voice interaction experience and a more secure device control experience in devices such as smart cars and smart homes.
[0004] With its wide range of application scenarios and important research value, the speaker recognition technology is gradually entering all aspects of our lives, and has broad development prospects and commercial value. Summary of the Invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] Therefore, the technical problem solved by the present invention is: the problems of noise interference and insufficient information in short-time speech recognition.
[0007] To solve the above technical problems, the present invention provides the following technical solution: A speaker recognition method for real scenarios, including:
[0008] Obtain external sound information, process it through multi-resolution convolutional kernels, extract features in the time domain and perform feature fusion only in the time domain;
[0009] After feature fusion in the time domain, group convolution is used to extract features in the frequency domain, process them, and perform feature fusion only in the frequency domain;
[0010] After frequency-domain feature fusion, masking processing is performed, and finally the recognition result is output;
[0011] The multi-resolution convolution kernel includes: setting parallel branch filters with three different scales, and each branch consists of two layers of one-dimensional convolution;
[0012] The first layer of one-dimensional convolution is used for primary filtering, and the second layer of one-dimensional convolution is used for dimension matching;
[0013] Feature fusion in the time domain includes: extracting features at different time resolutions through different convolution kernel sizes and strides, and splicing the features of three different time resolutions along the channel dimension;
[0014] Find the smallest time step and crop all features so that all feature time steps are the same;
[0015] Feature extraction in the frequency domain includes: dividing the input frequency channel features into 4 groups to obtain four grouped subsets X1, X2, X3, X4;
[0016] The first subset is X1, and the output is Y1;
[0017] The second subset is X2, which is processed by a 7×7 convolution kernel to capture broader context information, and the output is Y2;
[0018] The third subset is X3, Y2 is added to X3, and it is processed using a 5×5 convolution kernel to obtain the output Y3;
[0019] The fourth subset is X4, Y3 is added to X4, and it is processed using a 3×3 convolution kernel to obtain the output Y4;
[0020] By performing feature splicing on Y1, Y2, Y3, and Y4 in the frequency channel dimension, the combined feature vector is input into the classifier, and the final speaker recognition result is output through activation and normalization processing. As a preferred solution of the speaker recognition method for real scenarios described in the present invention, among them: performing masking processing on the time-frequency features includes: obtaining the acoustic features of the external sound information, predicting the enhanced acoustic features, and obtaining the feature map;
[0021] Applying a noise mask on the feature map to blur the irrelevant noise.
[0022] As a preferred solution of the speaker recognition method for real scenarios according to the present invention, the time-frequency features are masked as follows: extracting the utterance-level feature vectors, calculating the mean and standard deviation, and mapping the mean and standard deviation to the auxiliary context features;
[0023] Generating a noise mask using the auxiliary context features and the acoustic features, and calculating the final feature map, which is expressed by the formula:
[0024]
[0025] where e represents the auxiliary context feature vector, W3 represents the weight matrix, represents the transpose of the weight matrix; μ represents the mean vector calculated by statistical pooling, σ represents the standard deviation vector calculated by statistical pooling, b3 represents the bias term, and M *t represents the noise mask vector generated at the t-th frame, σ(·) represents the Sigmoid function, W2 represents the weight matrix, represents the transpose of the weight matrix; δ(·) represents the combination of the ReLU function and BN; W1 represents another weight matrix, represents the transpose of the weight matrix; F *t represents the feature vector at the t-th frame, b2 represents the bias term, represents the feature map after being masked by the noise mask, g(·) represents the transformation function, F(x) represents the feature representation; M represents the generated noise mask, and ⊙ represents multiplication.
[0026] As a preferred solution of the speaker recognition method for real scenarios according to the present invention, the final output of the recognition result includes: inputting the merged feature vectors into a classifier, and the classifier outputs the final speaker recognition result through activation and normalization processing.
[0027] A speaker recognition system for real scenarios, including:
[0028] A noise mask module that processes noise interference, extracts robust speaker feature vectors, and focuses on the target speaker and blurs irrelevant noise by dynamically predicting the noise mask;
[0029] A time-domain multi-resolution encoder that extracts features with different time resolutions in the time domain, sets multiple parallel branch filters, extracts features from the input speech signal through convolutional kernels of different scales, and performs feature fusion;
[0030] A frequency-domain multi-channel encoder that extracts features in the frequency domain by means of grouped convolution, processes short-time speech information, divides the frequency-domain features into multiple groups, and uses different convolutional kernels for progressive processing and fusion;
[0031] A classifier inputs the processed features into the classifier for speaker recognition and outputs the recognition result.
[0032] Advantages of the present invention: The speaker recognition method for real scenarios provided by the present invention enhances the extraction of target speaker features and reduces the influence of noise on the recognition result through a noise masking module and multi-resolution feature extraction. At the same time, in a complex noise environment, the system can effectively focus on the target speaker, improving the stability and reliability of recognition. Through multi-scale feature extraction in the time domain and frequency domain, more comprehensive feature information is provided, which helps to capture the subtle differences of speakers. Description of the Drawings
[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1 It is the overall flowchart of a speaker recognition method for real scenarios provided by the first embodiment of the present invention;
[0035] Figure 2 It is the system structure diagram of a speaker recognition method for real scenarios provided by the first embodiment of the present invention;
[0036] Figure 3 It is the MTFNM-TDNN model structure diagram of a speaker recognition method for real scenarios provided by the first embodiment of the present invention;
[0037] Figure 4 It is the noise masking module of a speaker recognition method for real scenarios provided by the first embodiment of the present invention;
[0038] Figure 5 It is the time-domain multi-resolution encoder of a speaker recognition method for real scenarios provided by the first embodiment of the present invention;
[0039] Figure 6 It is the ResfBlock structure diagram of a speaker recognition method for real scenarios provided by the first embodiment of the present invention;
[0040] Figure 7 It is the frequency-domain multi-channel encoder of a speaker recognition method for real scenarios provided by the first embodiment of the present invention. Detailed Embodiments
[0041] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] Example 1, referring to Figure 1 、 Figure 2 and Figure 3 , an embodiment of the present invention provides a speaker recognition method for a real scenario, including:
[0043] S1: Obtain external sound information, process it through a multi-resolution convolution kernel, and extract features in the time domain.
[0044] Aiming at the problems of various noise types and short speech duration in real scenarios, a noise masking method and a short-time speech information enhancement method based on a deep neural network are proposed, including: aiming at the problem that the existence of various noises in real scenarios affects the accuracy of the speaker recognition system, considering uniformly masking different types of noises, and applying the noise masking module to the deep layer and shallow layer of the network respectively. Applied to the shallow layer of the network, it can process non-speech information such as background noise. Applied to the deep layer of the network, it can process other speaker interference noises such as laughter and discussion sounds. This module solves the noise interference problem by omitting the speech and non-speech noises that interfere with the speaker. The utterance-level feature vector is extracted through statistical pooling to distinguish the target speaker from the global noise, and all frames are combined with the statistical pooling layer to assign different mask values to different channels and different frames.
[0045] Then, aiming at the problem of insufficient short-time speech information, consider enhancing short-time speech information, enhancing short-time speech information in two directions: time domain and frequency domain. A multi-resolution encoder is constructed in the time domain, and a multi-channel encoder is constructed in the frequency domain to extract features with different frequencies and different time resolutions, focusing on global information and local detail information. The time-domain multi-resolution encoder uses convolution kernels with different stride sizes for multi-scale convolution feature extraction. These convolution kernels extract features at different time resolutions respectively, and then the features with three different time resolutions are concatenated along the channel dimension to ensure that all features have the same time step. The multi-channel encoder adopts the idea of grouped convolution. Compared with ordinary single-layer convolution, the output result of grouped convolution will contain combinations of different receptive field sizes, obtaining information at different frequency scales respectively.
[0046] The degradation of recognition performance caused by noise has always been a long-term challenge in speaker recognition. For noise processing, previous methods usually perform denoising transformation on the speaker's feature vectors or enhance the speaker's feature vectors. These methods are lossy and inefficient for the extraction of speaker feature vectors. The present invention proposes a method for extracting robust speaker feature vectors, a noise masking module. The noise masking module enables the speaker recognition model to focus on the target speaker and blur the irrelevant noise.
[0047] The method of applying denoising transformation to the speaker feature vector is to use a statistical backend or a neural network backend to transform the noisy speaker feature vector into an enhanced speaker feature vector. The speaker feature vector is obtained by aggregating hundreds of frame-level features into a discourse-level speaker feature vector through a statistical pooling layer (such as mean pooling). During the aggregation stage, information loss is inevitable. Therefore, post-processing the noisy speaker feature vector after the aggregation step is lossy.
[0048] The method of filtering noise through a speech enhancement model, these methods train the model based on noise-clean speech pairs or perform unsupervised speech enhancement through a generative adversarial network. However, these methods are not specific to the speaker recognition task, and the training data for speech enhancement and speaker feature vectors may belong to different domains, which also greatly increases the training cost.
[0049] In the speaker recognition task, the model input usually adopts filter bank frequency band (Fbank) features and Mel frequency cepstral coefficients (MFCC) features. However, for the speaker recognition task of short-time speech, due to the short speech time length and limited information contained, using Fbank features and MFCC features as input will lead to a decline in model performance.
[0050] Fbank features have some deficiencies in processing short-time speech. The short duration of short-time speech results in less spectral information in the Fbank features, which may make the features insufficient to comprehensively represent the characteristics of the speaker, thereby reducing the recognition accuracy. Fbank features mainly reflect the spectral information of speech, but in short-time speech, due to limited temporal information, the dynamic changes of speaker features may not be fully captured.
[0051] For MFCC features, their extraction process depends on a fixed-length time window. For short-time speech, this fixed time window may not fully adapt to the changes in short-time speech, thus affecting the accuracy of the features. In addition, during the extraction of MFCC features, the spectral information is compressed through the discrete cosine transform (DCT), which will lose a part of the high-frequency information. For short-time speech, the already limited information is further compressed, which may lead to the loss of important speaker features.
[0052] To solve the problem of insufficient speech information caused by short-time speech, the present invention designs a time-domain multi-resolution encoder, which sets parallel branch filters of three different scales. The convolutional kernels of these branch filters have different sizes and strides. Each branch consists of two layers of one-dimensional convolution, as Figure 5 shown. The first layer of one-dimensional convolution is used for primary filtering, and the second one-dimensional convolution performs dimension matching, aiming to compensate for the resulting dimensional differences.
[0053] The detailed settings of the time-domain multi-resolution encoding are shown in Table 1, and TDNN is used as a frame aggregator to perform dimension matching with the subsequent structure. The numbers in parentheses are the convolution parameters: representing the number of channels, size, and stride of the convolutional kernel respectively.
[0054] Table 1 Multi-resolution Encoder Design Table
[0055]
[0056] The time-domain multi-resolution encoder extracts features at different time resolutions through the convolutional kernels of three different branches. The features of three different time resolutions are concatenated along the channel dimension through torch.cat, the smallest time step is found, and all features are cropped to ensure that all features have the same time step.
[0057] Furthermore, in this way, the time-domain multi-resolution encoder can effectively capture the features of different time resolutions in the input speech signal, fuse these features, and increase the information of short-time speech for subsequent processing.
[0058] S2: After extracting features in the time domain, use grouped convolution to extract features in the frequency domain, perform processing and feature fusion.
[0059] Short-time speech may contain limited frequency information, especially in the high-frequency part, where information may be lost or insufficient. This may cause the model to fail to capture the complete spectral features of the speech signal. Secondly, currently commonly used frequency-domain feature extraction methods (such as MFCC or Fbank) may not be able to fully capture the frequency-domain features of the speech signal, especially in the case of complex background noise or drastic speech changes, and their expression ability is limited. In addition, the current frequency-domain features may not be able to fully express the features of the speaker. The frequency-domain information in the speech signal may have different importance in different frequency ranges, and existing methods may not be able to effectively capture these differences.
[0060] To solve the above problems, the present invention designs a frequency-domain multi-channel encoder, which applies the idea of grouped convolution to the frequency dimension. Aiming at the problem that the short duration of short-time speech leads to insufficient speech information, the frequency-domain multi-channel encoder based on the idea of grouped convolution divides the frequency-domain features into multiple groups, and each group can be regarded as features at different scales. This enables the model to capture the frequency-domain information at different scales more comprehensively, thereby improving the diversity and richness of the features.
[0061] In addition, grouped convolution processes the features at different granularities in the frequency domain, avoiding the risk of information loss. Each group can extract features within a specific frequency range through an independent convolutional kernel, thereby retaining more frequency-domain information. Finally, the features within each group can communicate and integrate information through residual connections or other means. This communication mechanism can help effective information transfer and learning between features, enhancing the information expression ability of short-time speech in the frequency domain.
[0062] As Figure 6 and Figure 7 shown, the input feature X is first grouped in the dimension of frequency channels, dividing the original C-dimensional data into 4 groups of C'-dimensional data, thus obtaining four grouped subsets X1, X2, X3, X4, and the frequency channel dimension of each subset is Thereby realizing grouped processing of frequencies. For the first subset X1, it is processed using BN and ReLU activation functions to obtain the output Y1, aiming to improve the stability of the model training process and accelerate convergence. For the second subset X2, it is first processed using a 7×7 convolutional kernel to capture more extensive context information, and then it is also processed through BN and ReLU to output Y2. This step enables the model to pay attention to broader feature information while maintaining the local details of the features. Then, Y2 is added to the third subset X3 as the input of the next convolutional operation. Here, a 5×5 convolutional kernel is used for processing, aiming to balance the width of the receptive field and the specificity of the features to obtain a more fine-grained feature representation. After processing, it is also processed through BN and ReLU to output Y3. Finally, Y3 is added to the fourth subset X4, and a 3×3 convolutional kernel is used for processing. Aiming to further extract finer local features to further enhance the model's ability to capture details. After passing through BN and ReLU, the final output Y4 is obtained.
[0063] In summary, through a series of grouped convolutions and gradually integrating the feature processing at different scales, the finally obtained multi-scale grouped convolution output Y is formed by feature concatenation of Y1, Y2, Y3, and Y4 in the frequency channel dimension.
[0064] In the experimental stage, the Equal Error Rate (EER) and the minimum Detection Cost Function (minDCF) are used to evaluate the overall performance of the speaker recognition model. The equal error rate represents the value when the False Acceptance Rate (FAR) is equal to the False Rejection Rate (FRR). Among them, the calculation formulas of FAR and FRR are shown in Equations (10) and (11) respectively. When the two error values are equal, the value is EER.
[0065]
[0066] Among them, n FA represents the number of false acceptances, n oth represents the total number of test speech segments of other speakers, n FR represents the number of false rejections, n tar represents the total number of test speech segments of the target speaker.
[0067] The DCF minimum detection cost function, compared with the equal error rate, takes into account the prior probability and different costs, and is calculated as shown in Equation (12).
[0068] DCF = C FRR ×FRR×P target +C FAR ×FAR×(1 - P target ) (12)
[0069] Among them, C FAR represents the risk coefficient of false acceptance samples, P target represents the prior probability of positive example pairs, C FRR represents the risk coefficient of false rejection samples.
[0070] Furthermore, the frequency-domain multi-channel encoder not only enhances the model's ability to process features of different frequency scales, but also obtains a rich feature representation by processing and gradually fusing the outputs of each group step by step. These features contain information of multiple scales, enabling the model to capture subtle differences in the frequency domain and effectively enhancing the information of short-time speech in the frequency domain.
[0071] S3: After extracting the time-domain and frequency-domain features, masking processing is performed, and the recognition result is finally output.
[0072] Such as Figure 4As shown, the noise masking module enables the speaker recognition model to handle noise, thereby extracting robust speaker feature vectors. The noise masking module is applied earlier than the aggregation step to avoid information loss. The noise masking module extracts auxiliary context feature vectors to distinguish the features of the target speaker from the noise, so that the speaker recognition model can focus on the target speaker and blur the irrelevant noise.
[0073] The specific design of the noise masking module is as follows:
[0074] If the input acoustic feature is represented as x, then the output feature map of the selected hidden layer is given by Equation (1):
[0075] g(F(x)) (1)
[0076] where g(·) represents the transformation of the hidden layer and F(·) represents the transformation before the hidden layer.
[0077] Based on previous speech enhancement methods, the first step is to predict the enhanced acoustic feature, denoted as Then the output feature map of the corresponding hidden layer becomes Equation (2):
[0078]
[0079] The noise mask is applied to the feature map using a ratio for masking. If a mask that satisfies the following condition in Equation (3) can be found then it is possible to obtain an effect similar to speech enhancement:
[0080]
[0081] where ⊙ represents element-wise multiplication, and the values in are normalized to (0, 1).
[0082] The following section will demonstrate how to estimate a suitable noise mask. To achieve this goal, the present invention utilizes the feature vectors of the target speaker and the noise, which can be obtained from the auxiliary context feature vectors denoted as e. As Figure 3 shown, this module predicts the noise mask frame by frame:
[0083] F = F(x) (4)
[0084]
[0085] where σ(·) represents the Sigmoid function, δ(·) represents the combination of the ReLU function and BN, M is the predicted noise mask, M *t and F *t represent the current frame t. Denote the feature map after noise masking, and the auxiliary context feature e is used as a dynamic vector to control the activation threshold.
[0086] The noise masking module dynamically extracts the auxiliary context feature vector without the need to provide additional clean speech of the target speaker as a reference. The noise masking module will automatically find the target speaker during forward propagation.
[0087] When any of the following conditions is met, the noise masking module will identify a certain speaker as the target speaker: (1) This speaker occupies most of the speech in the utterance; (2) The volume of this speaker is significantly higher than that of other speakers. The speakers appearing in this utterance are called interfering speakers. In the speaker recognition task, it is beneficial to blur the speech and non-speech noise of the interfering speakers.
[0088] A simple way to distinguish the target speaker from the noise is to extract the utterance-level feature vector. First, combine all frames with the statistical pooling layer:
[0089]
[0090] After the statistical pooling layer, the FC layer maps the mean and standard deviation vectors into the auxiliary context features:
[0091]
[0092] Among them, e represents the auxiliary context feature vector, W3 represents the weight matrix, μ represents the mean vector calculated through statistical pooling, σ represents the standard deviation vector calculated through statistical pooling, b3 represents the bias term, and M *t represents the noise masking vector generated at the t-th frame, σ(·) represents the Sigmoid function, W2 represents the weight matrix, δ(·) represents the combination of the ReLU function and BN; W1 represents another weight matrix, and F *t represents the feature vector at the t-th frame, and b2 represents the bias term. Denote the feature map after noise masking processing, g(·) represents the transformation function, F(x) represents the feature representation; M represents the generated noise masking, and ⊙ represents multiplication.
[0093] On the other hand, this embodiment also provides a speaker recognition system for real scenarios, which includes:
[0094] A noise masking module that processes noise interference, extracts robust speaker feature vectors, focuses on the target speaker by dynamically predicting the noise masking, and blurs the irrelevant noise.
[0095] A time-domain multi-resolution encoder extracts features with different time resolutions in the time domain, sets multiple parallel branch filters, extracts features from the input speech signal through convolutional kernels of different scales, and performs feature fusion.
[0096] A frequency-domain multi-channel encoder extracts features in the frequency domain by means of grouped convolution, processes short-time speech information, divides the frequency-domain features into multiple groups, and uses different convolutional kernels for progressive processing and fusion.
[0097] A classifier inputs the processed features into the classifier for speaker recognition and outputs the recognition result.
[0098] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0099] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0100] More specific examples (nonexhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer diskettes (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0101] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
[0102] Example 2. To improve the accuracy of speaker recognition, the present invention designs a set of base model comparison experiments to verify the influence of different methods on the accuracy of speaker recognition. The purpose of this experiment is to verify whether the MTFNM-TDNN model proposed by the present invention is more competitive compared with the selected mainstream speaker recognition model, the ECAPA-TDNN model. The recognition effects of the ECAPA-TDNN model and the MTFNM-TDNN model proposed in this paper are compared in terms of two evaluation metrics: equal error rate and minimum detection cost function.
[0103] The experiment analyzes and compares the equal error rate and minimum cost test function of the two models, and the results are shown in Table 2. As can be seen from the table, the noise masking module and the multi-scale time-frequency channel module that increases short-time speech information in the present invention are beneficial to improving the performance of the speaker recognition system in a real scenario. An offline speaker recognition system has been designed on Jetson Nano, which has functions such as recording registration, recording recognition, and model selection. The experimental results are shown in Table 2.
[0104] Table 2 Comparison of the performance of the two models
[0105]
[0106] The present invention proposes a speaker recognition method for real scenarios and designs an offline speaker recognition system. This method addresses the problem of the performance degradation of the speaker recognition system caused by noise diversity in real scenarios through a noise masking module, and the problem of insufficient short-time speech information through a multi-scale time-frequency channel module. Moreover, the present invention implements an offline speaker recognition system, creating conditions for future model deployment.
[0107] The method of the present invention not only improves the performance of the speaker recognition system in real scenarios, but also conducts offline deployment of the model, and is applicable to speaker recognition applications in real scenarios, such as intelligent conference recording, in-vehicle human-computer interaction, etc. The technical solution of the present invention provides technical support for the development of speaker recognition technology in real scenarios, and has broad development prospects and commercial value.
[0108] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A speaker recognition method for real scenarios, characterized in that: include: Acquire external sound information, process it through multi-resolution convolution kernels, extract features in the time domain and perform feature fusion only in the time domain; After feature fusion in the time domain, group convolution is used to extract features in the frequency domain, process them, and perform feature fusion only in the frequency domain; After the frequency domain features are fused, mask processing is performed and the recognition results are finally output; The multi-resolution convolution kernel includes: setting three parallel branch filters of different scales, each branch consisting of two layers of one-dimensional convolution; The first layer of one-dimensional convolution is used for primary filtering, and the second layer of one-dimensional convolution is used for dimension matching; Feature fusion in the time domain includes: extracting features at different time resolutions through different convolution kernel sizes and step sizes, and splicing the features of three different time resolutions along the channel dimension; Find the minimum time step and clip all features so that all features have the same time step; Extracting features in the frequency domain includes: dividing the input frequency channel features into 4 groups, obtaining four grouping subsets X1, X2, X3, and X4; The first subset is X1, and the output is Y1; The second subset is X2, which is processed by a 7×7 convolution kernel to capture broader context information, and the output is Y2; The third subset is X3. Y2 is added to X3 and processed using a 5×5 convolution kernel to obtain output Y3. The fourth subset is X4. Y3 is added to X4 and processed using a 3×3 convolution kernel to obtain the output Y4. By concatenating the features of Y1, Y2, Y3, and Y4 in the frequency channel dimension, the combined feature vector is input into the classifier, and the final speaker recognition result is output through activation and normalization processing.
2. The real-scene speaker recognition method according to claim 1, characterized in that: Masking the time-frequency features includes: obtaining acoustic features of external sound information, predicting enhanced acoustic features, and obtaining a feature map; Apply a noise mask on the feature map to blur out irrelevant noise.
3. The real-scene speaker recognition method according to claim 2, characterized in that: Masking the time-frequency features also includes: extracting a discourse-level feature vector, calculating a mean and a standard deviation, and mapping the mean and the standard deviation to auxiliary context features; Auxiliary context features and acoustic features are used to generate noise masks and calculate the final feature map. The formula is expressed as: Where e represents the auxiliary context feature vector, W3 represents the weight matrix, represents the transpose of the weight matrix; μ represents the mean vector calculated by statistical pooling, σ represents the standard deviation vector calculated by statistical pooling, b3 represents the bias term, and M *t represents the noise mask vector generated in the tth frame, σ(·) represents the Sigmoid function, W2 represents the weight matrix, represents the transpose of the weight matrix; δ(·) represents the combination of ReLU function and BN; W1 represents another weight matrix, represents the transpose of the weight matrix; F *t represents the feature vector of the tth frame, b2 represents the bias term, represents the feature map after noise mask processing, g(·) represents the transformation function, F(x) represents the feature representation; M represents the generated noise mask, and ⊙ represents multiplication.
4. The real-scene speaker recognition method according to claim 3, characterized in that: The final output of the recognition result includes: inputting the combined feature vector into a classifier, and the classifier outputs the final speaker recognition result through activation and normalization processing.
5. A real-scene speaker recognition system using the method according to any one of claims 1 to 4, characterized in that: The noise mask module processes noise interference, extracts robust speaker feature vectors, and focuses on the target speaker and blurs irrelevant noise by dynamically predicting the noise mask; The time domain multi-resolution encoder extracts features of different time resolutions in the time domain, sets multiple parallel branch filters, extracts features from the input speech signal through convolution kernels of different scales, and performs feature fusion; The frequency domain multi-channel encoder extracts features in the frequency domain by grouping convolution to process short-term speech information, divides the frequency domain features into multiple groups, and uses different convolution kernels for step-by-step processing and fusion; Classifier, input the processed features into the classifier for speaker recognition and output the recognition results.
Citation Information
Patent Citations
Voice separation method based on time-frequency cross-domain feature selection
CN113113041A
Call scene speaker identification method
CN117690441A