Speaker recognition method based on channel attention in communication scene

By embedding the channel attention mechanism based on cross-network layer feature aggregation in the speaker recognition model, the problem of insufficient distinction between speaker representation in communication scenarios is solved, and more efficient feature selection and differentiated modeling is achieved, which improves the accuracy and robustness of speaker recognition.

CN120220694AActive Publication Date: 2025-06-27NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510625594.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-27
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In communication scenarios, the characteristics of different channel dimensions of neural networks change dynamically with the input sample, resulting in insufficient distinction between speaker representation and limiting the promotion of speaker recognition technology. The existing attention mechanisms are insufficient in dealing with multi-level information and cross-layer multi-semantics, resulting in insufficient generalization performance and robustness of the model.

Method used

The channel attention mechanism based on cross-network layer feature aggregation is adopted, and multiple channel attention network modules are embedded in the speaker recognition model. The learnable dictionary encoding unit and information aggregation unit are used to extract the comprehensive representation of features within a specific network layer and the global information representation across the network layer, and calculate the scaling coefficient and translation coefficient for feature calibration.

Benefits of technology

The model's accurate perception of the importance of different channel characteristics is improved, the distinction between speaker representation and speaker recognition is improved, and the model's robustness in communication scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220694A_ABST
    Figure CN120220694A_ABST
Patent Text Reader

Abstract

The invention relates to a speaker recognition method based on channel attention in a communication scene. The method comprises the following steps: constructing a speaker recognition model, wherein the model comprises a representation extraction backbone network and a speaker classification network which are connected in sequence; a channel attention mechanism based on cross-network layer feature aggregation is embedded into the representation extraction backbone network in the form of a plurality of channel attention network modules, and each channel attention network module comprises a learnable dictionary coding unit and an information aggregation unit; and carrying out optimization training on the speaker recognition model embedded with the channel attention network module, and executing a speaker recognition task in a communication scene by adopting the trained speaker recognition model. According to the method, the feature importance of each channel in the network can be accurately perceived by expressing the hidden layer feature information in a multi-level manner, so that feature selection and differentiation modeling are more efficiently carried out, and the method has an important value for improving the discrimination of speaker characterization and the accuracy of speaker recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of speaker recognition, and particularly to a speaker recognition method based on channel attention in a communication scenario. Background Art

[0002] Speaker recognition is a task of determining whether a given speech belongs to a specific speaker, and has wide applications in scenarios such as criminal investigation and surveillance, voice retrieval of suspect personnel, and security access control in high-sensitivity places. In recent years, with the rise of deep learning, the most advanced speaker recognition systems are mainly implemented through neural networks. The key core is to design an efficient neural network to extract highly discriminative speaker feature vectors from speech. In the process of speaker feature extraction, the mainstream method usually takes the speech spectrogram feature as the network input and uses the convolution algorithm to capture the acoustic texture clues in the time-frequency domain, which involves extracting speaker features in multiple channel dimensions. However, in scenarios such as communication, affected by factors such as device and channel differences, the features of different channel dimensions of the neural network show different importance with the dynamic change of the input samples. Processing these differential features in the same way will lead to insufficient discriminability of the speaker representation, restricting the popularization of speaker recognition technology in the communication scenario.

[0003] Currently, the attention mechanism is mainly used in speaker recognition to solve the above problems, including the squeeze-and-excitation attention and the global context vector attention mechanism. Among them, the squeeze-and-excitation attention mechanism has a simple algorithm and is convenient to deploy, and has good channel feature selection ability. However, this mechanism has too single a representation of the features in the middle layer of the network and lacks the representation of multi-level information. When there are a large number of irrelevant interferences in the training data, the model is prone to overfitting, resulting in insufficient generalization performance of the model. The global context vector attention mechanism has high operation efficiency and higher attention to the low-frequency feature region containing richer speaker clues in the time-frequency domain. However, in the presence of strong noise interference, this method is prone to introduce high-energy noise into the global context acquisition process, resulting in bias in estimating the importance of different channel features, affecting the attention distribution, and further reducing the robustness of the model. In addition, the above two methods have not fully considered the multi-level information utilization of single-layer multi-perspectives and cross-layer multi-semantics in the neural network, and the effectiveness of attention needs to be further improved, resulting in the need to improve the discriminability of the speaker representation and the accuracy of speaker recognition. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a speaker recognition method based on channel attention in a communication scenario, which can represent the hidden layer feature information at multiple levels, accurately perceive the importance of each channel feature in the network, and thus perform feature selection and differential modeling more efficiently, which has important value for improving the discriminability of the speaker representation and the accuracy of speaker recognition.

[0005] A speaker recognition method based on channel attention in a communication scenario, the method comprising: Construct a speaker recognition model, which is composed of a feature extraction backbone network and a speaker classification network connected in sequence; Embed a channel attention mechanism based on cross-network layer feature aggregation into the feature extraction backbone network in the form of multiple channel attention network modules; the channel attention network module includes a learnable dictionary encoding unit and an information aggregation unit; wherein, the learnable dictionary encoding unit is designed to obtain a comprehensive representation of the features within a specific network layer in the model, and is used to encode the acoustic features input by the previous network layer to obtain an encoded vector; the information aggregation unit is designed to obtain a global information representation of the cross-network layer features in the model, and is used to aggregate the encoded vectors of multiple network layers to obtain the global channel information representation of the current network layer, and according to the scaling coefficient and translation coefficient calculated from the global channel information representation, perform feature calibration on the acoustic features input by the previous network layer, and use the calibrated acoustic features as the input of the next network layer; Obtain a speaker voice dataset and input it into the speaker recognition model embedded with channel attention network modules for optimization training, and the model parameters are calculated by the speaker classification network to calculate the loss and optimize and update until a trained speaker recognition model is obtained; Use the trained speaker recognition model to perform the speaker recognition task in the communication scenario.

[0006] In one embodiment, the feature extraction backbone network in the model includes two processes: frame-level speaker feature extraction and segment-level speaker feature extraction; wherein, the frame-level speaker feature extraction includes four residual stages, and each residual stage is composed of four residual network blocks, and the channel attention network module is embedded after the last residual network block of each residual stage; the segment-level speaker feature extraction is composed of a statistical pooling layer and two fully connected networks connected in sequence; the speaker classification network is a one-layer AM-Softmax network.

[0007] In one embodiment, all networks in the model are constructed based on the TensorFlow deep learning framework, optimized using the Adam optimizer, and the model is trained using the AM-softmax loss function.

[0008] In one embodiment, the model training set is composed of the VoxCeleb 1 / 2 dataset, and the data augmentation algorithm provided by the Kaldi toolkit is used, and the test set includes the SITW dataset and the VOiCES dataset.

[0009] In one embodiment, the learnable dictionary encoding unit encodes the acoustic features input by the previous network layer to obtain an encoded vector, including: Assume that the acoustic features input by the previous network layer are ; where R represents the real number space, and the superscripts T , F and C represent the time dimension, frequency dimension, and channel dimension respectively; The learnable dictionary encoding unit first performs channel dimensionality reduction on through the convolution operation of to obtain the dimension-reduced acoustic features , and the expression is: ; where and represent the weight matrix and bias vector of respectively, and N represents the total number of preset dictionary components; Then, the frequency dimensions of and are merged into the channel dimension to obtain T the acoustic feature vector of the frame and T the dimension-reduced acoustic feature vector of the frame ; where is the number of frames, and ; Meanwhile, three sets of learnable parameter sets are defined, namely the mean vector , the projection vector , and the weight coefficient , and they are respectively initialized as random vectors, all-1 vectors, and 0; where represents the dictionary component serial number; According to the three sets of learnable parameter sets, perform learnable dictionary encoding on to obtain the encoded vector finally output by the learnable dictionary encoding unit, which is expressed as: ; ; ; where the encoded vector is composed of N dictionary component vectors , represents the weight assigned to the th dictionary component; , and respectively represent the i th mean vector, projection vector, and weight coefficient; is the transpose of , is the transpose of .

[0010] In one embodiment, the information aggregation unit aggregates the encoded vectors of multiple network layers to obtain the global channel information representation of the current network layer, including: Define that the memory unit aggregates the feature representations from the previous layer networks, and its calculation formula is as follows: ; where represents the aggregated features of the previous layer networks, represents the encoded vector output by the th layer network, represents the aggregated features of the previous layer networks, represents the non-linear projection process, and its calculation formula is as follows: ; where , and b are all learnable parameters, representing the first-layer weight matrix, the second-layer weight matrix, and the bias vector of the non-linear projection process respectively, is 's transpose; x is the input vector, is the activation function; For the th layer network, the expression of the aggregated global channel information representation is: ; ; where represents the feature dimension concatenation operation, represents and 's concatenation result, and respectively represent 's weight matrix and bias vector, is 's transpose.

[0011] In one embodiment, the information aggregation unit calculates the scaling coefficient and the translation coefficient based on the global channel information representation, calibrates the acoustic features input by the previous network layer, and uses the calibrated acoustic features as the input of the next network layer, including: According to the Global channel information representation obtained by aggregating layers of the network , calculate the scaling factor of the th layer of the network and the translation factor , which are respectively expressed as: ; ; wherein, represents the sigmoid activation function, and respectively represent the weight matrix and bias vector of , and respectively represent the weight matrix and bias vector of , and respectively represent the biases of and ; Using the scaling factor and the translation factor to calibrate the acoustic features input to the th layer of the network, and using the calibrated acoustic features as the input to the th layer of the network. The calibrated acoustic features are expressed as: .

[0012] In one embodiment, the acoustic features are 64-dimensional log Mel filter bank energy features, and after being processed by the Kaldi toolkit, the acoustic features are randomly cropped to a length of 2 to 4 seconds, and 64 speech segments with the same number of frames are grouped into a mini-batch.

[0013] The above method for speaker recognition based on channel attention in a communication scenario has the following beneficial effects: 1. Embed the channel attention mechanism based on cross-network layer feature aggregation into the speaker recognition model in the form of multiple channel attention network modules. Based on the learnable dictionary encoding unit in the network module, it can extract a comprehensive representation of the features within a specific network layer, enabling the speaker representation vector to have the ability to represent the input feature map from multiple perspectives. And based on the information aggregation unit in the network, it can fuse the features across network layers, thereby obtaining a more rich and comprehensive feature representation, which can improve the model's accurate perception ability of the importance of different channel features and enhance the discrimination of the speaker representation in the communication scenario.

[0014] 2. In the information aggregation unit, the weighted sum and bias coefficients are adaptively calculated through more accurate feature representation, and the scaling and translation operations are integrated in the feature calibration mechanism, so that attention can be applied at different levels, making the attention mechanism have a more efficient feature selection process and effectively improving the accuracy of speaker recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 FIG. is a schematic flowchart of a speaker recognition method based on channel attention in a communication scenario in an embodiment; Figure 2 FIG. is a schematic diagram of the overall architecture of a channel attention network module in an embodiment; Figure 3 FIG. is a schematic diagram of the structure of a learnable dictionary encoding unit in an embodiment; Figure 4 FIG. is a schematic diagram of the structure of an information aggregation unit in an embodiment; Figure 5 FIG. is a schematic diagram of the embedded deployment of a channel attention network module in a speaker recognition model in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0017] In one embodiment, as Figure 1 shown, a speaker recognition method based on channel attention in a communication scenario is provided, including the following steps: Step S1, constructing a speaker recognition model, which is composed of a feature extraction backbone network and a speaker classification network connected in sequence.

[0018] Step S2, embedding a channel attention mechanism based on cross-network layer feature aggregation into the feature extraction backbone network in the form of multiple channel attention network modules; the channel attention network module includes a learnable dictionary encoding unit and an information aggregation unit; wherein, the learnable dictionary encoding unit is designed to obtain a comprehensive representation of the features within a specific network layer in the model, encode the acoustic features input by the previous network layer to obtain an encoded vector; the information aggregation unit is designed to obtain a global information representation of the cross-network layer features in the model, aggregate the encoded vectors of multiple network layers to obtain the global channel information representation of the current network layer, and perform feature calibration on the acoustic features input by the previous network layer according to the scaling coefficient and translation coefficient calculated based on the global channel information representation, and use the calibrated acoustic features as the input of the next network layer.

[0019] Specifically, the overall architecture of the channel attention network module is as follows Figure 2 As shown, this network module can provide a modular and efficient channel attention method under the condition of limited growth of the number of model parameters and computational complexity. It can be embedded in the network layer of any speaker recognition model in the form of multiple network modules to efficiently utilize the important features in the network, thereby enhancing the discrimination ability of the extracted speaker representations and improving the recognition performance of the model. As Figure 2 shown, for the neural network hidden layer 1 in any speaker recognition model, its output is first input into the learnable dictionary encoding unit to obtain the encoded vector . At the same time, according to the aggregated features of the previous layer network, then and are respectively input into the information aggregation unit to jointly calculate the scaling coefficient and translation coefficient, and perform feature calibration on and then use it as the input of the neural network hidden layer 2. After transforming the existing speaker recognition model in the above way, more accurate speaker representation vectors can be extracted for the speaker recognition task.

[0020] The structure of the learnable dictionary encoding unit is as Figure 3 shown. Its process of encoding the acoustic features input from the previous network layer to obtain the encoded vector specifically includes: Assume that the acoustic features input from the previous network layer are ; where R represents the real number space, and the superscripts T , F and C represent the time dimension, frequency dimension, and channel dimension respectively; The learnable dictionary encoding unit first performs channel dimensionality reduction on through the convolution operation of to obtain the dimension-reduced acoustic features , and the expression is: ; where and represent the weight matrix and bias vector of respectively, and N represents the total number of preset dictionary components; Then, the frequency dimensions of and are merged into the channel dimension to obtain the acoustic feature vector T of frames and the dimension-reduced acoustic feature vector T of frames; where is the number of frames, and ; Meanwhile, three sets of learnable parameter sets are defined, namely the mean vector , the projection vector , and the weight coefficient , and they are respectively initialized as random vectors, all - 1 vectors, and 0; where represents the dictionary component serial number; According to the three sets of learnable parameter sets, is encoded with a learnable dictionary to obtain the encoded vector finally output by the learnable dictionary encoding unit, which is expressed as: ; ; ; Among them, the encoded vector is composed of N dictionary component vectors , represents is assigned to the weight of the th dictionary component; , and respectively represent the i th mean vector, projection vector, and weight coefficient; is the transpose of , is the transpose of . In the calculation formula of the above - mentioned encoding process, by introducing N dictionary components and using them as the central vectors, weighted statistical information is calculated, enabling the speaker representation vector to have the ability to represent the input feature map from multiple perspectives.

[0021] The structure of the information aggregation unit is as shown in Figure 4 , and the specific process of its feature aggregation and feature calibration includes: First, a memory unit is defined to aggregate the feature representations from the previous layers of the network, and its calculation formula is as follows: ; Among them, represents the aggregated features of the previous layers of the network, represents the encoded vector output by the th layer of the network, represents the aggregated features of the previous layers of the network, represents the non - linear projection process, and the calculation formula is as follows: ; Among them, 、 and b are all learnable parameters, representing the first-layer weight matrix, the second-layer weight matrix, and the bias vector of the non-linear projection process respectively, is the transpose of; x is the input vector, is the activation function; For the -th layer network, the global channel information representation aggregated is expressed as: ; ; Among them, represents the feature dimension concatenation operation, represents and the concatenation result of, and respectively represent the weight matrix and bias vector of, is the transpose of. In the above feature aggregation process, the memory unit calculates the global feature representations of each shallow network in a recursive manner, so as to obtain multi-level semantic features with complementarity across network layers, which have more robust and rich information than the feature representations extracted in a single-layer network.

[0022] The feature calibration mechanism aims to achieve attention selection for each channel feature according to importance. The core is how to calculate the feature scaling coefficient and translation coefficient by combining the global channel information representation obtained by the information aggregation unit. Specifically, according to the global channel information representation aggregated by the -th layer network, calculate the scaling coefficient and translation coefficient of the -th layer network, which are respectively expressed as: ; ; Among them, represents the sigmoid activation function, and respectively represent the weight matrix and bias vector of, and respectively represent the weight matrix and bias vector of, and respectively represent and Bias; Using a scaling factor and a translation factor for the acoustic features input to the layer network for feature calibration, and using the calibrated acoustic features as the input to the layer network. The calibrated acoustic features are expressed as: .

[0023] In the above feature calibration process, by weighting in the channel dimension, more importance can be given to important features, and the bias can give important features a more significant translation amount, so as to perform multi-faceted feature calibration operations on more important features in the channel dimension, thereby enhancing the effectiveness of attention.

[0024] Specifically, the present application embeds a channel attention mechanism based on cross-network layer feature aggregation in the form of multiple channel attention network modules into the Figure 5 speaker recognition model shown. The feature extraction backbone network in this model includes two processes: frame-level speaker feature extraction and segment-level speaker feature extraction. Among them, frame-level speaker feature extraction includes four residual stages, and each residual stage consists of four residual network blocks. The channel attention network module is embedded after the last residual network block in each residual stage. Segment-level speaker feature extraction consists of a statistical pooling layer and two fully connected networks connected in sequence. The speaker classification network is a one-layer AM-Softmax (Softmax with additional margin) network. AM-Softmax can enhance feature discriminability by introducing a fixed margin parameter to force an increase in the inter-class distance.

[0025] Furthermore, in this embodiment, the acoustic features are 64-dimensional log Mel filter bank energy features, and after being processed by the Kaldi toolkit, the acoustic features are randomly cropped to a length of 2 to 4 seconds, and 64 speech segments with the same number of frames are grouped into a mini-batch.

[0026] Step S3: Obtain a speaker speech dataset and input it into the speaker recognition model embedded with the channel attention network module for optimization training. The model parameters are calculated by the speaker classification network to calculate the loss and optimize the update until a trained speaker recognition model is obtained.

[0027] Specifically, all networks in the speaker recognition model in this embodiment are constructed based on the TensorFlow deep learning framework, optimized using the Adam optimizer, and the model is trained using the AM-softmax loss function. The hyperparameters s and m in the loss function of this embodiment are set to 30 and 0.15 respectively, and the learning rate is gradually reduced from 1e-3 to 1e-4. The model training set consists of the VoxCeleb1 / 2 dataset, and the data augmentation algorithm provided by the Kaldi toolkit is used. The test set includes the SITW dataset and the VOiCES dataset. Among them, the VoxCeleb1 / 2 dataset is a large-scale speaker recognition dataset, containing the speech data of celebrity interview videos. The SITW dataset is also called the real-scenario speaker dataset, which is used to test the speaker recognition performance in a real and complex environment. The VOiCES dataset is also called the noisy environment speech dataset, which is a test dataset specifically designed for noisy environments.

[0028] Step S4: Use the trained speaker recognition model to perform the speaker recognition task in the communication scenario.

[0029] In the above method for speaker recognition based on channel attention in a communication scenario, by embedding a channel attention mechanism based on cross-network layer feature aggregation in the form of a channel attention network module in the speaker model, it is possible to extract a comprehensive representation of the features within a specific network layer based on the learnable dictionary encoding unit in this network module, enabling the speaker representation vector to have the ability to represent the input feature map from multiple perspectives. And based on the information aggregation unit in this network, it is possible to fuse the features across network layers, thereby obtaining a more rich and comprehensive feature representation, improving the accuracy of the speaker representation extracted by the model in the communication scenario, and enabling more accurate speaker recognition.

[0030] Furthermore, to verify the effectiveness of the method proposed in this application, a comparative experiment was further conducted. A total of four types of systems were constructed for comparison in the experiment, including the baseline system and the system applying this application, and two types of comparative systems that only use attention methods, namely Comparative System 1 (squeeze-and-excitation attention mechanism) and Comparative System 2 (global context vector attention mechanism). The embedding positions and parameter settings of these two plug-and-play systems in the speaker recognition model are the same as those of the method proposed in this application. Unless otherwise specified, the experimental settings of all systems are the same as those of the baseline system.

[0031] In all systems, the parameter NSet to 4. The backend scoring uses the PLDA (Probabilistic Linear Discriminant Analysis) model, and its training and operation processes are both implemented by Kaldi. Before the operation, the speaker representation needs to be length-normalized, centered, and dimension-reduced by LDA (Linear Discriminant Analysis). The final experimental results are evaluated using EER (Equal Error Rate) and minDCF (Minimum Normalized Detection Cost Function). The lower the values of these two metrics, the better the performance. When calculating minDCF, the prior target probability is set to 0.01.

[0032] The performance comparison results of the four types of systems on the SITW and VOiCES public evaluation datasets are shown in Table 1 and Table 2 respectively. As shown in Table 1, on the SITW dataset containing a wide range of open recording scenarios, the system applying the present application shows significant advantages in both the EER and minDCF metrics, proving that the method proposed in the present application has excellent scene generalization. As shown in Table 2, on the VOiCES dataset containing strong reverberation noise, the system applying the present application also outperforms the comparison systems in various metrics of the development set and the test set. Especially in the test set, a significant breakthrough in performance is achieved, verifying the higher robustness of the method proposed in the present application.

[0033] Table 1 Performance comparison results of four types of systems on the SITW public evaluation dataset

[0034] Table 2 Performance comparison results of four types of systems on the VOiCES public evaluation dataset

[0035] Table 3 lists the comparison results of the four types of systems in terms of the number of parameters and the amount of computation under the condition of 100-frame feature input. It can be seen from Table 3 that although the system applying the present application has a higher number of parameters and a larger amount of computation, the growth of the computing resource requirements is overall controllable and has limited impact on the actual application deployment. In addition, due to the significant performance advantages of the method proposed in the present application in multi-scenarios and strong noise interference, it is more popularizable and universal.

[0036] Table 3 Comparison results of the number of parameters and the amount of computation of four types of systems

[0037] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0038] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A speaker recognition method based on channel attention in a communication scenario, characterized in that: The method comprises: Build a speaker recognition model, which consists of a representation extraction backbone network and a speaker classification network connected in sequence; The channel attention mechanism based on cross-network layer feature aggregation is embedded in the representation extraction backbone network in the form of multiple channel attention network modules; the channel attention network module includes a learnable dictionary encoding unit and an information aggregation unit; wherein the learnable dictionary encoding unit is intended to obtain a comprehensive representation of features in a specific network layer in the model, and is used to encode the acoustic features of the previous network layer input to obtain a coding vector; the information aggregation unit is intended to obtain a global information representation of cross-network layer features in the model, and is used to aggregate the coding vectors of multiple network layers to obtain a global channel information representation of the current network layer, and feature calibrate the acoustic features of the previous network layer input according to the scaling coefficient and translation coefficient calculated by the global channel information representation, and use the calibrated acoustic features as the input of the next network layer; Obtain a speaker speech dataset and input it into a speaker recognition model embedded with a channel attention network module for optimization training. The model parameters are optimized and updated by the speaker classification network in conjunction with the loss calculation until a trained speaker recognition model is obtained. The trained speaker recognition model is used to perform speaker recognition tasks in communication scenarios.

2. The method according to claim 1, characterized in that The representation extraction backbone network in the model includes two processes: frame-level speaker representation extraction and segment-level speaker representation extraction; wherein, frame-level speaker representation extraction includes four residual stages, and each residual stage consists of four residual network blocks, and the channel attention network module is embedded after the last residual network block of each residual stage; segment-level speaker representation extraction consists of a statistical pooling layer and two fully connected networks connected in sequence; the speaker classification network is a layer of AM-Softmax network.

3. The method according to claim 2, characterized in that All networks in the model are built based on the TensorFlow deep learning framework and optimized using the Adam optimizer, and the model is trained using the AM-softmax loss function.

4. The method according to claim 3, characterized in that The model training set consists of the VoxCeleb 1 / 2 dataset and uses the data augmentation algorithm that comes with the Kaldi toolkit. The test set includes the SITW dataset and the VOiCES dataset.

5. The method according to claim 1, characterized in that The learnable dictionary encoding unit encodes the acoustic features input by the previous network layer to obtain an encoding vector, including: Assume that the acoustic features of the previous network layer input are ;in, R denotes the real number space, the superscript T , F and C Represent the time dimension, frequency dimension and channel dimension respectively; The learnable dictionary encoding unit first passes through The convolution operation of Perform channel dimensionality reduction to obtain acoustic features after dimensionality reduction , the expression is: ; in, and Respectively The weight matrix and bias vector of N Indicates the total number of preset dictionary components; Then, and The frequency dimension of is merged into the channel dimension, and we get T Acoustic feature vector of the frame as well as T Acoustic feature vector after frame dimension reduction ;in, is the number of frames, and ; At the same time, three sets of learnable parameter sets are defined, namely, the mean vector , projection vector And the weight coefficient , and initialized to random vector, all 1 vector and 0 respectively; among them, Indicates the dictionary component number; According to three sets of learnable parameter sets Perform learnable dictionary encoding to obtain the encoding vector finally output by the learnable dictionary encoding unit , expressed as: ; ; ; Among them, the encoding vector Depend on N dictionary component vector composition, express Assigned to The weights of the dictionary components; , and Respectively represent i mean vectors, projection vectors and weight coefficients; for The transpose of for The transpose of .

6. The method according to claim 1, characterized in that The information aggregation unit aggregates the encoding vectors of multiple network layers to obtain a global channel information representation of the current network layer, including: The definition memory unit aggregates the previous The feature representation of the layer network is calculated as follows: ; in, Before Aggregate features of layer networks, Indicates The encoded vector output by the layer network, Before Aggregate features of layer networks, Represents the nonlinear projection process, and the calculation formula is as follows: ; in, , and b are learnable parameters, representing the first layer weight matrix, the second layer weight matrix and the bias vector of the nonlinear projection process, respectively. for The transpose of x is the input vector, is the activation function; For Layer network, aggregated global channel information representation The expression is: ; ; in, represents the feature dimension concatenation operation, express and The splicing result is and Respectively The weight matrix and bias vector of for The transpose of .

7. The method according to claim 6, characterized in that The information aggregation unit performs feature calibration on the acoustic features input by the previous network layer according to the scaling coefficient and the translation coefficient calculated by the global channel information representation, and uses the calibrated acoustic features as the input of the next network layer, including: According to Global channel information representation obtained by layer network aggregation , calculate the Scaling factor of the layer network and translation coefficient , respectively expressed as: ; ; in, represents the sigmoid activation function, and Respectively The weight matrix and bias vector of and Respectively The weight matrix and bias vector of and Respectively and Bias of Using the zoom factor and translation coefficient For Acoustic features of the layer network input Perform feature calibration and convert the calibrated acoustic features into As the Input of the layer network, calibrated acoustic features It is expressed as: 。 8. The method according to claim 1, 5 or 7, characterized in that: The acoustic features are 64-dimensional logarithmic Mel filter bank energy features, and after being processed by the Kaldi toolkit, the acoustic features are randomly cropped to a length of 2 to 4 seconds, and 64 speech segments with the same number of frames are grouped into a mini-batch.

Citation Information

Patent Citations

  • Multi-attention feature fusion speaker recognition method

    CN113763965A

  • Speech wakeup method and apparatus, and storage medium and system

    WO2022206602A1