Voiceprint recognition method based on spatial cross-learning multi-scale attention feature module

By adopting the method of multi-scale attention feature module based on spatial cross-learning in voiceprint recognition technology, the problem of inefficiency in traditional technology is solved, and more efficient and accurate voiceprint recognition is achieved.

CN119091887BActive Publication Date: 2025-05-16CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411266942.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2025-05-16
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

Traditional voiceprint recognition technology requires a lot of manual processing time, resulting in low recognition efficiency.

Method used

The voiceprint recognition method based on spatial cross-learning multi-scale attention feature module is adopted. By extracting the two-dimensional spectrum features of the audio, multi-level feature recognition and weighted fusion processing are carried out, and the time domain, frequency domain and global features are fully utilized.

Benefits of technology

The accuracy and efficiency of voiceprint recognition are improved, and the feature expression ability is improved through multi-level feature recognition and weighted fusion processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091887B_ABST
    Figure CN119091887B_ABST
Patent Text Reader

Abstract

The present application relates to a voiceprint recognition method, device, computer equipment, computer-readable storage medium and computer program product based on a spatial cross-learning multi-scale attention feature module, which can be used in the field of computer technology. The method includes: extracting two-dimensional spectrum features of audio; performing feature map recognition on the two-dimensional spectrum features to obtain a multi-channel three-dimensional feature map; grouping the multi-channel three-dimensional feature map to obtain an original sub-feature map group; performing time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature map group respectively; generating basic weights according to time domain features and frequency domain features, and performing weighted processing on the original sub-feature map group; generating target weights according to the target sub-feature map group and the global features; performing weighted fusion processing on the original sub-feature map group using the target weights to obtain a fused feature map; performing voiceprint recognition on the audio according to the fused feature map to obtain the voiceprint recognition result of the audio. The use of this method can improve the efficiency of voiceprint recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a voiceprint recognition method, apparatus, computer equipment, computer-readable storage medium and computer program product based on a spatial cross-learning multi-scale attention feature module. Background Art

[0002] With the development of computer technology, voiceprint recognition technology has important applications in many fields. How to perform voiceprint recognition efficiently has become an important research direction.

[0003] Traditional technology usually performs voiceprint recognition by manually extracting audio information; however, voiceprint recognition using this method requires a lot of manual processing time, resulting in low efficiency of voiceprint recognition. Summary of the invention

[0004] Based on this, it is necessary to provide a voiceprint recognition method, device, computer equipment, computer-readable storage medium and computer program product based on a spatial cross-learning multi-scale attention feature module that can improve the efficiency of voiceprint recognition in response to the above-mentioned technical problems.

[0005] In a first aspect, the present application provides a voiceprint recognition method based on a spatial cross-learning multi-scale attention feature module. The method comprises:

[0006] Extract two-dimensional spectrum features of audio;

[0007] Performing feature map recognition on the two-dimensional spectrum features through a feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio;

[0008] Grouping the multi-channel three-dimensional feature map by channel layer through a target residual block to obtain an original sub-feature map group of the audio;

[0009] Performing time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group respectively to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group;

[0010] Generate basic weights according to the time domain features and the frequency domain features, and use the basic weights to perform weighted processing on the original sub-feature graph group to obtain a target sub-feature graph group;

[0011] Generate a target weight according to the target sub-feature map group and the global feature;

[0012] Performing weighted fusion processing on the original sub-feature graph group using the target weight to obtain a fused feature graph of the audio;

[0013] According to the fused feature map, voiceprint recognition is performed on the audio to obtain a voiceprint recognition result of the audio.

[0014] In one of the embodiments, before performing voiceprint recognition on the audio according to the fused feature graph to obtain the voiceprint recognition result of the audio, the method further includes:

[0015] Using a residual network to perform feature recognition on the fused feature map to obtain a target feature map;

[0016] The performing voiceprint recognition on the audio according to the fused feature graph to obtain the voiceprint recognition result of the audio includes:

[0017] According to the target feature map, voiceprint recognition is performed on the audio to obtain the voiceprint recognition result.

[0018] In one embodiment, performing voiceprint recognition on the audio according to the target feature graph to obtain the voiceprint recognition result includes:

[0019] Using an attention pooling layer and a linear layer to perform feature recognition on the target feature map to obtain a voiceprint feature of the target feature map;

[0020] The audio is subjected to voiceprint recognition according to the voiceprint feature to obtain the voiceprint recognition result.

[0021] In one embodiment, the performing time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group respectively to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group includes:

[0022] Performing a time domain dimension pooling operation on the original sub-feature map group to obtain the time domain feature;

[0023] Performing frequency domain dimension pooling operation on the original sub-feature map group to obtain the frequency domain feature;

[0024] A global convolution operation is performed on the original sub-feature map group to obtain the global feature.

[0025] In one embodiment, generating a basic weight according to the time domain feature and the frequency domain feature includes:

[0026] Performing splicing processing on the time domain feature and the frequency domain feature to obtain a splicing feature;

[0027] Performing a convolution operation on the splicing feature to obtain a convolution operation result;

[0028] An activation process is performed on the convolution operation result to obtain the basic weight.

[0029] In one embodiment, extracting the two-dimensional frequency spectrum features of the audio includes:

[0030] Acquire a speech signal as the audio;

[0031] The audio is subjected to a Mel spectrum conversion process to obtain the two-dimensional spectrum feature.

[0032] In a second aspect, the present application also provides a voiceprint recognition device based on a spatial cross-learning multi-scale attention feature module. The device comprises:

[0033] A feature extraction module is used to extract two-dimensional spectrum features of audio;

[0034] A first recognition module, configured to perform feature map recognition on the two-dimensional spectrum feature through a feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio;

[0035] A feature grouping module, used for grouping the multi-channel three-dimensional feature map by channel layer through a target residual block to obtain an original sub-feature map group of the audio;

[0036] A second recognition module is used to perform time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group, to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group;

[0037] A first generating module, configured to generate basic weights according to the time domain features and the frequency domain features, and perform weighted processing on the original sub-feature graph group using the basic weights to obtain a target sub-feature graph group;

[0038] A second generating module, used for generating a target weight according to the target sub-feature graph group and the global feature;

[0039] A feature fusion module, used to perform weighted fusion processing on the original sub-feature map group using the target weight to obtain a fused feature map of the audio;

[0040] The voiceprint recognition module is used to perform voiceprint recognition on the audio according to the fusion feature map to obtain the voiceprint recognition result of the audio.

[0041] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0042] Extract two-dimensional spectrum features of audio;

[0043] Performing feature map recognition on the two-dimensional spectrum features through a feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio;

[0044] Grouping the multi-channel three-dimensional feature map by channel layer through a target residual block to obtain an original sub-feature map group of the audio;

[0045] Performing time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group respectively to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group;

[0046] Generate basic weights according to the time domain features and the frequency domain features, and use the basic weights to perform weighted processing on the original sub-feature graph group to obtain a target sub-feature graph group;

[0047] Generate a target weight according to the target sub-feature map group and the global feature;

[0048] Performing weighted fusion processing on the original sub-feature graph group using the target weight to obtain a fused feature graph of the audio;

[0049] According to the fused feature map, voiceprint recognition is performed on the audio to obtain a voiceprint recognition result of the audio.

[0050] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0051] Extract two-dimensional spectrum features of audio;

[0052] Performing feature map recognition on the two-dimensional spectrum features through a feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio;

[0053] Grouping the multi-channel three-dimensional feature map by channel layer through a target residual block to obtain an original sub-feature map group of the audio;

[0054] Performing time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group respectively to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group;

[0055] Generate basic weights according to the time domain features and the frequency domain features, and use the basic weights to perform weighted processing on the original sub-feature graph group to obtain a target sub-feature graph group;

[0056] Generate a target weight according to the target sub-feature map group and the global feature;

[0057] Performing weighted fusion processing on the original sub-feature graph group using the target weight to obtain a fused feature graph of the audio;

[0058] According to the fused feature map, voiceprint recognition is performed on the audio to obtain a voiceprint recognition result of the audio.

[0059] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0060] Extract two-dimensional spectrum features of audio;

[0061] Performing feature map recognition on the two-dimensional spectrum features through a feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio;

[0062] Grouping the multi-channel three-dimensional feature map by channel layer through a target residual block to obtain an original sub-feature map group of the audio;

[0063] Performing time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group respectively to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group;

[0064] Generate basic weights according to the time domain features and the frequency domain features, and use the basic weights to perform weighted processing on the original sub-feature graph group to obtain a target sub-feature graph group;

[0065] Generate a target weight according to the target sub-feature map group and the global feature;

[0066] Performing weighted fusion processing on the original sub-feature graph group using the target weight to obtain a fused feature graph of the audio;

[0067] According to the fused feature map, voiceprint recognition is performed on the audio to obtain a voiceprint recognition result of the audio.

[0068] The voiceprint recognition method, device, computer equipment, computer-readable storage medium and computer program product based on the spatial cross-learning multi-scale attention feature module extract the two-dimensional spectrum features of the audio; perform feature map recognition on the two-dimensional spectrum features through the feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio; group the multi-channel three-dimensional feature map by channel layer through the target residual block to obtain the original sub-feature map group of the audio; perform time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature map group respectively to obtain the time domain features, frequency domain features and global features of the original sub-feature map group; generate basic weights according to the time domain features and the frequency domain features, and use the basic weights to weight the original sub-feature map group to obtain a target sub-feature map group; generate target weights according to the target sub-feature map group and the global features; use the target weights to weightedly fuse the original sub-feature map group to obtain a fused feature map of the audio; perform voiceprint recognition on the audio according to the fused feature map to obtain a voiceprint recognition result of the audio. Through multi-level feature recognition and weighted fusion processing, this solution is conducive to making full use of the time domain features, frequency domain features and global features of the audio, thereby improving the expressiveness of the features and improving the accuracy and efficiency of voiceprint recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0070] Figure 1 1. It is a flowchart of a voiceprint recognition method based on a spatial cross-learning multi-scale attention feature module in one embodiment;

[0071] Figure 2 A system structure diagram of a voiceprint recognition method based on a spatial cross-learning multi-scale attention feature module in one embodiment;

[0072] Figure 3 2 is a schematic diagram of the structure of a multi-scale attention feature fusion module based on spatial cross-learning in one embodiment;

[0073] Figure 4 is a schematic diagram of an improved residual block in one embodiment;

[0074] Figure 5 It is a structural block diagram of a voiceprint recognition device based on a spatial cross-learning multi-scale attention feature module in one embodiment;

[0075] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0076] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0077] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0078] In an exemplary embodiment, Figure 1 As shown, a voiceprint recognition method based on a spatial cross-learning multi-scale attention feature module is provided. This embodiment uses the method applied to a terminal as an example for illustration; it can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers, etc.; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. In this embodiment, the method includes the following steps:

[0079] Step S101, extracting two-dimensional frequency spectrum features of audio.

[0080] Step S102, performing feature map recognition on the two-dimensional spectrum features through a feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio.

[0081] Step S103, grouping the multi-channel three-dimensional feature maps by channel layer through the target residual block to obtain an original sub-feature map group of the audio.

[0082] Step S104, performing time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group respectively, to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group.

[0083] Step S105, generating basic weights according to the time domain features and the frequency domain features, and performing weighted processing on the original sub-feature graph group using the basic weights to obtain a target sub-feature graph group.

[0084] Step S106, generating a target weight according to the target sub-feature graph group and the global feature.

[0085] Step S107, using the target weights to perform weighted fusion processing on the original sub-feature map group to obtain a fused feature map of the audio.

[0086] Step S108: Perform voiceprint recognition on the audio according to the fused feature graph to obtain a voiceprint recognition result of the audio.

[0087] The two-dimensional spectrum feature may be a representation of the audio signal in two dimensions of time and frequency. For example, the two-dimensional spectrum feature may be a Mel-spectrogram or a spectrogram.

[0088] The feature map recognition layer may be a neural network layer for extracting higher-level feature representations from input features, for example, the feature map recognition layer may be a convolutional layer.

[0089] Among them, the multi-channel three-dimensional feature map can be a three-dimensional feature representation of multiple channels obtained after being processed by the feature map recognition layer. For example, the multi-channel three-dimensional feature map can be a convolution feature map with multiple filter outputs.

[0090] The target residual block may be a neural network module for learning residual mapping, for example, the target residual block may be an improved convolution residual module integrating time-frequency difference information.

[0091] The original sub-feature map group may be a plurality of groups of sub-feature maps obtained by grouping the multi-channel three-dimensional feature map by channel.

[0092] The time domain feature may be a feature that describes the characteristics of the audio signal changing over time.

[0093] The frequency domain feature may be a feature that describes the frequency distribution characteristics of the audio signal.

[0094] The global feature may be a feature that describes the overall characteristics of the entire audio signal. For example, the global feature may be the average energy or spectrum flatness of the audio.

[0095] Among them, the basic weight can be a weight coefficient used to preliminarily weight the original sub-feature map group. For example, the basic weight can be an attention weight calculated based on time domain features and frequency domain features.

[0096] The target sub-feature graph group may be a sub-feature graph group obtained by weighting with a basic weight. For example, the target sub-feature graph group may be the product of the original sub-feature graph group and the basic weight.

[0097] Among them, the target weight can be a weight coefficient used to further weightedly fuse the target sub-feature map group. For example, the target weight can be an attention weight calculated based on the target sub-feature map group and the global feature.

[0098] Among them, the fused feature map can be the final feature representation obtained after weighted fusion with the target weight. For example, the fused feature map can be the weighted sum of the target sub-feature map group and the target weight.

[0099] The voiceprint recognition result may be a judgment result of the speaker of the input audio, for example, the voiceprint recognition result may be a pre-registered speaker identifier or a prompt message indicating that the voiceprint does not match.

[0100] Optionally, the terminal preprocesses the input audio signal, including noise reduction and framing operations; extracts two-dimensional spectral features of the audio, such as a Mel spectrum map; performs a convolution operation on the two-dimensional spectral features through a feature map recognition layer to obtain a multi-channel three-dimensional feature map; uses a target residual block to group the multi-channel three-dimensional feature map by channel to form an original sub-feature map group; applies a time domain feature recognition module, a frequency domain feature recognition module and a global feature recognition module to the original sub-feature map group to obtain corresponding time domain features, frequency domain features and global features; generates basic weights according to the time domain features and the frequency domain features, and uses the basic weights to weight the original sub-feature map group to obtain a target sub-feature map group; generates target weights according to the target sub-feature map group and the global features, and uses the target weights to weightedly fuse the original sub-feature map group to obtain a fused feature map; performs voiceprint recognition on the audio based on the fused feature map to obtain a voiceprint recognition result of the audio.

[0101] In the above-mentioned voiceprint recognition method based on spatial cross-learning multi-scale attention feature module, two-dimensional spectrum features of audio are extracted; feature map recognition is performed on the two-dimensional spectrum features through the feature map recognition layer to obtain a multi-channel three-dimensional feature map of audio; the multi-channel three-dimensional feature map is grouped according to the channel layer through the target residual block to obtain the original sub-feature map group of audio; time domain feature recognition, frequency domain feature recognition and global feature recognition are performed on the original sub-feature map group respectively to obtain the time domain features, frequency domain features and global features of the original sub-feature map group; basic weights are generated according to the time domain features and frequency domain features, and the original sub-feature map group is weightedly processed using the basic weights to obtain the target sub-feature map group; target weights are generated according to the target sub-feature map group and the global features; the original sub-feature map group is weightedly fused using the target weights to obtain the fused feature map of audio; voiceprint recognition is performed on the audio according to the fused feature map to obtain the voiceprint recognition result of the audio. Through multi-level feature recognition and weighted fusion processing, this solution is conducive to making full use of the time domain features, frequency domain features and global features of the audio, thereby improving the expressiveness of the features and improving the accuracy and efficiency of voiceprint recognition.

[0102] In an exemplary embodiment, before performing voiceprint recognition on the audio according to the fused feature map to obtain the voiceprint recognition result of the audio, the following contents are also included: performing feature recognition on the fused feature map using a residual network to obtain a target feature map; performing voiceprint recognition on the audio according to the fused feature map to obtain the voiceprint recognition result of the audio, which specifically includes the following contents: performing voiceprint recognition on the audio according to the target feature map to obtain a voiceprint recognition result.

[0103] The residual network may be a deep neural network structure, for example, the residual network may include a target residual block and an original residual block.

[0104] Among them, the target feature map can be a higher-level feature representation obtained after being processed by the residual network. For example, the target feature map can be the output feature map of the last convolutional layer of the residual network.

[0105] Optionally, the terminal uses a residual network to perform feature recognition on the fused feature map to obtain a target feature map; based on the target feature map, voiceprint recognition is performed on the audio to obtain a voiceprint recognition result.

[0106] For example, the terminal further extracts and recognizes features of the fused feature map through the residual network; obtains the target feature map after processing by the residual network; and performs voiceprint recognition on the audio according to the target feature map to obtain the voiceprint recognition result.

[0107] The technical solution provided in this embodiment further processes the fused feature map through a residual network, which is conducive to extracting deeper voiceprint features, thereby helping to improve the accuracy of voiceprint recognition.

[0108] In an exemplary embodiment, voiceprint recognition is performed on audio according to a target feature map to obtain a voiceprint recognition result, which specifically includes the following contents: using an attention pooling layer and a linear layer to perform feature recognition on the target feature map to obtain the voiceprint features of the target feature map; performing voiceprint recognition on the audio according to the voiceprint features to obtain a voiceprint recognition result.

[0109] Among them, the attention pooling layer can be a layer for feature aggregation by learning the importance weights of different features. For example, the attention pooling layer can be a self-attention mechanism or a multi-head attention mechanism.

[0110] Among them, the linear layer can be a fully connected layer, which is used to perform linear transformation and dimensionality reduction on features.

[0111] The voiceprint feature may be a feature vector obtained after being processed by the attention pooling layer and the linear layer, and is used to characterize the voice characteristics of the speaker. For example, the voiceprint feature may be a 256-dimensional vector.

[0112] Optionally, the terminal performs weighted aggregation of features at different time and frequency positions in the target feature map through an attention pooling layer to obtain voiceprint information; uses a linear layer to perform dimensionality reduction and linear transformation on the aggregated features (voiceprint information) to obtain a final voiceprint feature vector as the voiceprint feature of the target feature map; performs voiceprint recognition on the audio according to the voiceprint feature to obtain a voiceprint recognition result.

[0113] The technical solution provided in this embodiment processes the target feature map through the attention pooling layer and the linear layer, which is conducive to capturing key voiceprint information and compressing the feature dimension, thereby facilitating the extraction of more refined voiceprint features and improving the accuracy and efficiency of voiceprint recognition.

[0114] In an exemplary embodiment, time domain feature recognition, frequency domain feature recognition and global feature recognition are performed on the original sub-feature graph group respectively to obtain the time domain features, frequency domain features and global features of the original sub-feature graph group, which specifically includes the following contents: time domain dimension pooling operation is performed on the original sub-feature graph group to obtain time domain features; frequency domain dimension pooling operation is performed on the original sub-feature graph group to obtain frequency domain features; global convolution operation is performed on the original sub-feature graph group to obtain global features.

[0115] Among them, the time domain dimension pooling operation can be a feature aggregation operation performed on the original sub-feature map group in the time dimension. For example, the time domain dimension pooling operation can be a maximum pooling or average pooling operation performed on the time axis.

[0116] Among them, the frequency domain dimension pooling operation can be a feature aggregation operation performed on the original sub-feature map group in the frequency dimension. For example, the frequency domain dimension pooling operation can be a maximum pooling or average pooling operation performed on the frequency axis.

[0117] Among them, the global convolution operation can be a large-scale feature extraction operation performed on the original sub-feature map group to capture overall context information, such as a two-dimensional convolution operation.

[0118] Among them, the time domain feature can be a feature representation obtained by time domain dimension pooling operation, which reflects the time change characteristics of the audio signal.

[0119] Among them, the frequency domain feature can be a feature representation obtained by frequency domain dimension pooling operation, which reflects the frequency distribution characteristics of the audio signal.

[0120] The global feature may be a feature representation obtained through a global convolution operation and reflecting the overall characteristics of the audio signal.

[0121] Optionally, the terminal performs feature aggregation on the original sub-feature graph group in the time dimension through a pooling operation in the time domain dimension to obtain time domain features that reflect the time variation characteristics of the audio signal; performs feature aggregation on the original sub-feature graph group in the frequency dimension through a pooling operation in the frequency domain dimension to obtain frequency domain features that reflect the frequency distribution characteristics of the audio signal; and performs large-range (preset range) feature extraction on the original sub-feature graph group through a global convolution operation to obtain global features that reflect the overall characteristics of the audio signal.

[0122] The technical solution provided in this embodiment is beneficial for simultaneously capturing the time domain changes, frequency domain distribution and overall characteristics of the audio signal by performing time domain dimension pooling, frequency domain dimension pooling and global convolution operations on the original sub-feature graph group, thereby facilitating the extraction of more comprehensive and multi-dimensional voiceprint features and improving the accuracy of voiceprint recognition.

[0123] In an exemplary embodiment, basic weights are generated according to time domain features and frequency domain features, specifically including the following contents: splicing the time domain features and frequency domain features to obtain splicing features; performing convolution operation on the splicing features to obtain convolution operation results; and activating the convolution operation results to obtain basic weights.

[0124] Among them, the splicing process can be an operation of merging two or more feature tensors in a preset dimension. For example, the splicing process can be connecting the time domain features and the frequency domain features in the channel dimension to form a new feature tensor with more channels.

[0125] The splicing feature may be a feature tensor obtained through splicing processing. For example, the splicing feature may be a three-dimensional tensor whose number of channels is equal to the sum of the number of channels of the time domain feature and the frequency domain feature.

[0126] The convolution operation may be an operation of extracting and transforming local features of input features. For example, the convolution operation may be a two-dimensional convolution operation performed on the splicing features using a convolution kernel of size 3×3 or 5×5.

[0127] The convolution operation result may be a feature map obtained after the convolution operation. For example, the convolution operation result may be a three-dimensional tensor whose number of channels is determined by the number of convolution kernels.

[0128] The activation process may be an operation of performing a nonlinear transformation on a feature to introduce nonlinear properties. For example, the activation process may be a nonlinear transformation of a convolution operation result using an activation function.

[0129] Optionally, the terminal concatenates the time domain features and the frequency domain features, integrates the feature information of two different dimensions together to form a concatenated feature; performs a convolution operation on the concatenated feature, uses a convolution kernel to extract a higher-level feature combination, and obtains a convolution operation result; and activates the convolution operation result through an activation function to obtain a basic weight.

[0130] The technical solution provided in this embodiment is conducive to fusing voiceprint information of different dimensions and extracting high-level feature combinations by splicing, convolution and activation processing of time domain features and frequency domain features, thereby facilitating the generation of more accurate basic weights and improving the accuracy of voiceprint recognition.

[0131] In an exemplary embodiment, extracting two-dimensional spectrum features of audio specifically includes the following contents: acquiring a speech signal as audio; and performing Mel spectrum conversion processing on the audio to obtain two-dimensional spectrum features.

[0132] The speech signal may be data containing the speaker's voice information, for example, the speech signal may be a collected original audio waveform, or pre-processed audio data.

[0133] The Mel spectrum conversion process may be a process of converting an audio signal into a spectrum representation based on human ear perception characteristics. For example, the Mel spectrum conversion process may include steps such as short-time Fourier transform, Mel filter bank processing, and logarithmic compression.

[0134] The two-dimensional spectrum feature may be a two-dimensional matrix representing the time-frequency characteristics of the audio signal, for example, the two-dimensional spectrum feature may be a Mel-spectrogram.

[0135] Optionally, the terminal obtains a speech signal as the audio to be processed; and performs Mel spectrum conversion on the audio to be processed to obtain a two-dimensional spectrum feature of the audio.

[0136] The technical solution provided in this embodiment, by acquiring the speech signal and performing Mel spectrum conversion processing, is conducive to obtaining more accurate two-dimensional spectrum features of the audio, and provides more effective feature input for subsequent voiceprint recognition.

[0137] The following is an application example to illustrate the voiceprint recognition method based on spatial cross-learning multi-scale attention feature module provided by this application. This application example uses the method applied to a terminal as an example.

[0138] The voiceprint recognition involved in this application example is also called speaker recognition. This application example designs a method to enhance speaker recognition performance based on the Efficient Multi-Scale Attention Module with Cross-SpatialLearning (EMA) feature fusion module and convolutional neural network, that is, a multi-scale parallel sub-network is used to extract and spatially cross-fuse the residual output of the network module and the original input in the time domain, frequency domain and global features, and guide the voiceprint network to learn speaker information from multiple scales, so as to improve the performance of speaker recognition across channels and even in general scenarios.

[0139] The following will introduce the focus of this application example - the multi-scale attention feature fusion module of cross-space learning and its application method. The speaker recognition system used in this application example mainly includes two stages: feature extraction and network classification. Figure 2 Shown is a system architecture diagram describing speaker classification learning using a depth-first residual convolutional network, which includes original speech, feature extraction, two-dimensional spectrogram, convolutional layer, downsampling, attention pooling, fully connected layer, and speaker classification prediction.

[0140] In the feature extraction stage, short-time Fourier transform is first used to convert the speech signal from the time domain to the time-frequency domain, with a frame length and frame shift of 25ms (milliseconds) and 10ms respectively. Since the Mel domain is more in line with the auditory characteristics of the human ear, the linear spectrum is converted to an 80-dimensional Mel spectrum and the logarithm is taken, and finally the acoustic features that need to be fed into the backbone network are obtained.

[0141] Then the two-dimensional spectrogram passes through the first convolutional layer to obtain a multi-channel three-dimensional feature map with a shape of [c, f, t], where c represents the number of channels, f is the frequency dimension of the spectrogram (here is 80), and t is the time dimension of the spectrogram. The convolutional layer includes modules such as convolution, batch normalization, and ReLU (rectified linear unit) activation function.

[0142] The residual block of the depth-first residual convolutional neural network includes three convolutional layers, downsampling operations, and residual connections. The general method is to feed the multi-channel feature map into several convolutional residual blocks to extract deep features. Under the goal of ensuring that information is not lost and performance is not degraded, this method improves and updates the first three residual blocks. The structure of a single residual block is as follows: Figure 2 The part within the dotted box is shown.

[0143] Specifically, the improvement method within a residual block is that after the residual block calculates the output, the feature map of each channel of the output result Y is first divided into g groups (g represents the total number of groups), and the dimension of each sub-feature map is [c / / g, f, t]. There are three branch operations for the sub-feature map. The first two are average pooling in the time domain and frequency domain dimensions. The pooled results are concatenated and then passed through a 1×1 convolution layer and an activation function to obtain a weight w 1 And act on the original sub-feature map; the third branch operation is to directly extract the global information x from the sub-feature map through a 3×3 convolutional layer 2 The sub-feature map is denoted as x, and the specific implementation is:

[0144] ,

[0145] ,

[0146] Among them, * represents multiplication, x 1 represents the first output feature map, Sigmoid represents the S-type activation function, Conv 1x1 represents a 1×1 convolution operation, Concat represents a concatenation operation, and Avg f Represents the average pooling along the frequency dimension, Avg t represents the average pooling along the time dimension, Conv 3x3 Represents a 3×3 convolution operation.

[0147] The second stage is spatial cross-learning, x 1 Perform global average pooling, after the Softmax function (normalized exponential function), the same as x 2 Perform matrix multiplication weighting; x 2 Do global average pooling, after Softmax function, the same as x 1 Perform matrix multiplication weighting, and finally add the two weighted results, obtain the final attention weight through the activation function, and weight the original sub-feature map group. The specific implementation formula of this stage is:

[0148] ,

[0149] ,

[0150] ,

[0151] Among them, Softmax represents the normalized exponential function, x 3 represents the third intermediate feature, Avg represents the average operation, GroupNorm represents the group normalization operation, and x 4Represents the fourth intermediate feature, output represents the final output, w represents the weight, Sigmoid represents the S-type activation function, and Matmul represents the matrix multiplication operation.

[0152] Multi-scale attention feature fusion module structure reference based on spatial cross learning Figure 3 , which includes input residual, grouping, time domain average pooling, frequency domain average pooling, 3×3 convolution, splicing + 1×1 convolution, activation function, weighting, normalization, global average pooling, normalized exponential function and matrix multiplication. Improved residual block reference Figure 4 , which includes downsampling, convolutional layers, and multi-scale attention feature fusion modules based on spatial cross-learning. After an improved residual block, compared with the original output, the multi-scale speaker fusion information in the time domain, frequency domain, and global is fused. Generally speaking, there are dozens of residual blocks in a deep convolutional neural network. If each residual block is improved as mentioned above, it will bring a lot of additional computational burden. In addition, due to the characteristics of neural networks, the closer to the classification layer, the more the network pays attention to the detailed information of the audio. In the speaker recognition task, the speaker features are more of an overall feature than the speech content. Performing the above processing at a shallow level can guide the model to pay attention to the speaker information to the maximum extent while minimizing the increase in the number of model parameters.

[0153] After several improved residual blocks and several original residual blocks, the attention pooling layer and the linear layer are used to enhance the expressiveness of the deep features and obtain the speaker representation corresponding to the audio. Finally, the linear classification layer is used in combination with the classification loss function to perform speaker classification training.

[0154] After several improved residual blocks and several original residual blocks, the attention pooling layer and the linear layer are used to enhance the expressiveness of the deep features and obtain the speaker representation corresponding to the audio. Finally, the linear classification layer is used in combination with the classification loss function to perform speaker classification training.

[0155] For example, the process of this application example is as follows:

[0156] First, extract the two-dimensional logarithmic Mel spectrum of the audio as the acoustic feature;

[0157] Second, through the first convolution layer, a multi-channel three-dimensional feature spectrum is obtained;

[0158] Third, the residuals output by the first residual block are grouped by channel layer, and time domain and frequency domain pooling and global convolution are performed on the sub-feature map groups to extract information;

[0159] Fourth, extract the weights of the features after time-frequency domain pooling and weight the original sub-feature map group;

[0160] Fifth, the weighted time-frequency feature map group in the fourth step and the global information feature map group in the third step are spatially cross-learned, and the final weight is obtained by matrix multiplication;

[0161] Sixth, use the final weight to weight the original sub-feature map and integrate it into a new residual, and add it to the residual block input to get the output of the residual block;

[0162] Seventh, repeat steps 3 to 6 several times, and then send them into other subsequent original residual modules of the network to obtain the classification results.

[0163] The technical solution provided by this application example, through multi-level feature recognition and weighted fusion processing, is conducive to making full use of the time domain features, frequency domain features and global features of the audio, thereby improving the expressiveness of the features and improving the accuracy and efficiency of voiceprint recognition.

[0164] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0165] Based on the same inventive concept, the embodiment of the present application also provides a voiceprint recognition device based on a spatial cross-learning multi-scale attention feature module for implementing the voiceprint recognition method based on a spatial cross-learning multi-scale attention feature module involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more embodiments of the voiceprint recognition device based on a spatial cross-learning multi-scale attention feature module provided below can refer to the limitations of the voiceprint recognition method based on a spatial cross-learning multi-scale attention feature module above, and will not be repeated here.

[0166] In an exemplary embodiment, Figure 5 As shown, a voiceprint recognition device based on a spatial cross-learning multi-scale attention feature module is provided. The voiceprint recognition device 500 based on a spatial cross-learning multi-scale attention feature module may include:

[0167] A feature extraction module 501 is used to extract two-dimensional frequency spectrum features of audio;

[0168] A first recognition module 502 is used to perform feature map recognition on the two-dimensional spectrum features through a feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio;

[0169] A feature grouping module 503 is used to group the multi-channel three-dimensional feature map by channel layer through the target residual block to obtain an original sub-feature map group of the audio;

[0170] The second recognition module 504 is used to perform time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group, and obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group;

[0171] The first generating module 505 is used to generate basic weights according to the time domain features and the frequency domain features, and use the basic weights to perform weighted processing on the original sub-feature graph group to obtain a target sub-feature graph group;

[0172] A second generating module 506, configured to generate a target weight according to the target sub-feature graph group and the global feature;

[0173] A feature fusion module 507 is used to perform weighted fusion processing on the original sub-feature map group using the target weight to obtain a fused feature map of the audio;

[0174] The voiceprint recognition module 508 is used to perform voiceprint recognition on the audio according to the fused feature graph to obtain the voiceprint recognition result of the audio.

[0175] In an exemplary embodiment, the device 500 also includes: a third recognition module, which is used to use a residual network to perform feature recognition on the fused feature map to obtain a target feature map; a voiceprint recognition module 508, which is also used to perform voiceprint recognition on the audio according to the target feature map to obtain a voiceprint recognition result.

[0176] In an exemplary embodiment, the voiceprint recognition module 508 is also used to perform feature recognition on the target feature map using the attention pooling layer and the linear layer to obtain the voiceprint features of the target feature map; perform voiceprint recognition on the audio according to the voiceprint features to obtain the voiceprint recognition result.

[0177] In an exemplary embodiment, the second recognition module 504 is also used to perform time domain dimension pooling operation on the original sub-feature graph group to obtain time domain features; perform frequency domain dimension pooling operation on the original sub-feature graph group to obtain frequency domain features; perform global convolution operation on the original sub-feature graph group to obtain global features.

[0178] In an exemplary embodiment, the first generation module 505 is also used to perform splicing processing on the time domain features and the frequency domain features to obtain splicing features; perform convolution operation on the splicing features to obtain convolution operation results; and perform activation processing on the convolution operation results to obtain basic weights.

[0179] In an exemplary embodiment, the feature extraction module 501 is further used to obtain a speech signal as audio; perform Mel spectrum conversion on the audio to obtain a two-dimensional spectrum feature.

[0180] Each module in the above-mentioned voiceprint recognition device based on spatial cross-learning multi-scale attention feature module can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0181] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be realized through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a voiceprint recognition method based on a spatial cross-learning multi-scale attention feature module is realized. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0182] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0183] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0184] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0185] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0186] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0187] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0188] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A voiceprint recognition method based on spatial cross-learning multi-scale attention feature module, characterized in that: The method comprises: Extract two-dimensional spectrum features of audio; Performing feature map recognition on the two-dimensional spectrum features through a feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio; the multi-channel three-dimensional feature map is a three-dimensional feature representation of multiple channels obtained after being processed by the feature map recognition layer, and the multi-channel three-dimensional feature map includes a convolution feature map with multiple filter outputs; Grouping the multi-channel three-dimensional feature map by channel layer through a target residual block to obtain an original sub-feature map group of the audio; Performing time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group respectively to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group; Generate basic weights according to the time domain features and the frequency domain features, and use the basic weights to perform weighted processing on the original sub-feature graph group to obtain a target sub-feature graph group; Generate a target weight according to the target sub-feature map group and the global feature; Performing weighted fusion processing on the original sub-feature graph group using the target weight to obtain a fused feature graph of the audio; According to the fused feature map, voiceprint recognition is performed on the audio to obtain a voiceprint recognition result of the audio.

2. The method according to claim 1, characterized in that: Before performing voiceprint recognition on the audio according to the fused feature graph to obtain the voiceprint recognition result of the audio, the method further includes: Using a residual network to perform feature recognition on the fused feature map to obtain a target feature map; The performing voiceprint recognition on the audio according to the fused feature graph to obtain the voiceprint recognition result of the audio includes: According to the target feature map, voiceprint recognition is performed on the audio to obtain the voiceprint recognition result.

3. The method according to claim 2, characterized in that The performing voiceprint recognition on the audio according to the target feature graph to obtain the voiceprint recognition result includes: Using an attention pooling layer and a linear layer to perform feature recognition on the target feature map to obtain a voiceprint feature of the target feature map; The audio is subjected to voiceprint recognition according to the voiceprint feature to obtain the voiceprint recognition result.

4. The method according to claim 1, characterized in that The performing time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group respectively to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group comprises: Performing a time domain dimension pooling operation on the original sub-feature map group to obtain the time domain feature; Performing frequency domain dimension pooling operation on the original sub-feature map group to obtain the frequency domain feature; A global convolution operation is performed on the original sub-feature map group to obtain the global feature.

5. The method according to claim 1, characterized in that The generating a basic weight according to the time domain feature and the frequency domain feature comprises: Performing splicing processing on the time domain feature and the frequency domain feature to obtain a splicing feature; Performing a convolution operation on the splicing feature to obtain a convolution operation result; An activation process is performed on the convolution operation result to obtain the basic weight.

6. The method according to any one of claims 1 to 5, characterized in that: The extracting of two-dimensional frequency spectrum features of the audio comprises: Acquire a speech signal as the audio; The audio is subjected to a Mel spectrum conversion process to obtain the two-dimensional spectrum feature.

7. A voiceprint recognition device based on spatial cross-learning multi-scale attention feature module, characterized in that: The device comprises: A feature extraction module is used to extract two-dimensional spectrum features of audio; A first recognition module is used to perform feature map recognition on the two-dimensional spectrum feature through a feature map recognition layer to obtain a multi-channel three-dimensional feature map of the audio; the multi-channel three-dimensional feature map is a three-dimensional feature representation of multiple channels obtained after being processed by the feature map recognition layer, and the multi-channel three-dimensional feature map includes a convolution feature map with multiple filter outputs; A feature grouping module, used for grouping the multi-channel three-dimensional feature map by channel layer through a target residual block to obtain an original sub-feature map group of the audio; A second recognition module is used to perform time domain feature recognition, frequency domain feature recognition and global feature recognition on the original sub-feature graph group, respectively, to obtain the time domain feature, frequency domain feature and global feature of the original sub-feature graph group; A first generating module, configured to generate basic weights according to the time domain features and the frequency domain features, and perform weighted processing on the original sub-feature graph group using the basic weights to obtain a target sub-feature graph group; A second generating module, used for generating a target weight according to the target sub-feature graph group and the global feature; A feature fusion module, used to perform weighted fusion processing on the original sub-feature map group using the target weight to obtain a fused feature map of the audio; The voiceprint recognition module is used to perform voiceprint recognition on the audio according to the fused feature map to obtain a voiceprint recognition result of the audio.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Voiceprint feature extraction method and device, electronic equipment and storage medium

    CN111833884A

  • Microphone array-oriented channel attention weighted speech enhancement method

    CN112151059A