Speech recognition method, device, computer-readable storage medium, and computer equipment

By calculating the sparseness value of the speech recognition feature vector and determining the self-attention calculation feature vector, the problem of increasing the computational complexity of the self-attention mechanism is solved, and the efficiency of speech recognition is improved.

CN113823264BActive Publication Date: 2025-06-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110731479.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-30
Publication Date
2025-06-24
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

Under the self-attention mechanism, as the length of the input sequence increases, the computational complexity increases greatly, resulting in low speech recognition efficiency.

Method used

By extracting the speech information to be recognized, the sparsity value of each feature vector is calculated, and the feature vector whose sparsity value is greater than the preset threshold is determined for self-attention calculation. Other feature vectors do not perform self-attention calculation, and the target matrix is ​​generated and input to the classification network for classification processing.

Benefits of technology

This reduces the amount of computing, improves the efficiency of speech recognition, and reduces the use of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113823264B_ABST
    Figure CN113823264B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a speech recognition method, apparatus, computer-readable storage medium, and computer device. The method extracts features from the speech information to be recognized to obtain a plurality of feature vectors; calculates the sparsity value of each feature vector, where the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence; determines a first feature vector with a sparsity value greater than a preset threshold and a second feature vector with a sparsity value not greater than the preset threshold; determines a target matrix according to the self-attention calculation result of the first feature vector and the second feature vector; inputs the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain a recognition result corresponding to the speech information to be recognized. In this way, the present application adopts a deep learning method to reduce the computational complexity of the self-attention mechanism in the speech recognition process, thereby improving the efficiency of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and in particular, to a speech recognition method, apparatus, computer-readable storage medium, and computer device. Background Art

[0002] Automatic Speech Recognition (ASR) technology is a technology that enables a machine to convert a speech signal into a corresponding text or command through a recognition and understanding process. The speech recognition technology mainly includes three aspects: feature extraction technology, pattern matching criterion, and model training technology.

[0003] In recent years, the automatic speech recognition technology has developed rapidly, and its applications have also penetrated into various fields of people's lives. Among them, the End-to-End (E2E) automatic speech recognition technology has been widely favored for its simplified architecture and excellent performance. The transducer and the attention-based codec are two popular E2E frameworks, which can directly convert the input audio stream features into text results, and have certain advantages over traditional speech recognition models in terms of resource consumption and accuracy.

[0004] However, under the self-attention mechanism, as the length of the input sequence increases, the computational complexity will increase significantly, resulting in low speech recognition efficiency. Summary of the Invention

[0005] Embodiments of the present application provide a speech recognition method, apparatus, computer-readable storage medium, and computer device, and the method can improve the efficiency of speech recognition.

[0006] The first aspect of the present application provides a speech recognition method, including:

[0007] Performing feature extraction on the speech information to be recognized to obtain a plurality of feature vectors;

[0008] Calculating the sparsity value of each feature vector, where the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence;

[0009] Determining a first feature vector with a sparsity value greater than a preset threshold and a second feature vector with a sparsity value not greater than the preset threshold;

[0010] Determining a target matrix according to the self-attention calculation result of the first feature vector and the second feature vector;

[0011] Inputting the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain a recognition result corresponding to the speech information to be recognized.

[0012] Correspondingly, a second aspect of the present application provides a voice recognition device, which includes:

[0013] An extraction unit for extracting features from the voice information to be recognized to obtain a plurality of feature vectors;

[0014] A calculation unit for calculating the sparsity value of each feature vector, where the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence;

[0015] A first determination unit for determining a first feature vector with a sparsity value greater than a preset threshold and a second feature vector with a sparsity value not greater than the preset threshold;

[0016] A second determination unit for determining a target matrix according to the self-attention calculation result of the first feature vector and the second feature vector;

[0017] An identification unit for inputting the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain an identification result corresponding to the voice information to be recognized.

[0018] In some embodiments, the calculation unit includes:

[0019] A first calculation subunit for calculating the self-attention score sequence of each feature vector;

[0020] A second calculation subunit for calculating the relative entropy between the distribution of each score sequence and the uniform distribution to obtain the sparsity value of the feature vector corresponding to each score sequence.

[0021] In some embodiments, the calculation unit includes:

[0022] A selection subunit for randomly selecting a target number of feature vectors from the plurality of feature vectors to generate a key matrix;

[0023] A third calculation subunit for calculating the sparsity value of each feature vector according to the target number, the plurality of feature vectors, and the key matrix.

[0024] In some embodiments, the extraction unit includes:

[0025] A division subunit for dividing the voice information to be recognized into multiple frames of voice signals;

[0026] A transformation subunit for performing a discrete Fourier transform on each frame of voice signal to obtain the spectral information corresponding to each frame of voice signal;

[0027] The first processing subunit is configured to perform Mel cepstrum processing on the spectral information corresponding to each frame of voice signal to obtain multiple feature vectors of the voice information to be recognized.

[0028] In some embodiments, the apparatus further includes:

[0029] A noise reduction unit, configured to perform noise reduction processing on the voice information to be recognized;

[0030] A pre-emphasis unit, configured to perform pre-emphasis processing on the voice information after noise reduction processing.

[0031] In some embodiments, the second determination unit includes:

[0032] A first acquisition subunit, configured to acquire a target self-attention score sequence corresponding to each first feature vector;

[0033] A fourth calculation subunit, configured to perform weighted calculation according to the target self-attention score sequence and the value vector corresponding to each first feature vector to obtain a third feature vector corresponding to each first feature vector;

[0034] A determination subunit, configured to determine a target matrix according to the third feature vector and the second feature vector.

[0035] In some embodiments, the first acquisition subunit includes:

[0036] A calculation module, configured to calculate the dot product result of each target feature vector and each feature vector in the first feature vector;

[0037] A processing module, configured to normalize the dot product result to obtain a target self-attention score sequence corresponding to each first feature vector.

[0038] In some embodiments, the recognition unit includes:

[0039] A second acquisition subunit, configured to acquire a label sequence corresponding to the speech recognition result text;

[0040] An extraction subunit, configured to extract features from the label sequence to obtain a label feature vector corresponding to the label sequence;

[0041] A second processing subunit, configured to process the label feature vector by using an artificial neural network to obtain a feature matrix corresponding to the label sequence;

[0042] A recognition subunit, configured to combine the target matrix and the feature matrix by using a multi-layer fully connected layer, and input the combination result into a classification network for decoding to obtain a recognition result corresponding to the voice information to be recognized.

[0043] In a third aspect of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the voice recognition method provided in the first aspect of the present application.

[0044] In a fourth aspect of the present application, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the voice recognition method provided in the first aspect of the present application are implemented.

[0045] In a fifth aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a storage medium. A processor of a computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the voice recognition method provided in the first aspect.

[0046] The voice recognition method provided in the embodiments of the present application extracts features from the voice information to be recognized to obtain multiple feature vectors; calculates the sparsity value of each feature vector, where the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence; determines a first feature vector with a sparsity value greater than a preset threshold and a second feature vector with a sparsity value not greater than the preset threshold; determines a target matrix according to the self-attention calculation result of the first feature vector and the second feature vector; inputs the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain a recognition result corresponding to the voice information to be recognized. In this way, by calculating the sparsity value of the feature vector and then determining that the feature vector with a sparsity value less than the preset threshold does not need to perform self-attention calculation, the amount of calculation can be reduced, thereby improving the efficiency of voice recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 It is a schematic diagram of the scenario of the voice recognition method provided in the present application;

[0049] Figure 2 It is a schematic flow diagram of the voice recognition method provided in the present application;

[0050] Figure 3Schematic diagram of the automatic speech recognition framework for the Transducer model;

[0051] Figure 4 Another flowchart of the speech recognition method provided by this application;

[0052] Figure 5 Performance comparison chart of the speech recognition method of this application with the baseline model.

[0053] Figure 6 Schematic diagram of the structure of the speech recognition device provided by this application;

[0054] Figure 7 Schematic diagram of the structure of the computer device provided by this application. Specific implementation manners

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] The embodiments of the present invention provide a speech recognition method, device, computer-readable storage medium, and computer device. Among them, the speech recognition method can be used in a speech recognition device. The speech recognition device can be integrated in a computer device, and the computer device can be a terminal or a server. Among them, the terminal includes but is not limited to mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0057] Such as Figure 1As shown, it is a schematic diagram of the scenario of the speech recognition method provided by this application. As shown in the figure, after receiving the speech information to be recognized, the computer device extracts features from the speech information to be recognized to obtain multiple feature vectors corresponding to the speech information to be recognized; then, the computer device calculates the sparsity of each feature vector to obtain the sparsity value of each feature vector. Further, the feature vectors are distinguished according to the sparsity values of each feature vector, and the feature vectors with sparsity values greater than the preset threshold are determined as the first feature vectors, and the feature vectors with sparsity values not greater than the preset threshold are determined as the second feature vectors. For the first feature vectors with sparsity values greater than the preset threshold, self-attention calculation is performed on them; for the second feature vectors with sparsity values not greater than the preset threshold, self-attention calculation is not performed on them. Then, the target matrix is determined according to the self-attention calculation results of the first feature vectors and the second feature vectors. Finally, the target matrix and the feature matrix corresponding to the label sequence are input into the classification network for classification processing to determine the output label, and further, the recognition result corresponding to the speech information to be recognized is determined according to the output label.

[0058] It should be noted that Figure 1 The schematic diagram of the speech recognition scenario shown is only an example. The speech recognition scenario described in the embodiments of this application is to more clearly illustrate the technical solution of this application, and does not constitute a limitation on the technical solution provided by this application. Those of ordinary skill in the art know that with the evolution of speech recognition and the emergence of new business scenarios, the technical solution provided by this application is equally applicable to similar technical problems.

[0059] Based on the above implementation scenarios, the following will be described in detail respectively.

[0060] The embodiments of this application will be described from the perspective of a speech recognition device, which can be integrated in a computer device. Among them, the computer device can be a terminal or a server, and the terminal includes but is not limited to mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc. As Figure 2 As shown, it is a flowchart of the speech recognition method provided by this application, and the method includes:

[0061] Step 101: Extract features from the speech information to be recognized to obtain multiple feature vectors.

[0062] Among them, with the continuous development of speech recognition technology, using deep learning technology for speech recognition has greatly improved the accuracy of speech recognition. Deep Learning (DL) is a new research direction in the field of Machine Learning (ML). Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analysis and learning like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies.

[0063] Among the numerous solutions for speech recognition using deep learning technology, the E2E automatic speech recognition technology is widely favored for its simplified architecture and excellent performance. The Transducer and the attention-based codec are two relatively commonly used E2E automatic speech recognition frameworks. The Transducer can directly convert the features of the input audio stream into text results, and it has great advantages over traditional speech recognition models in terms of resource consumption and accuracy.

[0064] To introduce the technical solution of this application more clearly, the framework of the Transducer model is briefly introduced below.

[0065] As Figure 3As shown in the figure, it is a schematic diagram of the automatic speech recognition framework of the Transducer model. This framework includes an encoder 10 (Encoder), a prediction network 20 (Prediction Network), a fully connected layer 31 (Joint Network), and a classification layer 32 (Softmax Layer). In the encoder 10, first, through the downsampling and position embedding layer 12, the speech feature 41 is mapped into the vector space to obtain the feature vector corresponding to the speech information to be recognized. Then, the obtained feature vector passes through the M-layer Convolution-Augmented Transformer (Conformer) layer 11, and the output matrix of the encoder is output. Among them, compared with the recurrent neural network with Long Short-Term Memory (LSTM), the Transformer Block model has higher accuracy and efficient computing power because the Transformer has the ability to model longer global contexts. However, the Transformer has poor ability to capture local information, and this ability is necessary for speech recognition. In order to balance the ability to capture local information and the ability to obtain global information, the Conformer model is proposed. The Conformer model combines convolution and self-attention mechanisms, uses convolution to capture local information, and uses the self-attention mechanism to process global information. The Conformer model can obtain better speech recognition accuracy and recognition efficiency compared with the Transformer Block model. Therefore, the M-layer Conformer model can be adopted in the encoder of the Transducer model to improve the speech recognition ability.

[0066] In the prediction network 20, first, the label sequence is mapped into the vector space through the embedding layer 21 to obtain the feature vector corresponding to the label sequence, and then it is processed through the N-layer recurrent neural network layer 22 to output the output result of the prediction network.

[0067] Then, the output matrix of the encoder 10 and the output result of the prediction network 20 can be input into the fully connected layer 31 for combination. Finally, the combined result is decoded through the classification layer 32 to obtain the predicted label 43. After obtaining the predicted label, the predicted label is added to the label sequence 42, and then further speech recognition is performed until all labels are recognized to obtain the speech recognition result.

[0068] In the above speech recognition framework, since the Conformer model combines convolution and self-attention mechanisms. However, under the self-attention mechanism, each output is a weighted combination of the entire feature vector sequence. Thus, when the speech information to be recognized is long, resulting in a large number of input feature vectors, the computational complexity of the self-attention mechanism will increase exponentially, leading to the consumption of a large amount of computing resources and a decrease in speech recognition efficiency. To solve the problem of low long speech recognition efficiency caused by the self-attention mechanism in the above Conformer model, this application proposes a speech recognition method, which will be described in detail below.

[0069] In the embodiment of this application, the proposed speech recognition method is still based on the above-mentioned speech recognition framework based on Transducer. Therefore, after receiving the speech information to be recognized, it is still necessary to first extract the features of the speech information and map the extracted features to the vector space to obtain a plurality of feature vectors.

[0070] In some embodiments, extracting features from the speech information to be recognized to obtain a plurality of feature vectors includes:

[0071] 1. Divide the speech information to be recognized into multiple frames of speech signals;

[0072] 2. Perform a discrete Fourier transform on each frame of speech signal to obtain the spectral information corresponding to each frame of speech signal;

[0073] 3. Perform Mel cepstrum processing on the spectral information corresponding to each frame of speech signal to obtain a plurality of feature vectors of the speech information to be recognized.

[0074] Among them, since the received speech information is a non-stationary and time-varying signal, however, within a short time range, it can be considered that the signal is stationary and time-invariant, and this short time is generally 10 - 30 ms. Therefore, when performing speech recognition, to reduce the influence of the overall non-stationarity and time-variation of the speech signal, it is necessary to segment the speech signal. Each segment is called a frame, and the frame length can be taken as 25 ms. Further, in order to make the frames transition smoothly and maintain their continuity, an overlapping segmentation method can be adopted to ensure that adjacent two frames overlap partially. The time difference between the starting positions of adjacent two frames is called the frame shift, and the frame shift can be taken as 10 ms.

[0075] After dividing the speech information to be recognized into multiple frames of speech signals, perform a discrete Fourier transform (DFT) on each frame of speech signal obtained by the division, and convert the speech signal obtained after the division from a time-domain signal to a frequency-domain signal. In some embodiments, the fast Fourier transform (FFT) can be used to reduce the time complexity of the calculation, thereby further improving the speech recognition efficiency.

[0076] In some embodiments, before performing the fast Fourier transform on the multiple frames of speech signals obtained by the division, the multiple frames of speech signals obtained by the division can also be windowed. Among them, windowing is to process the multiple frames of speech signals obtained by the division using a window function, or a weighting function. Since the requirement of the fast Fourier transform is that the signal is a periodic signal. And the speech signal obtained after the division is non-periodic. Using a non-periodic speech signal for the fast Fourier transform will cause the occurrence of the frequency leakage problem. Therefore, in order to minimize this leakage error, it is necessary to use a weighting function, or a window function to process the speech signals obtained by the division. Among them, frequency leakage means that there are frequency components that did not originally exist in the analysis result. For example, for a pure sine wave of 50 Hz (hertz), there is originally only one frequency component, but the analysis result contains other frequency components close to the 50 Hz frequency.

[0077] After performing the discrete Fourier transform on the speech signals obtained by the division to obtain the spectral information corresponding to each frame of speech signal, the spectral information corresponding to each frame of speech signal can be further subjected to mel cepstrum processing. Among them, mel cepstrum processing includes converting the spectral information from a frequency scale to a mel scale and cepstrum processing. Since the result obtained by the discrete Fourier transform or the fast Fourier transform is the amplitude on each frequency band, and humans have different perception abilities for speech of different frequencies. Specifically, below 1 kHz (kilohertz), the perception ability is linearly related to the frequency; above 1 kHz, the perception ability is logarithmically related to the frequency. And the higher the frequency, the worse the perception ability. The mel scale is a non-linear scale unit that represents the human ear's perception of pitch changes and is defined based on frequency. In the mel frequency domain, the human perception ability is linearly related to the frequency. If the mel frequencies of two pieces of speech differ by two times, then the human perception also differs by two times. After converting the spectral information of each frame of speech signal to the spectral information on the mel scale, perform cepstrum processing on the spectral information on the mel scale to obtain the mel cepstrum coefficients corresponding to each frame of speech signal, and obtain the speech feature parameters of the speech information to be recognized. Further, map the obtained feature parameters to a vector space to obtain multiple feature vectors corresponding to the speech information to be recognized.

[0078] In some embodiments, before dividing the speech information to be recognized into multiple frames of speech signals, the following steps are further included:

[0079] A. Perform noise reduction processing on the speech information to be recognized;

[0080] B. Perform pre-emphasis processing on the speech information after noise reduction processing.

[0081] Among them, before performing feature extraction on the speech information to be recognized, noise reduction processing can be performed on the speech information to be recognized first. Specifically, when performing speech recognition on the speech information collected in a relatively harsh environment, noise reduction processing can be performed first. For the speech information collected in an environment with weak noise and clean speech, noise reduction processing may not be required because the speech recognition technology based on deep learning itself has strong anti-noise ability.

[0082] Furthermore, after performing noise reduction processing on the speech information to be recognized, pre-emphasis processing can be further performed on the speech information to be recognized. Since in the audio recording process, high-frequency signals are more likely to attenuate, and the pronunciation of some factors such as vowels contains more components of high-frequency signals. The loss of high-frequency signals will cause the formants of the factors to be not obvious, and then lead to the weak modeling ability of the acoustic model for these factors. Pre-emphasis is a first-order high-pass filter, which can increase the energy of the high-frequency part of the signal, thereby reducing the attenuation of high-frequency signals, and then improving the accuracy of speech recognition.

[0083] Step 102: Calculate the sparsity value of each feature vector.

[0084] Among them, after extracting a plurality of feature vectors from the speech information to be recognized, the sparsity value of each feature vector is calculated. Among them, the sparsity value corresponding to each feature vector is the relative entropy between the actual distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence. Among them, relative entropy (RE), also known as Kullback-Leibler divergence or information divergence, is a measure of the asymmetry of the difference between two probability distributions. In the embodiments of the present application, the sparsity value of each feature vector is the measure of the asymmetry of the difference between the actual distribution and the uniform distribution of the self-attention score sequence of each feature vector. Specifically, for example, four feature vectors are extracted from the speech information to be recognized, and the actual distribution of the self-attention scores of a certain feature vector is 0.4, 0.3, 0.2, and 0.1. Then the sparsity value of this feature vector is the relative entropy between the sequence {0.4, 0.3, 0.2, 0.1} and the uniform distribution sequence {0.25, 0.25, 0.25, 0.25}. To solve the relative entropy between two sequences, the relative entropy calculation formula can be used for specific calculation, which will be introduced in detail in the following text of the present application.

[0085] In some embodiments, calculating the sparsity value of each feature vector includes:

[0086] 1. Calculate the self-attention score sequence of each feature vector;

[0087] 2. Calculate the relative entropy between the distribution of each score sequence and the uniform distribution to obtain the sparsity value of the feature vector corresponding to each score sequence.

[0088] Among them, the multiple feature vectors extracted from the speech information to be recognized can form an input matrix. Then, a linear transformation projection is performed on the input matrix to obtain a query matrix, a key matrix, and a value matrix corresponding to the input matrix respectively. It can be understood that the query matrix contains multiple query vectors, the key matrix contains multiple key vectors, and the value matrix also contains multiple value vectors. Calculating the self-attention score sequence corresponding to each feature vector can be calculating the self-attention score sequence corresponding to each query vector. And calculating the self-attention score sequence corresponding to each query vector can multiply the query vector by each key vector to obtain a plurality of dot product results. Then, the obtained plurality of dot product results are normalized to obtain the self-attention score corresponding to each dot product result. These self-attention scores constitute the self-attention score sequence of the query vector. Then, each query vector is traversed to obtain the self-attention score sequence corresponding to each query vector, that is, the self-attention score sequence corresponding to each feature vector is obtained.

[0089] After calculating the self-attention score sequence corresponding to each feature vector, calculate the relative entropy between the distribution of the self-attention score sequence corresponding to each feature vector and the uniform distribution, so as to obtain the sparsity value corresponding to each feature vector.

[0090] Step 103, determine the first feature vectors with sparsity values greater than the preset threshold and the second feature vectors with sparsity values not greater than the preset threshold.

[0091] Among them, due to the certain sparsity of the feature matrix, it means that it is not necessary to perform self-attention calculations on all feature vectors. Generally, if the distribution of the self-attention scores of the feature vectors follows a uniform distribution, the output of the self-attention mechanism degenerates into the average value of all feature vectors, and thus the attention ability is lost. Therefore, only the feature vectors whose self-attention score sequence distributions are far from the uniform distribution need to perform self-attention calculations.

[0092] In this way, a sparsity threshold can be set. When the sparsity value corresponding to a feature vector is greater than this sparsity threshold, it is determined that this feature vector is the first feature vector that needs to perform self-attention calculations; when the sparsity value corresponding to a feature vector is not greater than this sparsity threshold, it is determined that this feature vector is the second feature vector that does not need to perform self-attention calculations.

[0093] Step 104, determine the target matrix according to the self-attention calculation results of the first feature vectors and the second feature vectors.

[0094] Among them, after determining the first feature vectors that need to perform self-attention calculations and the second feature vectors that do not need to perform self-attention calculations, perform self-attention calculations on the first feature vectors to obtain the self-attention calculation results corresponding to each first feature vector, where the self-attention calculation result corresponding to each first feature vector is also a vector. Then, generate the target matrix according to the vectors obtained by the self-attention calculations of each first feature vector and the second feature vectors. This target matrix is the output of the Conformer in the encoder of the Transducer model. Then, this output can be continuously input into the next Conformer layer for processing until the final encoder output target matrix is obtained after being processed by multiple Conformer layers.

[0095] In some embodiments, determining the target matrix according to the self-attention calculation results of the first feature vectors and the second feature vectors includes:

[0096] 1. Obtain the target self-attention score sequence corresponding to each first feature vector;

[0097] 2. Perform weighted calculation based on the target self-attention score sequence and the value vector corresponding to each first feature vector to obtain the third feature vector corresponding to each first feature vector;

[0098] 3. Determine the target matrix according to the third feature vector and the second feature vector.

[0099] Among them, for the first feature vectors that need to perform self-attention calculation, self-attention calculation can be performed by first obtaining the target self-attention score sequence corresponding to each first feature vector. Specifically, as described in the foregoing step 102, a feature matrix composed of multiple feature vectors extracted from the speech information to be recognized can be linearly mapped to obtain a query matrix, a key matrix, and a value matrix, and then the target query vector corresponding to the first feature vector can be determined from the query matrix according to the first feature vector. Then, calculate the dot product result of each target query vector and each key vector in the key matrix one by one, and determine the self-attention score sequence of each target query vector according to the dot product result, that is, obtain the target self-attention score sequence corresponding to each first feature vector. Then, perform weighted processing on the value vectors in the value matrix according to the target self-attention score sequence corresponding to each first feature vector to obtain the third feature vector corresponding to each first feature vector. Finally, determine the target matrix according to the third feature vector and the second feature vector, where the target matrix is the matrix obtained by replacing the first feature vector in the input matrix with the third feature vector corresponding to each first feature vector.

[0100] In some embodiments, obtaining the target self-attention score sequence corresponding to each first feature vector includes:

[0101] A. Calculate the dot product result of each target feature vector in the first feature vector and each feature vector;

[0102] B. Normalize the dot product result to obtain the target self-attention score sequence corresponding to each first feature vector.

[0103] Among them, in the embodiments of the present application, obtaining the target self-attention score sequence corresponding to each first feature vector may be to calculate the dot product result of a target feature vector in the first feature vector and each feature vector extracted from the speech information to be recognized to obtain a dot product result sequence. Then, normalize the calculated dot product result sequence to obtain the self-attention score sequence corresponding to the target feature vector. Then, use the same method to calculate the self-attention score sequence corresponding to each feature vector in the first feature vector one by one to obtain the target self-attention score sequence corresponding to each first feature vector.

[0104] Step 105: Input the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain the recognition result corresponding to the speech information to be recognized.

[0105] After obtaining the target matrix output by the encoder, input the target matrix output by the encoder and the feature matrix corresponding to the label sequence processed according to the label sequence in the prediction network into the classification network for classification processing. The classification network includes a fully connected layer and a classification layer. In the fully connected layer, first combine the target matrix and the feature matrix corresponding to the label sequence, and then input the combined result into the classification layer for classification to obtain a predicted label. Then, add the predicted label to the label sequence for further recognition, and so on until all label texts are recognized, thereby completing the recognition process of the speech to be recognized and obtaining the recognition result corresponding to the speech information to be recognized.

[0106] In some embodiments, inputting the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain the recognition result corresponding to the speech information to be recognized includes:

[0107] 1. Obtain the label sequence corresponding to the speech recognition result text;

[0108] 2. Extract features from the label sequence to obtain the label feature vector corresponding to the label sequence;

[0109] 3. Process the label feature vector using an artificial neural network to obtain the feature matrix corresponding to the label sequence;

[0110] 4. Use a multi-layer fully connected layer to combine the target matrix and the feature matrix, and input the combined result into the classification network for decoding to obtain the recognition result corresponding to the speech information to be recognized.

[0111] After obtaining the target matrix obtained by encoding the speech information to be recognized by the encoder, obtain the already recognized label sequence. The initial value of the label sequence can be set to 0, and then after each prediction label is output by the classification layer, update the output prediction label to the label sequence to form a new label sequence, and then generate subsequent prediction labels according to the new label sequence. After obtaining the currently recognized label sequence, extract features from the label sequence and map the extracted features into a vector space to obtain the feature vector corresponding to the label sequence. Then, use multiple recurrent neural networks containing long short-term memory units to process the feature vector corresponding to the label sequence to obtain the feature matrix corresponding to the label sequence.

[0112] Then, the target matrix and the feature matrix corresponding to the label sequence are input into the fully connected layer for combination, and the result obtained by the combination is further input into the classification network for decoding, so as to obtain the predicted label. This cycle is repeated until all predicted labels are obtained, thereby completing the recognition of the speech information to be recognized.

[0113] According to the above description, it can be known that the speech recognition method provided by the embodiment of the present application extracts features from the speech information to be recognized to obtain a plurality of feature vectors; calculates the sparsity value of each feature vector, and the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence; determines the first feature vector with a sparsity value greater than the preset threshold and the second feature vector with a sparsity value not greater than the preset threshold; determines the target matrix according to the self-attention calculation result of the first feature vector and the second feature vector; inputs the target matrix and the feature matrix corresponding to the label sequence into the classification network for classification processing to obtain the recognition result corresponding to the speech information to be recognized. In this way, by calculating the sparsity value of the feature vector and then determining that the feature vector with a sparsity value less than the preset threshold does not need to perform self-attention calculation, the amount of calculation can be reduced, thereby improving the efficiency of speech recognition.

[0114] Correspondingly, the embodiment of the present application will further describe in detail the speech recognition method provided by the present application from the perspective of a computer device, where the computer device can be a terminal or a server. Among them, the terminal includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, etc. As Figure 4 shown, it is another schematic flow chart of the speech recognition method provided by the present application, and the method includes:

[0115] Step 201, the computer device extracts features from the speech information to be recognized and maps the extracted feature information to a vector space to obtain a feature matrix corresponding to the speech information to be recognized.

[0116] Among them, to more clearly elaborate the technical solution of the present application, the Transducer model can be further described in detail first. The Transducer model can directly model the relationship between the input speech feature x and the text sequence y given the input speech feature x. Please continue to refer to Figure 3, in the Transducer automatic speech recognition framework, the speech feature x is used as the input of the encoder. Here, the speech feature can be the feature extracted from the speech information to be recognized. To extract the speech feature from the speech information to be recognized, the method of extracting the Mel cepstral coefficients of the speech information to be recognized can be used for feature extraction. After extracting the speech feature x from the speech information to be recognized, the extracted speech feature is input into the encoder for downsampling and embedding processing, so as to map the speech feature into the vector space and obtain the feature vector. Then, the obtained feature vector is processed through multiple layers of Conformer to obtain a high-dimensional representation of the input speech feature x:

[0117] h enc = Encoder(x)

[0118] where h enc represents the output of the encoder, and Encoder(x) represents the encoding operation on the input speech feature x. Among them, the encoding operation includes the above embedding operation and multiple layers of Conformer operations.

[0119] The role of the prediction network in the Transducer model is to obtain a high-dimensional representation of the historical decoding result, where the historical decoding result is the prediction label output by the classification layer in the previous time. The prediction network usually consists of an embedding layer and multiple recurrent neural network layers containing long short-term memory units. The high-dimensional representation of the historical decoding result is:

[0120]

[0121] where is the output of the u-th time of the prediction network, and Prediction(y u-1 ) is the prediction operation on the prediction label y u-1 obtained by the (u - 1)-th prediction. Among them, the prediction operation includes the above embedding operation and processing through multiple layers of recurrent neural network layers.

[0122] The connection network contains multiple fully connected layers. The role of the fully connected layer is to combine the outputs of the encoder and the prediction network, which can be specifically expressed as follows:

[0123]

[0124] where h t,u is the result of combining the outputs of the encoder and the prediction network, W joint is the combination weight coefficient, U and V are the mapping matrices for combining the outputs of the encoder and the prediction network, and b is the bias coefficient.

[0125] Finally, the result output by the connection network is classified through a classification layer, and then decoded to obtain the final classification result. It is expressed as follows:

[0126] P(k|t, u) = (h t,u )

[0127] Among them, P(k|t, u) represents the classification result of the classification layer, and (h t,u ) represents performing Softmax processing on the output of the connection network. Among them, in the entire Transducer model, the forward-backward algorithm can be used to optimize the posterior probability distribution.

[0128] In the Transducer model framework provided by this application, the Conformer structure is adopted in the encoder to process the feature vectors, and the Conformer module contains four parts: a macaron-style feed-forward fully connected module (Feed-Forward Network, FFN), a multi-head self-attention module (Multi-Head Self Attention, MHSA), a convolution module (Convolution, CONV), and a second macaron-style feed-forward fully connected module.

[0129] Among them, when the MHSA in the Conformer module performs self-attention calculation on the input matrix X, the input matrix X is first projected through a linear transformation to obtain a query matrix Q, a key matrix K, and a value matrix V. Among them, Q = XW q , K = XW k , V = XW v . W q , W k , and W v are all projection matrices. Then, the dot-product form attention mechanism represented by a ratio can be expressed as:

[0130]

[0131] Among them, KT represents the transpose matrix of the key matrix K, and d is the dimension of the input matrix X.

[0132] To better describe the technical problems of the related technology and the effects of this application, the above formula is marked in its vector form:

[0133]

[0134] Among them, p(k j |q i ) is the attention score of the i-th query vector to the j-th key vector, and L is the sequence length of the input matrix X. Then, further, the query vector q iThe self-attention output regarding the key matrix K can be expressed as:

[0135]

[0136] For the above formula, it can be explained in detail that for each query vector q, the dot product of this vector and each key vector needs to be calculated. Then, the dot products of this vector and each key vector are normalized to obtain the self-attention score sequence for each query vector. Then, the value vectors are weighted and summed using this self-attention score sequence to obtain the representation of this query vector using the value vectors. Then, each query vector in the query matrix is traversed to obtain the representation of each query vector using the value vectors, and thus the self-attention output for the input matrix X is obtained. As can be seen from the above, the time complexity of the entire self-attention calculation process is O(L 2 ), where L is the sequence length of the input matrix X. In this way, when the sequence length of the input matrix increases, the time complexity of the self-attention calculation process will increase exponentially, resulting in a large consumption of computing resources and a decrease in the efficiency of speech recognition. To solve the above problems, the present application provides a speech processing method to optimize the calculation of the self-attention module in the Conformer module to improve the calculation speed and reduce the consumption of computing resources. The speech recognition method provided by the present application will be described in detail below.

[0137] In an embodiment of the present application, after receiving the speech information to be recognized, it is still necessary to extract the features of the speech information to be recognized, and then input the extracted speech features into the downsampling and embedding layer to obtain the feature matrix X corresponding to the speech information to be recognized. This feature matrix X is also the input matrix of the Conformer module.

[0138] Step 202, the computer device performs a linear projection transformation on the feature matrix to obtain a query matrix, a key matrix, and a value matrix corresponding to the feature matrix.

[0139] Among them, in the speech recognition method provided in the embodiment of the present application, it is still necessary to perform a linear projection transformation on the input matrix X to obtain a query matrix Q, a key matrix K, and a value matrix V corresponding to the input matrix X. Then, the query matrix Q, the key matrix K, and the value matrix V are respectively represented in vector form to obtain multiple query vectors q corresponding to the query matrix Q, multiple key vectors k corresponding to the key matrix K, and multiple value vectors v corresponding to the value matrix.

[0140] Step 203, the computer device calculates the self-attention score sequence for each query vector.

[0141] Among them, due to the certain sparsity of the query matrix, it means that it is not necessary to perform self-attention calculations on all query vectors to obtain the self-attention outputs of all query vectors. Generally speaking, if the distribution of the self-attention scores of the query vectors follows a uniform distribution, the output of the self-attention mechanism degenerates into the average value of all value vectors, thus losing the attention ability. Therefore, only when the distribution of the self-attention scores of the query vectors relative to the key matrix K is far from the uniform distribution, this query vector is valid and self-attention calculation is required. Thus, after obtaining multiple query vectors q corresponding to the query matrix Q corresponding to the input matrix, multiple key vectors k corresponding to the key matrix K, and multiple value vectors v corresponding to the value matrix, the self-attention score sequence of each query vector q can be calculated first, and then it can be determined whether this query vector needs to perform self-attention calculation according to the self-attention score sequence of each query vector.

[0142] Specifically, to calculate the self-attention score sequence of each query vector q, the dot product of the query vector q and each key vector can be calculated to obtain a dot product sequence. Then, the obtained dot product sequence is normalized to obtain the self-attention score of the query vector relative to each key vector, that is, the self-attention score sequence of the query vector corresponding to the key matrix K is obtained. Then, in this way, the self-attention score sequence of each query vector is calculated one by one to obtain the self-attention score sequence of each query vector for the key matrix K.

[0143] Step 204, the computer device calculates the sparsity value of each query vector.

[0144] Among them, after calculating the self-attention score sequence of each query vector, the sparsity value of each query vector can be further calculated according to the self-attention score sequence of each query vector. Among them, the sparsity value of each query vector is the KL divergence between the distribution of the self-attention score sequence of the query vector and the uniform distribution, or the relative entropy between the distribution of the self-attention score sequence of the query vector and the uniform distribution. Among them, the query vector q i The formula for the KL divergence between the actual distribution P of the self-attention score sequence of and the uniform distribution U is as follows:

[0145]

[0146] Among them, KL(P||U) is the KL divergence between the actual distribution P of the self-attention score sequence of the query vector q i and the uniform distribution U. Similarly, among them, L is the sequence length of the input matrix X, and d is the dimension of the input matrix X. k j is the jth key vector in the key matrix K.

[0147] Then, further, the ith query vector q can be obtainedi The expression for the KL divergence between the actual distribution and the uniform distribution of the self-attention score sequence with respect to the key matrix K is:

[0148]

[0149] where M sparse (q i , K) is the sparsity value of the i-th query vector q i , and L K is the sequence length of the key vector K.

[0150] In some embodiments, calculating the sparsity value of each feature vector includes:

[0151] 1. Randomly select a target number of feature vectors from multiple feature vectors to generate a key matrix;

[0152] 2. Calculate the sparsity value of each feature vector based on the target number, multiple feature vectors, and the key matrix.

[0153] Among them, when calculating the sparsity value of each query vector, it is still necessary to first calculate the dot product of the query vector and each key vector, and then normalize the dot product result to obtain the self-attention score sequence of each query vector. This process still consumes a large amount of computational effort. To further reduce the computational effort in this part, this application proposes to approximately express the sparsity value calculation formula of the query vector using a sampling method as follows:

[0154]

[0155] where is the approximate sparsity value of the query vector q i , is a new key matrix composed of randomly sampling a certain number of key vectors from the key matrix K. is the sampling number. Specifically, r sample is the sampling degree. The sampling degree is a constant, which is used to control the number of sampled samples. k j is the j-th key vector in the newly sampled key matrix.

[0156] According to the above formula, in the embodiments of the present application, only a certain number of feature vectors are sampled from the multiple feature vectors extracted from the speech information to be recognized to form a feature matrix, and then the feature matrix is linearly transformed to obtain a key matrix. Based on this key matrix, the sampling number, and the query vectors corresponding to the multiple feature vectors extracted from the speech information to be recognized, the sparsity value of each query vector can be directly calculated. In this way, only the dot product of the query vector and the sampled key vector needs to be calculated, without calculating the dot product of the query vector and each key vector, further reducing the amount of calculation and improving the calculation efficiency and speech recognition efficiency.

[0157] Step 205, the computer device determines the target query vectors that need to perform self-attention calculation according to the sparsity value of each query vector.

[0158] Among them, after calculating the sparsity value of each query vector, according to a preset sparsity threshold, the query vectors with sparsity values greater than the sparsity threshold can be determined as the target query vectors that need to perform self-attention calculation, while the query vectors with sparsity values not greater than the sparsity threshold do not need to perform self-attention calculation.

[0159] In some embodiments, a preset number can also be set, and then the preset number of query vectors with higher sparsity values are determined as the target query vectors that need to perform self-attention calculation. Specifically, a sparsity rate r sparse can be set, where r sparse < 1, and then according to the sequence length L of the input matrix and the sparsity rate r sparse the number L sparse of target query vectors is calculated as L sparse = r sparse L. Then, the L

[0160] query vectors with higher sparsity values are determined as the target query vectors that need to perform self-attention calculation.

[0161] Among them, after determining the target query vectors that need to perform self-attention calculation according to the sparsity value of each query vector, self-attention calculation can be further performed on the target query vectors, while for other query vectors outside the target query vectors, self-attention calculation does not need to be performed. Specifically, the output of the Conformer model can be calculated according to the following formula:

[0162]

[0163] where I sparse is the set of serial numbers of the target query vectors that need to perform self-attention calculation.

[0164] According to the above formula, for the target query vector that needs to perform self-attention calculation, calculate the dot product of the target query vector and each key vector respectively, and then normalize these dot product results to obtain the self-attention scores of the target query vector and each key vector, obtaining a sequence of self-attention scores. Finally, weight the value vectors according to the sequence of self-attention scores to obtain the output vector corresponding to each target query vector. For the query vector that does not need to perform self-attention calculation, directly determine its corresponding value vector as the output vector. Among them, the output matrix of the Conformer model is the matrix determined by the above output vector.

[0165] Step 207, the computer device repeats the output matrix through multiple Conformer modules for processing and outputs a target matrix.

[0166] Among them, the input matrix X is processed by the first Conformer module to obtain an output matrix, and then the output matrix of the first Conformer module can be used as the input matrix of the next Conformer module for further processing.

[0167] In the embodiment of the present application, in order to avoid calculating the sparsity value of the query vector of the input matrix in each Conformer layer, a method of sharing sparsity between layers can be adopted. For example, if there are a total of M Conformer layers in the encoder, then it can be set to calculate the sparsity value of the query vector every N layers. In the N - 1 layers after the calculation, the sparsity value of the query vector in this layer can be adopted. For example, after calculating the sparsity value of each query vector in the first Conformer layer, the sparsity value of each query vector in the 2nd to Nth Conformer layers can directly adopt the sparsity value of the query vector in the first Conformer layer. In this way, the calculation amount can be further reduced, thereby improving the efficiency of speech recognition. The input matrix X is output as a target matrix after being processed by multiple Conformer models in the encoder.

[0168] Step 208, the computer device inputs the target matrix and the output matrix of the prediction network into the connection network, and processes the output of the connection network through the classification layer to obtain a prediction label.

[0169] Among them, after obtaining the target matrix output by the encoder, the target matrix output by the encoder and the feature matrix output by the prediction network are input into the connection network for combination, and then the combined feature matrix is input into the Softmax classification layer for classification prediction to obtain a prediction label.

[0170] Step 209, the computer device determines the recognition result of the speech information to be recognized according to the prediction label.

[0171] Among them, after the classification layer outputs a predicted label, the predicted label can be decoded to obtain an identification result. After obtaining the identification result, the label sequence can be updated according to the identification result, and the updated label sequence can be re-input into the prediction network for processing to further perform speech recognition on the speech information to be recognized. This cycle continues until all the speech information to be recognized is completely recognized, and the speech recognition result corresponding to the speech information to be recognized is obtained.

[0172] As Figure 5 shown, it is a performance comparison diagram of the speech recognition method provided by this application and the existing speech recognition method. As shown in the figure, the experimental data uses the same test data set. The time consumed and the memory occupied by using the speech recognition method provided by this application for speech recognition of the same test data set are both reduced by more than 40% compared with the existing speech recognition method.

[0173] According to the above description, it can be known that the speech recognition method provided by the embodiments of this application extracts features from the speech information to be recognized to obtain multiple feature vectors; calculates the sparsity value of each feature vector, where the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence; determines the first feature vector with a sparsity value greater than the preset threshold and the second feature vector with a sparsity value not greater than the preset threshold; determines the target matrix according to the self-attention calculation result of the first feature vector and the second feature vector; inputs the target matrix and the feature matrix corresponding to the label sequence into the classification network for classification processing to obtain the recognition result corresponding to the speech information to be recognized. In this way, by calculating the sparsity value of the feature vector and then determining that the feature vector with a sparsity value less than the preset threshold does not need to perform self-attention calculation, the amount of calculation can be reduced, thereby improving the efficiency of speech recognition.

[0174] To better implement the above method, the embodiments of the present invention also provide a speech recognition device, which can be integrated in a terminal or a server.

[0175] For example, as Figure 6 shown, it is a schematic structural diagram of the speech recognition device provided by the embodiments of this application. The speech recognition device may include an extraction unit 301, a calculation unit 302, a first determination unit 303, a second determination unit 304, and an identification unit 305, as follows:

[0176] The extraction unit 301 is configured to extract features from the speech information to be recognized to obtain multiple feature vectors;

[0177] A computing unit 302, configured to calculate the sparsity value of each feature vector, where the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence;

[0178] A first determination unit 303, configured to determine a first feature vector with a sparsity value greater than a preset threshold and a second feature vector with a sparsity value not greater than the preset threshold;

[0179] A second determination unit 304, configured to determine a target matrix according to the self-attention calculation result of the first feature vector and the second feature vector;

[0180] An identification unit 305, configured to input the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing, and obtain an identification result corresponding to the speech information to be identified.

[0181] In some embodiments, the computing unit includes:

[0182] A first computing subunit, configured to calculate the self-attention score sequence of each feature vector;

[0183] A second computing subunit, configured to calculate the relative entropy between the distribution of each score sequence and the uniform distribution, and obtain the sparsity value of the feature vector corresponding to each score sequence.

[0184] In some embodiments, the computing unit includes:

[0185] A selection subunit, configured to randomly select a target number of feature vectors from multiple feature vectors to generate a key matrix;

[0186] A third computing subunit, configured to calculate the sparsity value of each feature vector according to the target number, multiple feature vectors, and the key matrix.

[0187] In some embodiments, the extraction unit includes:

[0188] A division subunit, configured to divide the speech information to be identified into multiple frames of speech signals;

[0189] A transformation subunit, configured to perform a discrete Fourier transform on each frame of speech signal to obtain the spectral information corresponding to each frame of speech signal;

[0190] A first processing subunit, configured to perform a Mel cepstrum process on the spectral information corresponding to each frame of speech signal to obtain multiple feature vectors of the speech information to be identified.

[0191] In some embodiments, the apparatus further includes:

[0192] A noise reduction unit, configured to perform noise reduction processing on the speech information to be identified;

[0193] A pre-emphasis unit for pre-emphasizing the voice information after noise reduction processing.

[0194] In some embodiments, the second determination unit includes:

[0195] A first acquisition subunit for acquiring a target self-attention score sequence corresponding to each first feature vector;

[0196] A fourth calculation subunit for performing weighted calculation according to the target self-attention score sequence and the value vector corresponding to each first feature vector to obtain a third feature vector corresponding to each first feature vector;

[0197] A determination subunit for determining a target matrix according to the third feature vector and the second feature vector.

[0198] In some embodiments, the first acquisition subunit includes:

[0199] A calculation module for calculating the dot product result of each target feature vector in the first feature vector and each feature vector;

[0200] A processing module for normalizing the dot product result to obtain a target self-attention score sequence corresponding to each first feature vector.

[0201] In some embodiments, the recognition unit includes:

[0202] A second acquisition subunit for acquiring a label sequence corresponding to the speech recognition result text;

[0203] An extraction subunit for extracting features from the label sequence to obtain a label feature vector corresponding to the label sequence;

[0204] A second processing subunit for processing the label feature vector by using an artificial neural network to obtain a feature matrix corresponding to the label sequence;

[0205] A recognition subunit for combining the target matrix and the feature matrix by using a multi-layer fully connected layer and inputting the combination result into a classification network for decoding to obtain a recognition result corresponding to the speech information to be recognized.

[0206] In specific implementation, the above units can be implemented as independent entities, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.

[0207] According to the above description, in the speech recognition method provided by the embodiment of the present application, the extraction unit 301 extracts features from the speech information to be recognized to obtain a plurality of feature vectors; the calculation unit 302 calculates the sparsity value of each feature vector, where the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence; the first determination unit 303 determines the first feature vectors with sparsity values greater than a preset threshold and the second feature vectors with sparsity values not greater than the preset threshold; the second determination unit 304 determines a target matrix according to the self-attention calculation result of the first feature vectors and the second feature vectors; the recognition unit 305 inputs the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain the recognition result corresponding to the speech information to be recognized. In this way, by calculating the sparsity values of the feature vectors and then determining that the feature vectors with sparsity values less than the preset threshold do not need to perform self-attention calculation, the amount of calculation can be reduced, thereby improving the efficiency of speech recognition.

[0208] The embodiment of the present application also provides a computer device, which can include the functions of a smart terminal, such as Figure 7 shown in the structural schematic diagram of the computer device provided by the present application. Specifically:

[0209] The computer device may include a processor 401 with one or more processing cores, a memory 402 with one or more storage media, a power supply 403, an input unit 404 and other components. Those skilled in the art can understand that Figure 7 the structure of the computer device shown in

[0210] does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine some components, or arrange different components. Among them:

[0211] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and audio processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, and web access, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 402 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.

[0212] The computer device further includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0213] The computer device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0214] Although not shown, the computer device may further include a display unit, a lighting unit, etc. The display unit is used to display the processing results of the processor 401. The display unit can also receive the touch operations of the user, generate touch instructions and transmit them to the processor for corresponding processing. The lighting unit can control the change of its lighting brightness according to the processing instructions of the processor.

[0215] Specifically in this embodiment, the processor 401 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to achieve various functions as follows:

[0216] Feature extraction is performed on the speech information to be recognized to obtain a plurality of feature vectors; the sparsity value of each feature vector is calculated, and the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence; the first feature vector with a sparsity value greater than the preset threshold and the second feature vector with a sparsity value not greater than the preset threshold are determined; a target matrix is determined according to the self-attention calculation result of the first feature vector and the second feature vector; the target matrix and the feature matrix corresponding to the label sequence are input into a classification network for classification processing to obtain the recognition result corresponding to the speech information to be recognized.

[0217] It should be noted that the computer device provided in the embodiments of the present application belongs to the same concept as the audio processing method in the above embodiments. The specific implementation of each of the above operations can be referred to the previous embodiments and will not be elaborated here.

[0218] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by instructions controlling related hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0219] Therefore, an embodiment of the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any one of the methods provided by the embodiments of the present invention. For example, the instructions can execute the following steps:

[0220] Feature extraction is performed on the speech information to be recognized to obtain a plurality of feature vectors; the sparsity value of each feature vector is calculated, and the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence; the first feature vector with a sparsity value greater than the preset threshold and the second feature vector with a sparsity value not greater than the preset threshold are determined; a target matrix is determined according to the self-attention calculation result of the first feature vector and the second feature vector; the target matrix and the feature matrix corresponding to the label sequence are input into a classification network for classification processing to obtain the recognition result corresponding to the speech information to be recognized.

[0221] The specific implementation of each of the above operations can be referred to the previous embodiments and will not be elaborated here.

[0222] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0223] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the methods provided in the embodiments of the present invention, the beneficial effects achievable by any of the methods provided in the embodiments of the present invention can be realized. For details, refer to the previous embodiments and will not be elaborated here.

[0224] Wherein, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a storage medium. The processor of the computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions, so that the computer device executes the above Figure 2 or Figure 4 methods provided in various alternative implementations.

[0225] The above has introduced in detail a voice recognition method, device, computer-readable storage medium and computer device provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A speech recognition method, characterized in that, The method includes: Performing feature extraction on the speech information to be recognized to obtain a plurality of feature vectors; Calculating the sparsity value of each feature vector, where the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence; Determining a first feature vector with a sparsity value greater than a preset threshold and a second feature vector with a sparsity value not greater than the preset threshold; Determining a target matrix according to the self-attention calculation result of the first feature vector and the second feature vector; Inputting the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain a recognition result corresponding to the speech information to be recognized; The step of inputting the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain a recognition result corresponding to the speech information to be recognized includes: Obtaining a label sequence corresponding to the speech recognition result text; Performing feature extraction on the label sequence to obtain a label feature vector corresponding to the label sequence; Processing the label feature vector by using an artificial neural network to obtain a feature matrix corresponding to the label sequence; Using a multi-layer fully connected layer to combine the target matrix and the feature matrix, and inputting the combination result into a classification network for decoding to obtain a recognition result corresponding to the speech information to be recognized.

2. The method according to claim 1, wherein The step of calculating the sparsity value of each feature vector includes: Calculating the self-attention score sequence of each feature vector; Calculating the relative entropy between the distribution of each score sequence and the uniform distribution to obtain the sparsity value of the feature vector corresponding to each score sequence.

3. The method according to claim 1, wherein The step of calculating the sparsity value of each feature vector includes: Randomly selecting a target number of feature vectors from the plurality of feature vectors to generate a key matrix; Calculating the sparsity value of each feature vector according to the target number, the plurality of feature vectors, and the key matrix.

4. The method according to any one of claims 1 to 3, characterized in that The step of performing feature extraction on the speech information to be recognized to obtain a plurality of feature vectors includes: Dividing the speech information to be recognized into multiple frames of speech signals; Performing discrete Fourier transform on each frame of speech signal to obtain the spectrum information corresponding to each frame of speech signal; Performing Mel cepstrum processing on the spectrum information corresponding to each frame of speech signal to obtain a plurality of feature vectors of the speech information to be recognized.

5. The method according to claim 4, characterized in that Before dividing the speech information to be recognized into multiple frames of speech signals, it further includes: Performing noise reduction processing on the speech information to be recognized; Performing pre-emphasis processing on the speech information after noise reduction processing.

6. The method according to claim 1, wherein The step of determining a target matrix according to the self-attention calculation result of the first feature vector and the second feature vector includes: Obtaining a target self-attention score sequence corresponding to each first feature vector; Performing weighted calculation according to the target self-attention score sequence and the value vector corresponding to each first feature vector to obtain a third feature vector corresponding to each first feature vector; Determining a target matrix according to the third feature vector and the second feature vector.

7. The method according to claim 6, characterized in that, The step of obtaining a target self-attention score sequence corresponding to each first feature vector includes: Calculating the dot product result of each target feature vector in the first feature vector and each feature vector; Normalize the dot product result to obtain the target self-attention score sequence corresponding to each first eigenvector.

8. A voice recognition device, characterized in that, The device includes: An extraction unit for extracting features from the speech information to be recognized to obtain a plurality of feature vectors; A calculation unit for calculating the sparsity value of each feature vector, where the sparsity value is the relative entropy between the distribution of the self-attention score sequence of each feature vector and the uniform distribution of the self-attention score sequence; A first determination unit for determining the first feature vectors with sparsity values greater than a preset threshold and the second feature vectors with sparsity values not greater than the preset threshold; A second determination unit for determining a target matrix according to the self-attention calculation result of the first feature vector and the second feature vector; An identification unit for inputting the target matrix and the feature matrix corresponding to the label sequence into a classification network for classification processing to obtain the recognition result corresponding to the speech information to be recognized; The identification unit is used for: Obtaining the label sequence corresponding to the speech recognition result text; Extracting features from the label sequence to obtain the label feature vector corresponding to the label sequence; Processing the label feature vector by using an artificial neural network to obtain the feature matrix corresponding to the label sequence; Combining the target matrix and the feature matrix by using a multi-layer fully connected layer and inputting the combination result into a classification network for decoding to obtain the recognition result corresponding to the speech information to be recognized.

9. The device according to claim 8, characterized in that, The calculation unit includes: A first calculation subunit for calculating the self-attention score sequence of each feature vector; A second calculation subunit for calculating the relative entropy between the distribution of each score sequence and the uniform distribution to obtain the sparsity value of the feature vector corresponding to each score sequence.

10. The device according to claim 8, characterized in that, The calculation unit includes: A selection subunit for randomly selecting a target number of feature vectors from the plurality of feature vectors to generate a key matrix; A third calculation subunit for calculating the sparsity value of each feature vector according to the target number, the plurality of feature vectors, and the key matrix.

11. The device according to any one of claims 8 to 10, characterized in that, The extraction unit includes: A division subunit for dividing the speech information to be recognized into multiple frames of speech signals; A transformation subunit for performing a discrete Fourier transform on each frame of speech signal to obtain the spectrum information corresponding to each frame of speech signal; A processing subunit for performing mel cepstrum processing on the spectrum information corresponding to each frame of speech signal to obtain a plurality of feature vectors of the speech information to be recognized.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the speech recognition method according to any one of claims 1 to 7.

13. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the speech recognition method according to any one of claims 1 to 7.

14. A computer program product, characterized in that, The computer program includes computer instructions stored in a storage medium. A processor of a computer device reads the computer instructions from the storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of the voice recognition method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Speaker identification method and device, electronic equipment and storage medium

    CN111785287A