Speech recognition method, device, computer equipment and storage medium

By extracting the timing and frequency characteristics of speech audio and using the speech recognition model to generate the phoneme probability distribution matrix, the problem of insufficient semantic fluency and accuracy of speech recognition in the prior art is solved, and higher recognition accuracy and semantic fluency are achieved.

CN114141237BActive Publication Date: 2025-05-06ZHAOLIAN CONSUMER FINANCE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111309392.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-06
Publication Date
2025-05-06
Estimated Expiration
2041-11-06

AI Technical Summary

Technical Problem

Existing speech recognition technologies have shortcomings in identifying semantic fluency and accuracy, especially in audio data of different pronunciators, the generalization ability of the model is weak.

Method used

By obtaining an audio set with the same audio sample rate, each audio is subjected to timing feature and frequency feature extraction processing to generate feature matrix data. Then, the feature matrix data is encoded using the speech recognition model to generate a phoneme probability distribution matrix. The model parameters are adjusted based on the phoneme alignment loss value, and the model is iteratively trained until the stop condition is met.

Benefits of technology

The speech recognition model's recognition accuracy of audio can be improved, and it can be accurate to the phoneme level, enhance semantic fluency and recognition accuracy, and improve the generalization performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114141237B_ABST
    Figure CN114141237B_ABST
Patent Text Reader

Abstract

The present application relates to a speech recognition method, device, computer equipment and storage medium. The method comprises: obtaining an audio set with the same audio sampling rate; extracting and processing the time series features and frequency features of each audio in the audio set to obtain feature matrix data corresponding to the audio including time series and frequency feature information; in the process of iteratively training the speech recognition model, for each audio, encoding the feature matrix data corresponding to the audio through the speech recognition model to obtain the phoneme probability distribution of each time frame to generate the phoneme probability distribution matrix corresponding to the audio; based on the phoneme probability distribution matrix and the text sentence annotated for the audio, determining the phoneme alignment loss value of the forward propagation of the speech recognition model; adjusting the model parameters of the speech recognition model based on the phoneme alignment loss value to continue iterating until the iteration stop condition is met to obtain a trained speech recognition model. The adoption of this method can improve the accuracy of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of machine learning, and in particular to a speech recognition method, apparatus, computer device and storage medium. Background Art

[0002] With the advancement of machine learning technology and the rapid popularization of mobile Internet, computer technology has been widely used in various fields of society, and with it comes the generation of massive data. Among them, speech recognition technology is about to enter various fields such as industry, home appliances, communications, automotive electronics, medical care, home services, and consumer electronics, so there is a wide demand for speech recognition technology.

[0003] The current speech recognition technology is obtained through training of speech recognition models based on probability statistics or recurrent neural network models. Different fields need to collect a large amount of relevant audio data to train the speech recognition model. However, the semantics recognized by these two models after training are not smooth enough, and the recognition accuracy is not high enough. Summary of the invention

[0004] Based on this, it is necessary to provide a speech recognition method, apparatus, computer device, storage medium and computer program product that can improve the accuracy of speech recognition in response to the above technical problems.

[0005] In a first aspect, the present application provides a speech recognition method. The method comprises:

[0006] Get an audio collection with the same audio sampling rate;

[0007] Performing time sequence feature and frequency feature extraction processing on each audio of the audio set to obtain feature matrix data corresponding to the audio including time sequence and frequency feature information;

[0008] In the process of iteratively training the speech recognition model using the audio set, for each audio, the feature matrix data corresponding to the audio is encoded by the speech recognition model to obtain the phoneme probability distribution of each time frame in the audio, so as to generate a phoneme probability distribution matrix corresponding to the audio;

[0009] Based on the phoneme probability distribution matrix and the text sentence annotated with respect to the audio, determining a phoneme alignment loss value of the forward propagation of the speech recognition model;

[0010] The model parameters of the speech recognition model are adjusted based on the phoneme alignment loss value to continue iteration until an iteration stop condition is met, thereby obtaining a trained speech recognition model.

[0011] In one of the embodiments, before performing the time series feature and frequency feature extraction process on each audio of the audio set, the method further includes:

[0012] Randomly extracting some audio from the audio collection;

[0013] Perform at least one of the following data enhancement processing on the randomly extracted audio:

[0014] Simulating a first difference between the voices of different speakers, and for the randomly selected audio, increasing or decreasing the volume of the audio based on the first difference;

[0015] A second difference between the speaking speeds of different speakers is simulated, and for the randomly selected audio, the speaking speed of the audio is accelerated or slowed down based on the second difference;

[0016] Simulating a third difference in the change of speech speed and rhythm of different speakers during speaking, and distorting the audio waveform data at a preset time frame based on the third difference for the randomly selected audio;

[0017] The fourth difference between the timbre frequencies of different speakers is simulated, and for the randomly selected audio, the audio frequency is distorted within a preset frequency range based on the fourth difference.

[0018] In one of the embodiments, before performing the time series feature and frequency feature extraction process on each audio of the audio set, the method further includes:

[0019] Statistics on the pronunciation of different speakers in a certain business scenario;

[0020] Extracting the corresponding occurrence probabilities in the business scenario for different levels of voice volume, speaking speed, speaking speed rhythm and timbre frequency produced by different speakers in the pronunciation situation;

[0021] With the purpose of achieving the same probability of occurrence and on the principle of comprehensive coverage, audios in the audio set are selected to perform volume perturbation, speed perturbation, time distortion and frequency distortion data enhancement.

[0022] In one embodiment, the extracting time sequence features and frequency features of each audio of the audio set to obtain feature matrix data corresponding to the audio including time sequence and frequency feature information comprises:

[0023] Based on a plurality of bandpass filters having triangular filtering characteristics, characteristic values ​​of different frequencies of the audio at each time frame are calculated to obtain Mel spectrum matrix data including frequency characteristic information;

[0024] The mel spectrum matrix data is subjected to dimensionality reduction processing to obtain feature matrix data.

[0025] In one embodiment, the speech recognition model to be trained includes a multi-layer convolutional neural network; the dimensionality reduction processing of the Mel spectrum matrix data to obtain feature matrix data includes:

[0026] Inputting the Mel spectrum matrix data into each layer of the convolutional neural network, triggering each layer of the convolutional neural network to perform two-dimensional convolution calculation, and obtaining feature matrix data after dimensionality reduction based on the calculation results;

[0027] The step of adjusting the model parameters of the speech recognition model based on the phoneme loss value to continue iteration includes:

[0028] The parameters of each layer of the convolutional neural network of the speech recognition model are adjusted based on the phoneme loss value to continue iteration.

[0029] In one embodiment, for each audio, encoding the feature matrix data corresponding to the audio through the speech recognition model to obtain the phoneme probability distribution of each time frame in the audio to generate the phoneme probability distribution matrix corresponding to the audio includes:

[0030] For each audio, the feature matrix data corresponding to the audio is input into a multi-layer multi-head attention model with a convolutional network of the speech recognition model, so as to extract features based on each layer of the multi-head attention model in turn for information concerned by the layer; each layer in the multi-head attention model concerns different information;

[0031] Concatenating the features extracted from each layer in the multi-head attention model to obtain a phoneme probability distribution for each time frame in the audio;

[0032] According to the phoneme probability distribution of each time frame in the audio, a phoneme probability distribution matrix including all time frames is generated.

[0033] In one embodiment, determining the phoneme alignment loss value of the forward propagation of the speech recognition model based on the phoneme probability distribution matrix and the text sentence annotated with respect to the audio includes:

[0034] Performing phoneme alignment on the audio-annotated text sentence and the phoneme probability distribution matrix;

[0035] After performing the phoneme alignment process, the phoneme alignment loss value of the forward propagation of the speech recognition model is determined.

[0036] In a second aspect, the present application also provides a speech recognition device. The device comprises:

[0037] Get a template for obtaining an audio collection with the same audio sampling rate;

[0038] A feature calculation module, used for performing time series feature and frequency feature extraction processing on each audio of the audio set to obtain feature matrix data corresponding to the audio including time series and frequency feature information;

[0039] A loss value calculation module is used to encode the feature matrix data corresponding to each audio through the speech recognition model in the process of iteratively training the speech recognition model using the audio set, obtain the phoneme probability distribution of each time frame in the audio, and generate a phoneme probability distribution matrix corresponding to the audio; based on the phoneme probability distribution matrix and the text sentence annotated for the audio, determine the phoneme alignment loss value of the forward propagation of the speech recognition model;

[0040] The optimization module is used to adjust the model parameters of the speech recognition model based on the phoneme alignment loss value to continue iterating until the iteration stop condition is met to obtain a trained speech recognition model.

[0041] In a third aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor executes the steps of the above-mentioned speech recognition method.

[0042] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program is used by a processor to execute the steps of the above-mentioned speech recognition method.

[0043] In a fifth aspect, the present application further provides a computer program product, wherein the computer program product comprises a computer program, and the computer program is used by a processor to execute the steps of the above-mentioned speech recognition method.

[0044] The above-mentioned speech recognition method, device, computer equipment, storage medium and computer program product obtain an audio set with the same audio sampling rate; extract and process the time series features and frequency features of each audio in the audio set to obtain the feature matrix data corresponding to the audio including the time series and frequency feature information. In the process of iteratively training the speech recognition model using the audio set, for each audio, the feature matrix data corresponding to the audio is encoded by the speech recognition model to obtain the phoneme probability distribution of each time frame in the audio, so as to generate the phoneme probability distribution matrix corresponding to the audio, so that the speech recognition model can accurately recognize the audio to the phoneme level. Based on the phoneme probability distribution matrix and the text sentence annotated for the audio, the phoneme alignment loss value of the forward propagation of the speech recognition model is determined. Based on the loss value, the model parameters of the speech recognition model are adjusted to continue iteration until the iteration stop condition is met to obtain a trained speech recognition model. Therefore, when using the language recognition model, the speech recognition model recognizes the audio at the phoneme level, improves the semantic fluency, and improves the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A diagram of an application environment of a speech recognition method in an embodiment;

[0046] Figure 2 A flowchart of a speech recognition method in one embodiment;

[0047] Figure 3 A schematic diagram of the principle of a speech recognition method in one embodiment;

[0048] Figure 4 is a structural block diagram of a speech recognition device in one embodiment;

[0049] Figure 5 is a structural block diagram of a feature calculation module in an embodiment;

[0050] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] The speech recognition method provided in this application can be applied to Figure 1In the application environment shown, the terminal 110 communicates with the server 120 through a network. The terminal 110 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server 120 may be implemented as an independent server or a server cluster consisting of multiple servers.

[0053] The terminal 110 may perform sampling rate conversion on the audio of the audio set to obtain an audio set with the same sampling rate, and send the audio set to the server 120. The server 120 obtains an audio set with the same audio sampling rate; performs time series feature and frequency feature extraction processing on each audio of the audio set to obtain feature matrix data corresponding to the audio including time series and frequency feature information. In the process of iteratively training the speech recognition model using the audio set, the server 120 encodes the feature matrix data corresponding to the audio through the speech recognition model for each audio, obtains the phoneme probability distribution of each time frame in the audio, and generates the phoneme probability distribution matrix corresponding to the audio. Based on the phoneme probability distribution matrix and the text sentence annotated for the audio, the server 120 determines the phoneme alignment loss value of the forward propagation of the speech recognition model, and adjusts the model parameters of the speech recognition model based on the phoneme alignment loss value to continue iterating until the iteration stop condition is met, and finally obtains the trained speech recognition model. The server 120 uses the trained speech recognition model to recognize the speech sent by the terminal 110, and sends the corresponding recognition result to the terminal 110.

[0054] In one embodiment, Figure 2 As shown, a speech recognition method is provided. This embodiment takes the method applied to a server as an example. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0055] S202, obtaining an audio set having the same audio sampling rate; performing time sequence feature and frequency feature extraction processing on each audio in the audio set to obtain feature matrix data corresponding to the audio including time sequence and frequency feature information.

[0056] The audio sampling rate refers to the frequency used to convert the analog audio signal into a digital signal during the audio generation process, such as 44,100 Hz or 48,000 Hz. The feature matrix data includes the feature values ​​of different frequencies of the audio at each time frame and the feature values ​​of the time series.

[0057] In one embodiment, before performing timing feature and frequency feature extraction processing on each audio in the audio set, the server may perform at least one data enhancement method including volume disturbance, speed disturbance, time distortion, frequency distortion and random noise on the audio of the audio set.

[0058] In one embodiment, for audios in an audio collection, the server randomly extracts audios for data enhancement.

[0059] In one embodiment, the server may obtain mel-spectrogram matrix data including frequency feature information based on feature values ​​of different frequencies of the audio in each time frame, and perform dimensionality reduction processing on the mel-spectrogram matrix data to obtain feature matrix data.

[0060] In one embodiment, the server may perform dimensionality reduction processing on the Mel spectrum matrix based on a three-layer convolutional neural network layer of a speech recognition model to obtain feature matrix data.

[0061] Specifically, the server obtains an audio set with the same audio sampling rate; performs timing feature and frequency feature extraction processing on each audio in the audio set in each time frame to obtain feature matrix data corresponding to the audio including timing and frequency feature information of each time frame.

[0062] S204, in the process of iteratively training the speech recognition model using the audio set, for each audio, the feature matrix data corresponding to the audio is encoded through the speech recognition model to obtain the phoneme probability distribution of each time frame in the audio, so as to generate the phoneme probability distribution matrix corresponding to the audio.

[0063] The encoding is used to process the feature matrix data to obtain the phoneme probability distribution of each time frame of the audio. The encoding is implemented by the encoder of the speech recognition model.

[0064] Among them, phoneme refers to the smallest unit of speech divided according to the natural attributes of speech. Phoneme probability distribution refers to the probability distribution of phoneme values, and phoneme probability distribution matrix refers to the probability distribution of phonemes in all time frames of audio.

[0065] Specifically, in the process of iteratively training the speech recognition model using the audio collection, the server can, for each audio, encode the feature matrix data corresponding to the audio through the encoder of the speech recognition model to obtain the phoneme probability distribution of each time frame in the audio, and generate the phoneme probability distribution matrix corresponding to the audio based on the phoneme probability distribution of all time frames.

[0066] In one embodiment, the speech recognition model is a Transformer model (a model that introduces an attention mechanism). In other embodiments, the speech recognition model may also be other models that can obtain the phoneme probability distribution value of the audio in each time frame.

[0067] In one embodiment, the server can obtain the phoneme probability distribution matrix corresponding to the audio based on the seven-layer multi-head attention model with convolution of the Transformer model.

[0068] S206, based on the phoneme probability distribution matrix and the text sentence annotated with respect to the audio, determine a phoneme alignment loss value of the forward propagation of the speech recognition model.

[0069] Among them, the text sentence is the text form of the content expressed by the audio, and the phoneme alignment loss value refers to the loss value calculated after aligning the real text phoneme label and all correct paths corresponding to the label in the phoneme probability distribution matrix.

[0070] In one embodiment, the server calculates the phoneme alignment loss value by performing phoneme alignment processing on the phoneme probability distribution matrix and the phonemes of the text sentence through a time series classifier of a neural network.

[0071] In one embodiment, the server calculates the phoneme alignment loss value through a CTC (Connectionist Temporal Classification) algorithm, an algorithm that can avoid manual alignment of input and output and is used to solve the classification problem of time series data.

[0072] Specifically, the server uses an algorithm model that can perform phoneme alignment to perform phoneme alignment processing on the phoneme probability distribution matrix and the text sentence annotated for the audio, and obtains the phoneme alignment loss value of the forward propagation of the speech recognition model.

[0073] S208, adjusting the model parameters of the speech recognition model based on the phoneme alignment loss value to continue iteration until the iteration stop condition is met to obtain a trained speech recognition model.

[0074] Specifically, the server takes derivatives on various parameters of the speech recognition model based on the phoneme alignment loss value to obtain the update gradient of the speech recognition model, thereby updating the speech recognition model and learning the parameters. The server cyclically uses the audio collection to iteratively train the speech recognition model until the phoneme loss value is stable, then stops training and obtains a trained speech recognition model.

[0075] In one embodiment, the server may re-perform data enhancement processing on the audio collection before each iteration of training.

[0076] In one embodiment, during each iterative training process, the server may gradually reduce the learning rate of the speech recognition model.

[0077] In one embodiment, during each iterative training process, the server may gradually reduce the learning rate of the multi-head attention model in the speech recognition model.

[0078] In one embodiment, during each iterative training process, the server may gradually reduce the learning rate of the three-layer convolutional neural network layer in the speech recognition model.

[0079] In one embodiment, the learning rate of the speech recognition model gradually decreases as the number of training steps increases.

[0080] In one embodiment, the learning rate of the speech recognition model is updated based on the inverse of the current training step number.

[0081] The above-mentioned speech recognition method, device, computer equipment and storage medium obtain an audio set with the same audio sampling rate; perform time series feature and frequency feature extraction processing on each audio in the audio set to obtain feature matrix data corresponding to the audio including time series and frequency feature information. In the process of iteratively training the speech recognition model using the audio set, for each audio, the feature matrix data corresponding to the audio is encoded by the speech recognition model to obtain the phoneme probability distribution of each time frame in the audio, so as to generate the phoneme probability distribution matrix corresponding to the audio, so that the speech recognition model can accurately recognize the audio to the phoneme level. Based on the phoneme probability distribution matrix and the text sentence annotated for the audio, the phoneme alignment loss value of the forward propagation of the speech recognition model is determined. Based on the loss value, the model parameters of the speech recognition model are adjusted to continue iteration until the iteration stop condition is met to obtain a trained speech recognition model. Therefore, when using the language recognition model, the speech recognition model recognizes the audio at the phoneme level, improves the semantic fluency, and improves the recognition accuracy.

[0082] In one embodiment, before performing timing feature and frequency feature extraction processing on each audio of the audio set, the method also includes: randomly extracting part of the audio from the audio set; performing at least one of the following data enhancement processing on the randomly extracted audio: simulating a first difference between the volume of voices of different speakers, and for the randomly extracted audio, enhancing or weakening the volume of the audio based on the first difference; simulating a second difference between the speaking speeds of different speakers, and for the randomly extracted audio, speeding up or slowing down the speaking speed of the audio based on the second difference; simulating a third difference in the changes in speaking speed and rhythm of different speakers during speaking, and for the randomly extracted audio, distorting the audio waveform data on a preset time frame based on the third difference; simulating a fourth difference between the timbre frequency of different speakers, and for the randomly extracted audio, distorting the audio frequency on a preset frequency range based on the fourth difference.

[0083] Specifically, before extracting the timing features and frequency features of each audio in the audio collection, the server randomly extracts part of the audio from the audio collection and performs at least one data enhancement process including volume disturbance, speed disturbance, time distortion, frequency distortion and random noise on the part of the audio.

[0084] In one embodiment, the process of simulating the first difference, the second difference, the third difference, and the fourth difference includes grading the corresponding indicator objects, obtaining the differences between the indicator objects of different levels, and simulating the difference features to the audio data enhancement process, wherein the indicator objects correspond to the sound volume, speech rate, speech rate rhythm, and timbre frequency, respectively. For example, simulating the first difference between the sound volumes of different speakers includes grading the sound volumes of different speakers, obtaining the first difference between the sound volumes of different levels, and simulating the first difference features to the audio volume data enhancement process.

[0085] In one embodiment, the volume disturbance process includes: the server can classify the voices of different speakers, obtain the first difference between the voices of different levels, and increase or decrease the volume of the audio according to the first difference. It can be understood that by increasing or decreasing the volume, the recognition ability of the speech recognition model for audios of different volumes can be enhanced.

[0086] In one embodiment, the speed disturbance process includes: the server can classify the speaking speeds of different speakers, obtain the second difference between the different levels of speaking speeds, and increase or decrease the speaking speed of the audio according to the second difference. It can be understood that by increasing or decreasing the speaking speed, the recognition ability of the speech recognition model for different speaking speeds can be enhanced.

[0087] In one embodiment, the time warping process includes: the server can classify the speech rate and rhythm during the speaking process, obtain the third difference between the speech rate and rhythm of different levels, and perform nonlinear warping on the audio waveform data on one or more random time frames based on the third difference. It can be understood that the recognition ability of the speech recognition model for speech rate and rhythm changes can be enhanced by warping the audio waveform.

[0088] In one embodiment, the frequency distortion process includes: the server can classify the timbre frequencies of different speakers, obtain the fourth difference between the timbre frequencies of different levels, and distort the audio frequency within a preset frequency range based on the fourth difference. It can be understood that the recognition ability of the speech recognition model for different timbres can be enhanced by distorting the frequency.

[0089] In one embodiment, the process of random noise includes: the server can randomly select a section of background noise from an existing background noise library and superimpose it on the audio data. It can be understood that superimposing noise on the audio can enhance the anti-interference ability of the speech recognition model under noise.

[0090] In this embodiment, by randomly performing at least one of volume disturbance, speed disturbance, time distortion, frequency distortion and random noise on the audio, richer audio can be obtained, so that richer audio feature matrix data can be obtained for training the speech recognition model to improve the generalization performance and recognition accuracy of the speech recognition model. In addition, in this embodiment, since diversified data enhancement is performed on the audio, a diversified data set with an increased scale can be obtained based on a small-scale data set, thereby improving the generalization performance and recognition accuracy of the model.

[0091] In one embodiment, before performing timing feature and frequency feature extraction processing on each audio in the audio set, the method also includes: calculating a first pronunciation range feature of the speech of a target speaker in a target business scenario; the first pronunciation range feature includes target probabilities of different levels of sound volume, speaking speed, speaking rate rhythm and timbre frequency appearing in the speech of the target speaker; for each audio in the audio set, performing volume perturbation, speed perturbation, time distortion and frequency distortion data enhancement on the audio, so that the second pronunciation range feature and the first pronunciation range feature of the audio and the corresponding enhanced audio are the same.

[0092] The target speaker may be one person or multiple persons. The target speaker's voice includes multiple audios, which are a minimum set of voice volume, speaking speed, speaking speed rhythm and timbre frequency at different levels in the target business scenario.

[0093] In one embodiment, different levels of voice volume, speaking speed, speech rhythm and timbre frequency are obtained based on grading the voice volume, speaking speed, speech rhythm and timbre frequency. For example, in terms of voice volume, the volume of a sound of 1 decibel is set to level 1, the volume of a sound of 10 decibels is set to level 2, the volume of a sound of 18 decibels is set to level 3, and so on.

[0094] In this embodiment, by calculating the first pronunciation range feature of the target speaker's voice in the target business scenario; for each audio in the audio set, the audio is enhanced by volume disturbance, speed disturbance, time distortion and frequency distortion data, so that the second pronunciation range feature of the audio and the corresponding enhanced audio are the same as the first pronunciation range feature, so as to obtain a richer and more comprehensive audio, so that a more comprehensive audio feature matrix data can be obtained for training the speech recognition model to improve the generalization performance and recognition accuracy of the speech recognition model. And in the case of difficulty in speech collection or high cost of speech collection, data enhancement can be performed on a small data volume audio set to obtain a rich and comprehensive audio set, reduce the cost of manual collection and solve the problem of difficulty in speech collection.

[0095] In one embodiment, each audio in an audio set is subjected to timing feature and frequency feature extraction processing to obtain feature matrix data corresponding to the audio including timing and frequency feature information, including: based on multiple bandpass filters with triangular filtering characteristics, characteristic values ​​of different frequencies of the audio in each time frame are calculated to obtain Mel spectrum matrix data including frequency feature information; and the Mel spectrum matrix data is subjected to dimensionality reduction processing to obtain feature matrix data.

[0096] Among them, a bandpass filter is a device that allows waves of a specific audio frequency band to pass through while shielding other audio frequency bands. Mel spectrum matrix data is used to represent the matrix data obtained by processing audio based on Mel spectrum.

[0097] Specifically, the server sets multiple bandpass filters within the frequency spectrum of the speech, each filter has a triangular filtering characteristic, and its center frequency is evenly distributed within the frequency range of human ear perception. The server calculates the characteristic values ​​of different frequencies of the audio at each time frame based on each filter, and obtains Mel spectrum matrix data including frequency feature information. The server performs reduction processing on the Mel spectrum matrix data to obtain feature matrix data.

[0098] In one embodiment, the number of filters may be 80.

[0099] In this embodiment, the eigenvalues ​​of different frequencies of the audio in each time frame are calculated to obtain mel spectrum matrix data including frequency feature information, and the mel spectrum matrix data is subjected to dimensionality reduction processing to obtain feature matrix data. In this way, the server can obtain data including audio frequency feature information, and perform dimensionality reduction on the data, thereby obtaining feature matrix data with comprehensive information and small data volume, reducing the training load of the speech recognition model and improving the accuracy.

[0100] In one embodiment, the speech recognition model to be trained includes a multi-layer convolutional neural network; reducing the dimension of the Mel-spectrogram matrix data to obtain feature matrix data includes: inputting the Mel-spectrogram matrix data into each layer of the convolutional neural network, triggering each layer of the convolutional neural network to perform a two-dimensional convolution calculation, and obtaining the reduced-dimensional feature matrix data based on the calculation results; adjusting the model parameters of the speech recognition model based on the phoneme alignment loss value to continue iteration includes: adjusting the parameters of each layer of the convolutional neural network of the speech recognition model based on the phoneme alignment loss value to continue iteration.

[0101] Specifically, the speech recognition model to be trained includes a multi-layer convolutional neural network, and each layer of the convolutional neural network can perform two-dimensional convolution calculations. The server inputs the Mel spectrum matrix data into the convolutional neural network of the speech recognition model, and each layer of the convolutional neural network performs a two-dimensional convolution calculation on the input Mel spectrum matrix data, and obtains the feature matrix data after dimensionality reduction based on the calculation results.

[0102] Specifically, after obtaining the phoneme alignment loss index through step S206, the server can also adjust the parameters of each layer of the convolutional neural network of the speech recognition model based on the phoneme alignment loss index.

[0103] In one embodiment, the server may set the convolution kernel of the convolutional neural network of the speech recognition model to (3, 3) and the step setting to (2, 2).

[0104] In this embodiment, the mel spectrum matrix data is input into each layer of the convolutional neural network of the speech recognition model for dimensionality reduction processing, and the parameters of each layer of the convolutional neural network are adjusted based on the phoneme loss value to meet the requirements of iterative training, thereby improving the accuracy of the speech recognition model.

[0105] In one embodiment, for each audio, the feature matrix data corresponding to the audio is encoded through a speech recognition model to obtain the phoneme probability distribution of each time frame in the audio, so as to generate a phoneme probability distribution matrix corresponding to the audio, including: for each audio, the feature matrix data corresponding to the audio is input into the multi-layer multi-head attention model with convolutional network of the speech recognition model, so as to extract features based on each layer of the multi-head attention model in turn for the information concerned by the layer; each layer in the multi-head attention model pays attention to different information; the features extracted from each layer in the multi-head attention model are spliced ​​to obtain the phoneme probability distribution of each time frame in the audio; according to the phoneme probability distribution of each time frame in the audio, a phoneme probability distribution matrix including all time frames is generated.

[0106] Among them, the multi-head attention model refers to a model that has multiple different attention processing for the input Mel spectrum matrix data and correspondingly has multiple layers of convolutional networks.

[0107] Specifically, the speech recognition model includes a multi-head attention model, and each layer in the multi-head attention model focuses on different information. For each audio, the server inputs the feature matrix data corresponding to the audio into the multi-layer multi-head attention model with convolutional network of the speech recognition model, and sequentially performs calculations including layer normalization, feedforward network, convolution, multi-head attention, and feedforward network based on each layer of the multi-head attention model, thereby extracting features of the information concerned by the layer, and splicing the extracted features to obtain the phoneme probability distribution of each time frame in the audio. The server generates a phoneme probability distribution matrix including all time frames based on the phoneme probability distribution of each time frame in the audio.

[0108] In one embodiment, the speech recognition model includes a seven-layer multi-head attention model with a convolutional network.

[0109] In one embodiment, the layer normalization calculation performed by each layer of the multi-head attention model is used to normalize the data input to all neurons in this layer.

[0110] In one embodiment, the feed-forward network computations performed by each layer of the multi-head attention model are used to implement linear connections between layers.

[0111] In one embodiment, the convolution calculation performed by each layer of the multi-head attention model is based on a one-dimensional convolution with a convolution kernel set to 32 and a stride set to 1.

[0112] In one embodiment, the multi-head attention calculation performed by each layer of the multi-head attention model selects multiple feature information for parallel calculation.

[0113] In this embodiment, a multi-head attention model with different information of each layer is used to extract features of the information of interest, and the features extracted by each layer are spliced ​​to obtain the phoneme probability distribution of each time frame in the audio, thereby generating a phoneme probability distribution matrix including all time frames, and then an accurate and comprehensive phoneme probability distribution matrix can be generated to improve the accuracy of the speech recognition model. In addition, in this embodiment, the processing is performed according to each time frame, so as to achieve sequential encoding and realize streaming speech recognition.

[0114] In one embodiment, based on the phoneme probability distribution matrix and the text sentence for the audio annotation, determining the phoneme alignment loss value of the forward propagation of the speech recognition model includes: performing phoneme alignment on the text sentence for the audio annotation and the phoneme probability distribution matrix; after performing the phoneme alignment process, determining the phoneme alignment loss value of the forward propagation of the speech recognition model.

[0115] Specifically, when the server calculates the loss value between the audio-tagged text sentence and the phoneme probability distribution matrix, it performs phoneme alignment processing to calculate the phoneme alignment loss value.

[0116] In one embodiment, the server may process the audio-tagged text sentence and the phoneme probability distribution matrix based on the CTC algorithm, perform phoneme alignment processing, and thereby calculate the phoneme alignment loss value.

[0117] In this embodiment, the text sentence annotated with the audio and the phoneme probability distribution matrix are phoneme aligned; after the phoneme alignment process is performed, the phoneme alignment loss value of the forward propagation of the speech recognition model is determined. This makes the calculation result of the loss value more accurate, so that when adjusting the parameters of the speech recognition model based on the phoneme alignment loss value, a more accurate reference standard object can be provided, thereby improving the accuracy of the speech recognition model.

[0118] In one embodiment, Figure 3As shown, a schematic diagram of the principle of the speech recognition method is provided. Specifically, the server resamples the audio in the audio set to obtain audio with the same sampling rate. The server can also perform data enhancement on the audio in the audio set, and the data enhancement processing mainly includes: at least one of volume disturbance, speed disturbance, time distortion, frequency distortion, random noise, etc., to improve the diversity of audio data and the generalization ability of the model, and make it possible to obtain a diversified and comprehensive data set through an audio set with a small amount of data. The server extracts feature data from the audio in the audio set, and downsamples through the three-layer convolutional neural network layer of the speech recognition model to obtain feature matrix data corresponding to the audio including time series and frequency feature information, so that the server can reduce the system load of the speech recognition model while generating training data that meets the preset requirements. The server also inputs the feature matrix data into the seven-layer multi-head attention model with convolutional network of the speech recognition model (only one layer is drawn in the attached figure, and 7x is used to represent 7 layers), so that each layer sequentially performs calculations including layer normalization, feedforward network, convolution, multi-head attention, and feedforward network, and each layer extracts different information of concern, and finally generates a phoneme probability distribution matrix after splicing. The server performs phoneme alignment processing on the phoneme probability distribution matrix and the text sentence annotated by the audio by using a time sequence classifier. It can be understood that the time sequence classifier can be a CTC algorithm. After the server obtains the phoneme alignment loss value based on the time sequence classifier, it adjusts the parameters of the speech model based on the phoneme alignment loss value and iteratively trains the speech recognition model.

[0119] It should be understood that, although each step in the flow chart in the embodiment of the present application is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is clear explanation in this article, the execution of these steps does not have strict order restriction, and these steps can be performed in other order. Moreover, at least a portion of the steps in the flow chart may include a plurality of steps or a plurality of stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.

[0120] Based on the same inventive concept, the embodiment of the present application also provides a speech recognition device for implementing the speech recognition method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more speech recognition device embodiments provided below can refer to the limitations on the speech recognition method above, and will not be repeated here.

[0121] In one embodiment, Figure 4As shown, a speech recognition device 400 is provided, comprising: an acquisition module 402, a feature calculation module 404, a loss value calculation module 406 and an optimization module 408, wherein:

[0122] The acquisition module 402 is used to acquire an audio set having the same audio sampling rate.

[0123] The feature calculation module 404 is used to extract the time sequence feature and frequency feature of each audio in the audio set to obtain feature matrix data corresponding to the audio including time sequence and frequency feature information.

[0124] The loss value calculation module 406 is used to encode the feature matrix data corresponding to the audio through the speech recognition model for each audio in the process of iteratively training the speech recognition model using the audio set, and obtain the phoneme probability distribution of each time frame in the audio to generate the phoneme probability distribution matrix corresponding to the audio; based on the phoneme probability distribution matrix and the text sentence annotated for the audio, determine the phoneme alignment loss value of the forward propagation of the speech recognition model.

[0125] The optimization module 408 is used to adjust the model parameters of the speech recognition model based on the phoneme alignment loss value to continue iterating until the iteration stop condition is met to obtain a trained speech recognition model.

[0126] In one embodiment, before performing timing feature and frequency feature extraction processing on each audio of the audio set, the feature calculation module 404 is also used to: randomly extract part of the audio from the audio set; perform at least one of the following data enhancement processing on the randomly extracted audio: simulate a first difference between the volume of voices of different speakers, and for the randomly extracted audio, enhance or weaken the volume of the audio based on the first difference; simulate a second difference between the speaking speeds of different speakers, and for the randomly extracted audio, speed up or slow down the speaking speed of the audio based on the second difference; simulate a third difference in the changes in speaking speed and rhythm of different speakers during speaking, and for the randomly extracted audio, distort the audio waveform data on a preset time frame based on the third difference; simulate a fourth difference between the timbre frequency of different speakers, and for the randomly extracted audio, distort the audio frequency on a preset frequency range based on the fourth difference.

[0127] In one embodiment, Figure 5 As shown, the feature calculation module 404 includes: a feature extraction module 404a, and a dimension reduction module 404b, wherein:

[0128] The feature extraction module 404a is used to calculate the feature values ​​of different frequencies of the audio in each time frame based on multiple bandpass filters with triangular filtering characteristics, and obtain Mel spectrum matrix data including time sequence and frequency feature information.

[0129] The dimension reduction module 404b is used to perform dimension reduction processing on the Mel spectrum matrix data to obtain feature matrix data.

[0130] In one embodiment, the speech recognition model to be trained includes a multi-layer convolutional neural network; the loss value calculation module 406 is also used to: input the Mel spectrum matrix data into each layer of the convolutional neural network, trigger each layer of the convolutional neural network to perform a two-dimensional convolution calculation, and obtain the reduced-dimensional feature matrix data based on the calculation results; the optimization module 408 is also used to adjust the parameters of each layer of the convolutional neural network of the speech recognition model based on the phoneme loss value to continue iteration.

[0131] In one embodiment, the loss value calculation module 406 is also used to input the feature matrix data corresponding to the audio into the multi-layer multi-head attention model with convolutional network of the speech recognition model for each audio, so as to extract features based on the information focused on by each layer of the multi-head attention model in turn; each layer in the multi-head attention model focuses on different information; the features extracted by each layer in the multi-head attention model are spliced ​​to obtain the phoneme probability distribution of each time frame in the audio; based on the phoneme probability distribution of each time frame in the audio, a phoneme probability distribution matrix including all time frames is generated.

[0132] In one embodiment, the loss value calculation module 406 is also used to: perform phoneme alignment on the audio-tagged text sentence and the phoneme probability distribution matrix; and after performing the phoneme alignment process, determine the phoneme alignment loss value of the forward propagation of the speech recognition model.

[0133] The above-mentioned speech recognition device obtains an audio set with the same audio sampling rate; performs time series feature and frequency feature extraction processing on each audio in the audio set, and obtains feature matrix data corresponding to the audio including time series and frequency feature information. In the process of iteratively training the speech recognition model using the audio set, for each audio, the feature matrix data corresponding to the audio is encoded by the speech recognition model, and the phoneme probability distribution of each time frame in the audio is obtained to generate the phoneme probability distribution matrix corresponding to the audio, so that the speech recognition model can accurately recognize the audio to the phoneme level. Based on the phoneme probability distribution matrix and the text sentence annotated for the audio, the phoneme alignment loss value of the forward propagation of the speech recognition model is determined. Based on the loss value, the model parameters of the speech recognition model are adjusted to continue iteration until the iteration stop condition is met, and a trained speech recognition model is obtained. Therefore, when using the language recognition model, the speech recognition model recognizes the audio at the phoneme level, improves the semantic fluency, and improves the recognition accuracy.

[0134] For the specific limitations of the above-mentioned speech recognition, please refer to the limitations of the above-mentioned speech recognition method, which will not be repeated here. Each module in the above-mentioned speech recognition device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0135] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store audio collection data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a speech recognition method is implemented.

[0136] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0137] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0138] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0139] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0140] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0141] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0142] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A speech recognition method, characterized in that: The method comprises: Get an audio collection with the same audio sampling rate; Performing time sequence feature and frequency feature extraction processing on each audio of the audio set to obtain feature matrix data corresponding to the audio including time sequence and frequency feature information; In the process of iteratively training the speech recognition model using the audio set, for each audio, encoding the feature matrix data corresponding to the audio through the speech recognition model, obtaining the phoneme probability distribution of each time frame in the audio, so as to generate a phoneme probability distribution matrix corresponding to the audio; Based on the phoneme probability distribution matrix and the text sentence annotated with respect to the audio, determining a phoneme alignment loss value of the forward propagation of the speech recognition model; The model parameters of the speech recognition model are adjusted based on the phoneme alignment loss value to continue iteration until an iteration stop condition is met, thereby obtaining a trained speech recognition model.

2. The method according to claim 1, characterized in that: Before performing the time series feature and frequency feature extraction process on each audio of the audio set, the method further includes: Randomly extracting some audio from the audio collection; Perform at least one of the following data enhancement processing on the randomly extracted audio: Simulating a first difference between the voices of different speakers, and for the randomly selected audio, increasing or decreasing the volume of the audio based on the first difference; A second difference between the speaking speeds of different speakers is simulated, and for the randomly selected audio, the speaking speed of the audio is accelerated or slowed down based on the second difference; Simulating a third difference in the change of speech speed and rhythm of different speakers during speaking, and distorting the audio waveform data at a preset time frame based on the third difference for the randomly selected audio; The fourth difference between the timbre frequencies of different speakers is simulated, and for the randomly selected audio, the audio frequency is distorted within a preset frequency range based on the fourth difference.

3. The method according to claim 1, characterized in that The extracting process of timing features and frequency features of each audio of the audio set to obtain feature matrix data corresponding to the audio including timing and frequency feature information includes: Based on a plurality of bandpass filters having triangular filtering characteristics, characteristic values ​​of different frequencies of the audio at each time frame are calculated to obtain Mel spectrum matrix data including frequency characteristic information; The mel spectrum matrix data is subjected to dimensionality reduction processing to obtain feature matrix data.

4. The method according to claim 3, characterized in that: The speech recognition model to be trained includes a multi-layer convolutional neural network; the dimensionality reduction processing of the Mel spectrum matrix data to obtain feature matrix data includes: Inputting the Mel spectrum matrix data into each layer of the convolutional neural network, triggering each layer of the convolutional neural network to perform two-dimensional convolution calculation, and obtaining feature matrix data after dimensionality reduction based on the calculation results; The step of adjusting the model parameters of the speech recognition model based on the phoneme loss value to continue iteration includes: The parameters of each layer of the convolutional neural network of the speech recognition model are adjusted based on the phoneme loss value to continue iteration.

5. The method according to claim 1, characterized in that For each audio, encoding the feature matrix data corresponding to the audio by the speech recognition model to obtain the phoneme probability distribution of each time frame in the audio to generate the phoneme probability distribution matrix corresponding to the audio includes: For each audio, the feature matrix data corresponding to the audio is input into a multi-layer multi-head attention model with a convolutional network of the speech recognition model, so as to extract features based on each layer of the multi-head attention model in turn for information concerned by the layer; each layer in the multi-head attention model concerns different information; Concatenating the features extracted from each layer in the multi-head attention model to obtain a phoneme probability distribution for each time frame in the audio; According to the phoneme probability distribution of each time frame in the audio, a phoneme probability distribution matrix including all time frames is generated.

6. The method according to claim 1, characterized in that The determining of the phoneme alignment loss value of the forward propagation of the speech recognition model based on the phoneme probability distribution matrix and the text sentence annotated with respect to the audio comprises: Performing phoneme alignment processing on the audio-annotated text sentence and the phoneme probability distribution matrix; After performing the phoneme alignment process, the phoneme alignment loss value of the forward propagation of the speech recognition model is determined.

7. A speech recognition device, characterized in that: The device comprises: Get a template for obtaining an audio collection with the same audio sampling rate; A feature calculation module, used for performing time series feature and frequency feature extraction processing on each audio of the audio set to obtain feature matrix data corresponding to the audio including time series and frequency feature information; A loss value calculation module is used to encode the feature matrix data corresponding to each audio through the speech recognition model in the process of iteratively training the speech recognition model using the audio set, obtain the phoneme probability distribution of each time frame in the audio, and generate a phoneme probability distribution matrix corresponding to the audio; based on the phoneme probability distribution matrix and the text sentence annotated for the audio, determine the phoneme alignment loss value of the forward propagation of the speech recognition model; The optimization module is used to adjust the model parameters of the speech recognition model based on the phoneme alignment loss value to continue iterating until the iteration stop condition is met to obtain a trained speech recognition model.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Speech recognition method and device and computer storage medium

    CN116844529A