Speech recognition method, training method of speech recognition model, equipment and medium
By performing frame reduction processing on the stream encoding characteristics of audio frames in the audio stream, combined with convolution and attention mechanisms, the problem of high speech recognition delay under long audio input is solved, low-latency and efficient speech recognition is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202410041891.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-18
AI Technical Summary
When the audio input by the user is long, existing voice recognition technologies are prone to generate a large amount of computing, resulting in a higher voice recognition delay and a poor user experience.
By defragmenting the stream encoding features of multiple audio frames in the audio stream, the number of audio frames is reduced, and the feature sequence length of non-stream encoding is reduced, thereby reducing the amount of non-stream encoding is used to improve feature expression capabilities and capture global information, and low-latency speech recognition is achieved.
It effectively reduces the delay of voice recognition, improves the efficiency and accuracy of voice recognition, and improves the user experience.
Smart Images

Figure CN120340468A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech processing, and in particular, to a speech recognition method, a training method for a speech recognition model, a device, and a medium. Background Art
[0002] Automatic speech recognition (ASR) is a technical means for converting speech into corresponding text. In recent years, speech recognition technology can be applied to various scenarios such as intelligent (artificial intelligence, AI) calls, voice assistants, and meeting records of electronic devices.
[0003] However, when the duration of the audio input by the user is relatively long, it is easy to generate a large amount of computation, which in turn leads to a relatively high latency in speech recognition, thus resulting in a poor user experience. Summary of the Invention
[0004] The present application provides a speech recognition method, a training method for a speech recognition model, a device, and a medium, which are used to provide a low-latency speech recognition solution and improve the user experience.
[0005] To achieve the above object, the present application adopts the following technical solutions:
[0006] In a first aspect, a speech recognition method is provided. The method includes: receiving an audio stream input by a user; performing streaming encoding on multiple audio frames in the audio stream to obtain streaming encoding features of the multiple audio frames in the audio stream.
[0007] Moreover, performing downsampling processing on the streaming encoding features of the multiple audio frames in the audio stream to obtain streaming encoding features of the multiple audio frames after downsampling processing. Wherein, the number of frames of the multiple audio frames after downsampling processing is less than the number of frames of the multiple audio frames in the audio stream. That is to say, through downsampling processing, the number of frames of the original audio frames in the audio stream is reduced.
[0008] Furthermore, performing non-streaming encoding on the streaming encoding features of the multiple audio frames after downsampling processing to obtain non-streaming encoding features of the audio stream. It should be understood that since downsampling processing reduces the number of frames of the original audio frames in the audio stream, it also reduces the number of audio frames to be non-streaming encoded, and can effectively shorten the length of the feature sequence to be non-streaming encoded.
[0009] Finally, based on the non-streaming encoding features of the audio stream, obtaining a speech recognition result of the audio stream.
[0010] In the above technical solution, by performing downsampling processing on the streaming encoding features of multiple audio frames in the audio stream, the number of original audio frames in the audio stream is reduced, and thus the number of audio frames to be non-streaming encoded is reduced. This can effectively shorten the length of the feature sequence to be non-streaming encoded, thereby reducing the amount of non-streaming encoding operations, effectively reducing the latency of speech recognition, and thus effectively improving the efficiency of speech recognition and enhancing the user experience.
[0011] In a possible implementation manner of the first aspect, after performing streaming encoding on multiple audio frames in the audio stream to obtain the streaming encoding features of the multiple audio frames in the audio stream, the method further includes: performing normalization processing on the streaming encoding features of the multiple audio frames to obtain the character probability distribution of the multiple audio frames.
[0012] Among them, the character probability distribution represents the probability values of the multiple audio frames on multiple candidate characters respectively. It should be understood that if the probability value of an audio frame on a certain candidate character is large, it means that the audio frame is very likely to correspond to that candidate character.
[0013] In this way, by performing normalization processing on the streaming encoding features of multiple audio frames, the character probability distribution of the multiple audio frames can be obtained, so as to use the character probability distribution of the multiple audio frames to perform the downsampling processing process subsequently.
[0014] Correspondingly, performing downsampling processing on the streaming encoding features of multiple audio frames in the audio stream to obtain the streaming encoding features of the multiple audio frames after downsampling processing includes: performing downsampling processing on the streaming encoding features of multiple audio frames in the audio stream according to the character probability distribution of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after downsampling processing.
[0015] In this possible implementation manner, an implementation manner of performing downsampling processing based on the character probability distribution of each audio frame is provided. In this way, the amount of information referred to in the downsampling processing is increased, and the effect of the downsampling processing can be effectively improved.
[0016] In another possible implementation of the first aspect, according to the character probability distribution of the multiple audio frames, downsampling processing is performed on the streaming encoding features of the multiple audio frames in the audio stream to obtain the streaming encoding features of the multiple audio frames after downsampling processing, including: performing convolutional processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolutional processing. That is to say, by performing convolutional processing to implement feature extraction, time-series features with enhanced local information can be obtained. Furthermore, from the streaming encoding features of the multiple audio frames after convolutional processing, the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions are extracted to obtain the streaming encoding features of the multiple audio frames after downsampling processing. In this way, the streaming encoding features of the extracted audio frames are also time-series features with enhanced local information, which can effectively improve the expression ability of the features.
[0017] In this possible implementation, a downsampling processing scheme based on convolution and frame selection is provided. By first performing convolutional processing on the streaming encoding features of the multiple audio frames, time-series features with enhanced local information can be obtained. Then, from the time-series features with enhanced local information, the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions are extracted, so that the streaming encoding features of the extracted audio frames are also time-series features with enhanced local information, which can effectively improve the expression ability of the features, thereby improving the accuracy of speech recognition.
[0018] In another possible implementation of the first aspect, according to the character probability distribution of the multiple audio frames, downsampling processing is performed on the streaming encoding features of the multiple audio frames in the audio stream to obtain the streaming encoding features of the multiple audio frames after downsampling processing, including: performing convolutional processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolutional processing. That is to say, by performing convolutional processing to implement feature extraction, time-series features with enhanced local information can be obtained. Then, from the streaming encoding features of the multiple audio frames after convolutional processing, the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions are extracted. In this way, the streaming encoding features of the extracted audio frames are also time-series features with enhanced local information, which can effectively improve the expression ability of the features. Furthermore, based on the attention mechanism, attention calculation is performed on the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions to obtain the streaming encoding features of the multiple audio frames after downsampling processing. In this way, considering that the local characteristics of convolution are likely to limit the capture of global information, through attention calculation, the streaming encoding features of multiple audio frames including global information can be combined to comprehensively generate the final time-series features, which can compensate for the information loss problem caused by frame loss.
[0019] In this possible implementation, a frame reduction processing solution based on convolution, frame selection, and attention mechanism is provided. By first performing convolution processing on the streaming encoding features of the multiple audio frames, temporal features with enhanced local information can be obtained. Then, from the temporal features with enhanced local information, the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions are extracted, so that the extracted streaming encoding features of the audio frames are also temporal features with enhanced local information, which can effectively improve the expression ability of the features. Furthermore, attention calculation is performed on the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions to capture global information, which can make up for the information loss caused by frame dropping, thereby enabling lossless performance low-latency speech recognition.
[0020] In another possible implementation of the first aspect, extracting the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions from the streaming encoding features of the multiple audio frames after convolution processing includes: determining the audio frames whose probability value of the preset character is less than the preset threshold according to the character probability distribution of the multiple audio frames.
[0021] Wherein, the preset character is at least one of a blank character or an interval character. It should be noted that if the probability value of the preset character of an audio frame is less than the preset threshold, it means that the probability of the audio frame corresponding to the blank character or the interval character is small, that is, the audio frame is very likely to generate other meaningful characters.
[0022] Furthermore, extracting the streaming encoding features of the audio frames whose probability value of the preset character is less than the preset threshold from the streaming encoding features of the multiple audio frames, the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions are obtained. That is, extracting the streaming encoding features of the audio frames that are very likely to generate other meaningful characters from the streaming encoding features of the multiple audio frames, which realizes the extraction of key audio frames or important audio frames, thereby realizing the frame reduction processing of the audio frames.
[0023] In another possible implementation of the first aspect, based on the attention mechanism, performing attention calculation on the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions to obtain the streaming encoding features of the multiple audio frames after frame reduction processing includes: performing matrix multiplication processing on the streaming encoding features of the multiple audio frames after convolution processing and the streaming encoding features of the multiple audio frames after frame reduction processing to obtain a similarity matrix. Wherein, the similarity matrix represents the similarity between the streaming encoding features of the multiple audio frames after convolution processing and the streaming encoding features of the multiple audio frames after frame reduction processing.
[0024] Moreover, the similarity matrix is subjected to Scale (normalization) processing to obtain the similarity matrix after Scale processing. Furthermore, the similarity matrix after Scale processing is normalized to obtain the normalized similarity matrix. Among them, the value of each element in the normalized similarity matrix is a numerical value between (0, 1). At this time, the normalized similarity matrix can be used as a weight matrix to perform matrix multiplication with the streaming encoded features of multiple audio frames after convolution processing, thereby completing the attention calculation.
[0025] Finally, matrix multiplication is performed on the streaming encoded features of multiple audio frames after convolution processing and the normalized similarity matrix to obtain the streaming encoded features of multiple audio frames after frame reduction processing.
[0026] In this way, considering that the local characteristics of convolution are likely to limit the capture of global information, through attention calculation, it is possible to combine the streaming encoded features of multiple audio frames including global information to comprehensively generate the final temporal features, which can compensate for the information loss problem caused by frame loss.
[0027] In another possible implementation manner of the first aspect, based on the non-streaming encoded features of the audio stream, the speech recognition result of the audio stream is obtained, including:
[0028] Based on the non-streaming encoded features of the audio stream and the text prediction results of multiple audio frames in the audio stream, the output probability distribution of the audio stream is determined. Among them, the text prediction result is predicted based on the speech recognition result of the previous audio frame of the audio frame. The output probability distribution represents the probability values of the multiple audio frames on the multiple candidate texts corresponding to the text prediction results. In this way, by comprehensively referring to the non-streaming encoded features of the audio stream and the text prediction results of multiple audio frames in the audio stream, not only can the efficiency of speech recognition be ensured, but also the accuracy of speech recognition can be improved.
[0029] Furthermore, the output probability distribution of the audio stream is decoded to obtain the speech recognition result of the audio stream.
[0030] In the second aspect, the present application provides a method for training a speech recognition model, and the method includes:
[0031] The initial model is iteratively trained based on audio training data, and at the end of the iterative training, the trained model is obtained as the speech recognition model.
[0032] Among them, the audio training data includes audio samples and the annotated text of the audio samples. That is to say, the initial model is iteratively trained based on audio samples and the annotated text of audio samples to obtain a speech recognition model with low speech recognition latency.
[0033] Among them, during any iteration training process, the audio training data is input into the model obtained after the previous iteration training, and multiple audio frames in the audio sample are stream-encoded to obtain the stream-encoding features of multiple audio frames in the audio sample.
[0034] Moreover, downsampling processing is performed on the stream-encoding features of multiple audio frames in the audio sample to obtain the stream-encoding features of multiple audio frames after downsampling processing. The number of frames of multiple audio frames after downsampling processing is less than the number of frames of multiple audio frames in the audio sample. Through downsampling processing, the number of original audio frames in the audio sample is reduced.
[0035] Furthermore, non-stream encoding is performed on the stream-encoding features of multiple audio frames after downsampling processing to obtain the non-stream encoding features of the audio sample. It should be understood that since downsampling processing reduces the number of original audio frames in the audio sample, it also reduces the number of audio frames to be non-stream encoded, and can effectively shorten the length of the feature sequence to be non-stream encoded.
[0036] Next, based on the non-stream encoding features of the audio sample, the speech recognition result of the audio sample is obtained. In this way, by performing downsampling processing on the stream-encoding features of multiple audio frames in the audio sample, the number of original audio frames in the audio sample is reduced, which also reduces the number of audio frames to be non-stream encoded, can effectively shorten the length of the feature sequence to be non-stream encoded, and further reduces the computational amount of non-stream encoding, can effectively reduce the latency of speech recognition, and thus can effectively improve the efficiency of speech recognition.
[0037] Finally, based on the speech recognition result and the annotated text, the model loss value is determined, and then the model parameters are adjusted based on the model loss value. In this way, by adjusting the model parameters according to the model loss value, the learning ability of the model can be improved, and thus a speech recognition model with better learning ability can be trained.
[0038] In a possible implementation manner of the second aspect, after stream-encoding multiple audio frames in the audio sample to obtain the stream-encoding features of multiple audio frames in the audio sample, the method further includes: performing normalization processing on the stream-encoding features of the multiple audio frames to obtain the character probability distribution of the multiple audio frames.
[0039] Among them, the character probability distribution represents the probability values of the multiple audio frames on multiple candidate characters respectively. It should be understood that if the probability value of an audio frame on a certain candidate character is large, it means that the audio frame is very likely to correspond to the candidate character.
[0040] In this way, by normalizing the streaming encoding features of multiple audio frames, the character probability distribution of the multiple audio frames can be obtained, so as to subsequently use the character probability distribution of the multiple audio frames to perform the process of frame dropping.
[0041] Correspondingly, frame dropping is performed on the streaming encoding features of multiple audio frames in the audio sample to obtain the streaming encoding features of the multiple audio frames after frame dropping, including: performing frame dropping on the streaming encoding features of multiple audio frames in the audio sample according to the character probability distribution of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after frame dropping.
[0042] In this possible implementation, an implementation of performing frame dropping based on the character probability distribution of each audio frame is provided. In this way, the amount of information referred to in frame dropping is increased, and the effect of frame dropping can be effectively improved.
[0043] In a possible implementation of the second aspect, frame dropping is performed on the streaming encoding features of multiple audio frames in the audio sample according to the character probability distribution of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after frame dropping, including: performing convolution processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolution processing. Furthermore, from the streaming encoding features of the multiple audio frames after convolution processing, the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions are extracted to obtain the streaming encoding features of the multiple audio frames after frame dropping.
[0044] In this possible implementation, a frame dropping processing scheme based on convolution and frame selection is provided. By first performing convolution processing on the streaming encoding features of the multiple audio frames, the temporal features with enhanced local information can be obtained. Then, from the temporal features with enhanced local information, the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions are extracted, so that the extracted streaming encoding features of the audio frames are also the temporal features with enhanced local information, which can effectively improve the expression ability of the features, thereby improving the accuracy of speech recognition.
[0045] In a possible implementation of the second aspect, according to the character probability distribution of the multiple audio frames, perform downsampling processing on the streaming encoding features of the multiple audio frames in the audio sample to obtain the streaming encoding features of the multiple audio frames after downsampling processing, including: performing convolution processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolution processing. Then, from the streaming encoding features of the multiple audio frames after convolution processing, extract the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions. Furthermore, based on the attention mechanism, perform attention calculation on the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions to obtain the streaming encoding features of the multiple audio frames after downsampling processing.
[0046] In this possible implementation, a downsampling processing scheme based on convolution, frame selection, and attention mechanism is provided. By first performing convolution processing on the streaming encoding features of the multiple audio frames, the temporal features with enhanced local information can be obtained. Then, from the temporal features with enhanced local information, extract the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions, so that the extracted streaming encoding features of the audio frames are also the temporal features with enhanced local information, which can effectively improve the expression ability of the features. Furthermore, perform attention calculation on the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions to capture global information, which can make up for the information loss problem caused by frame dropping, and thus can achieve lossless performance low-latency speech recognition.
[0047] In a possible implementation of the second aspect, extracting the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions from the streaming encoding features of the multiple audio frames after convolution processing includes: determining the audio frames whose probability value of the preset character is less than the preset threshold according to the character probability distribution of the multiple audio frames.
[0048] Wherein, the preset character is at least one of a blank character or an interval character. It should be noted that if the probability value of the preset character of an audio frame is less than the preset threshold, it means that the probability of the audio frame corresponding to the blank character or the interval character is small, that is, the audio frame is very likely to generate other meaningful characters.
[0049] Furthermore, extract the streaming encoding features of the audio frames whose probability value of the preset character is less than the preset threshold from the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions. That is, extract the streaming encoding features of the audio frames that are very likely to generate other meaningful characters from the streaming encoding features of the multiple audio frames, which realizes the extraction of key audio frames or important audio frames, and thus realizes the downsampling processing of the audio frames.
[0050] In a possible implementation of the second aspect, based on the attention mechanism, attention calculation is performed on the streaming encoding features of the audio frames whose character probability distribution satisfies the preset conditions, to obtain the streaming encoding features of the multiple audio frames after frame reduction, including: performing matrix multiplication on the streaming encoding features of the multiple audio frames after convolution processing and the streaming encoding features of the multiple audio frames after frame reduction, to obtain a similarity matrix. Wherein, the similarity matrix represents the similarity between the streaming encoding features of the multiple audio frames after convolution processing and the streaming encoding features of the multiple audio frames after frame reduction.
[0051] And, perform Scale processing on the similarity matrix to obtain the similarity matrix after Scale processing. Furthermore, perform normalization processing on the similarity matrix after Scale processing to obtain the similarity matrix after normalization processing. Wherein, the value of each element in the similarity matrix after normalization processing is a numerical value between (0, 1). At this time, the similarity matrix after normalization processing can be used as a weight matrix for matrix multiplication with the streaming encoding features of the multiple audio frames after convolution processing, thereby completing the attention calculation.
[0052] Finally, perform matrix multiplication on the streaming encoding features of the multiple audio frames after convolution processing and the similarity matrix after normalization processing to obtain the streaming encoding features of the multiple audio frames after frame reduction.
[0053] In this way, considering that the local characteristics of convolution are likely to limit the capture of global information, thus through attention calculation, it is possible to combine the streaming encoding features of multiple audio frames including global information to comprehensively generate the final temporal features, and it can make up for the information loss problem caused by frame loss.
[0054] In a possible implementation of the second aspect, based on the non-streaming encoding features of the audio sample, obtain the speech recognition result of the audio sample, including:
[0055] Based on the non-streaming encoding features of the audio sample and the text prediction results of multiple audio frames in the audio sample, determine the output probability distribution of the audio sample. Wherein, the text prediction results are predicted based on the speech recognition results of the previous audio frames of the audio frames. The output probability distribution represents the probability values of the multiple audio frames on the multiple candidate texts corresponding to the text prediction results. In this way, by comprehensively referring to the non-streaming encoding features of the audio sample and the text prediction results of multiple audio frames in the audio sample, it is not only possible to ensure the efficiency of speech recognition, but also improve the accuracy of speech recognition.
[0056] Furthermore, decode the output probability distribution of the audio sample to obtain the speech recognition result of the audio sample.
[0057] In a third aspect, the present application provides an electronic device, including: a processor and a memory. The memory is used to store program code, and the processor is used to call the program code stored in the memory, so as to implement any one of the methods provided in the first aspect or the second aspect.
[0058] In a fourth aspect, there is provided a computer-readable storage medium including program code, which, when running on an electronic device, causes the electronic device to execute any one of the methods provided in the first aspect or the second aspect.
[0059] In a fifth aspect, there is provided a computer program product including program code, which, when running on an electronic device, causes the electronic device to execute any one of the methods provided in the first aspect or the second aspect.
[0060] It should be noted that for the technical effects brought by any one of the implementation manners in the third aspect to the fifth aspect, reference may be made to the technical effects brought by the corresponding implementation manners in the first aspect or the second aspect, which will not be elaborated herein. Description of the Drawings
[0061] Figure 1 A schematic diagram of the framework of a speech recognition model provided for the related art;
[0062] Figure 2 Another schematic diagram of the framework of a speech recognition model provided for the related art;
[0063] Figure 3 A schematic diagram of an electronic device provided for an embodiment of the present application;
[0064] Figure 4 A schematic diagram of the hardware structure of an electronic device provided for an embodiment of the present application;
[0065] Figure 5 A schematic diagram of the software structure of an electronic device provided for an embodiment of the present application;
[0066] Figure 6 A schematic diagram of the flow of a speech recognition method provided for an embodiment of the present application;
[0067] Figure 7 A schematic diagram of the flow of a speech recognition method provided for an embodiment of the present application;
[0068] Figure 8 A schematic diagram of the framework of a speech recognition model provided for an embodiment of the present application;
[0069] Figure 9 A schematic diagram of the framework of a convolutional module provided for an embodiment of the present application;
[0070] Figure 10Schematic diagram of the framework of a downsampler provided by an embodiment of the present application;
[0071] Figure 11 Schematic diagram of the framework of an attention module provided by an embodiment of the present application;
[0072] Figure 12 Schematic diagram of the framework of another downsampler provided by an embodiment of the present application;
[0073] Figure 13 Schematic diagram of the process of a method for training a speech recognition model provided by an embodiment of the present application;
[0074] Figure 14 Schematic diagram of the process of a method for training a speech recognition model provided by an embodiment of the present application;
[0075] Figure 15 Schematic diagram of the model framework in the model training stage provided by an embodiment of the present application;
[0076] Figure 16 Schematic diagram of the framework of a speech recognition device provided by an embodiment of the present application;
[0077] Figure 17 Schematic diagram of the framework of a device for training a speech recognition model provided by an embodiment of the present application. Detailed implementation manners
[0078] In the description of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B. "And / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" means one or more, and "multiple" means two or more. The terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily mean different.
[0079] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0080] The speech recognition method provided by the embodiments of the present application can be applied to speech recognition scenarios of electronic devices, such as various scenarios including intelligent calls, voice assistants, and meeting records of electronic devices.
[0081] Speech recognition is a technical means to convert speech into corresponding text. Currently, electronic devices usually deploy speech recognition systems, and the speech recognition systems provide speech recognition models with a unified architecture. Correspondingly, when the electronic device triggers the speech recognition function, it can use the speech recognition model with the unified architecture to perform speech recognition.
[0082] Exemplarily, the speech recognition system can be a Transducer-based speech recognition system. In this Transducer-based speech recognition system, a speech recognition model as shown Figure 1 can be provided. Figure 1 It is a schematic diagram of the framework of a speech recognition model provided by the related art. Refer to Figure 1 and this speech recognition model can include three modules: an encoder, a prediction network, and a fusion network.
[0083] Among them, the encoder is used to process the feature sequence of the audio to extract effective features from the feature sequence of the audio. Generally, the encoder can be set with more model parameters to implement the above function of extracting effective features. For example, the encoder can adopt an architecture based on Transformer or an architecture based on Conformer (adaptive), both of which are based on the attention mechanism to process the feature sequence of the audio. The prediction network is used to predict future text based on the historical speech recognition results (i.e., the recognized text) of the audio. The fusion network provides a function of feature fusion and is used to perform fusion processing on the output results of the encoder and the output results of the prediction network, so as to obtain the final speech recognition result.
[0084] Furthermore, in some possible implementation manners, in addition to the above Figure 1 shown speech recognition model, in this Transducer-based speech recognition system, a speech recognition model as shown Figure 2 can also be provided. Figure 2 It is a schematic diagram of the framework of another speech recognition model provided by the related art. Refer to Figure 2 and this speech recognition model can include a cascaded encoder composed of a streaming encoder and a non-streaming encoder (or called a high-latency encoder), a prediction network, and a fusion network.
[0085] Among them, the streaming encoder is used to process the feature sequence of real-time input audio frames to extract effective features from the feature sequence of audio frames. The output result of the streaming encoder can be used as the input of the non-streaming encoder. Correspondingly, the non-streaming encoder is used to process the output result of the streaming encoder to further extract the effective features of the audio. The prediction network is used to predict future text based on the historical speech recognition results (i.e., the recognized text) of the audio. The fusion network provides a function of feature fusion. For example, in some possible implementation manners, the fusion network is used to perform fusion processing on the output result of the streaming encoder and the output result of the prediction network, so as to obtain the speech recognition result determined based on the streaming encoder. Another example is that in some other possible implementation manners, the fusion network is used to perform fusion processing on the output result of the non-streaming encoder and the output result of the prediction network, so as to obtain the speech recognition result determined by combining the streaming encoder and the non-streaming encoder.
[0086] It should be noted that the above Figure 2 provides a speech recognition model with both a streaming encoder and a non-streaming encoder. In some possible implementation manners, speech recognition can be performed solely based on the streaming encoder, so that real-time display of speech recognition results can be achieved. In some other possible implementation manners, speech recognition can also be performed by combining the streaming encoder and the non-streaming encoder, so that higher speech recognition accuracy can be achieved.
[0087] It should be noted that the above Figure 2 provided speech recognition model can also take into account electronic devices with different computing power deployment capabilities. For example, for an electronic device with low computing power deployment capabilities, speech recognition can be performed on the audio received by the electronic device based on the streaming encoder. Another example is that for an electronic device with high computing power deployment capabilities, speech recognition can be performed on the audio received by the electronic device by combining the streaming encoder and the non-streaming encoder.
[0088] In addition, it should also be noted that in the Figure 2 shown speech recognition model, the streaming encoder and the non-streaming encoder share the prediction network and the fusion network, which can effectively reduce the number of parameters of the speech recognition model, thereby reducing the computational amount and power consumption of the speech recognition model.
[0089] However, when the duration of the audio input by the user is long, it is easy to generate a large computational amount, which in turn leads to a high latency in speech recognition, thus resulting in a poor user experience. Exemplarily, taking the Figure 2 provided speech recognition model as an example, when the duration of the audio input by the user is long, the feature sequence received by the non-streaming encoder will also become longer, and since the time complexity of the attention mechanism of the non-streaming encoder is (T 2), which will bring a huge amount of computation and power consumption, resulting in a relatively high latency in speech recognition, and thus a poor user experience.
[0090] In view of this, the embodiments of the present application provide a speech recognition method. By performing downsampling processing on the streaming encoded features of multiple audio frames in the audio stream, the number of original audio frames in the audio stream is reduced, that is, the number of audio frames to be non-streaming encoded is reduced, which can effectively shorten the length of the feature sequence to be non-streaming encoded, and further reduce the amount of non-streaming encoding computation, effectively reducing the latency of speech recognition, and thus effectively improving the efficiency of speech recognition and enhancing the user experience.
[0091] In a possible implementation manner, the speech recognition method provided by the embodiments of the present application can be applied to an electronic device 300 as shown in Figure 3 FIG. Exemplarily, Figure 3 is a schematic diagram of an electronic device provided by the embodiments of the present application.
[0092] Among them, the electronic device 300 may be a terminal device 301. Exemplarily, the terminal device 301 may be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, a virtual reality terminal, an augmented reality terminal, a wireless terminal, and a laptop portable computer.
[0093] In some embodiments, the terminal device 301 may provide a speech recognition function. Refer to Figure 3 . By operating on the terminal device 301, the user can trigger the terminal device 301 to receive the audio stream (such as speech) input by the user. Furthermore, the terminal device 301 performs speech recognition on the audio stream input by the user to obtain the speech recognition result of the audio stream.
[0094] Alternatively, the electronic device 300 may also be a server 302. Exemplarily, the server 302 may be an independent physical server, such as a general server, a graphics processing unit (GPU) server, a data processing unit (DPU) server, an artificial intelligence (AI) server, etc., or at least one of cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data or artificial intelligence platforms. The embodiments of the present application do not limit this.
[0095] In some embodiments, the server 302 may provide a speech recognition function. Refer to Figure 3, the user can trigger the server 302 to receive the audio stream (such as voice) input by the user. For example, the user can upload the audio stream to the server 302 through the terminal device. Furthermore, the server 302 performs speech recognition on the audio stream input by the user to obtain the speech recognition result of the audio stream.
[0096] In the embodiment of the present application, the electronic device 300 is used to receive the audio stream input by the user; perform streaming encoding on multiple audio frames in the audio stream to obtain the streaming encoding features of the multiple audio frames in the audio stream; perform downsampling processing on the streaming encoding features of the multiple audio frames in the audio stream to obtain the streaming encoding features of the multiple audio frames after downsampling processing; perform non-streaming encoding on the streaming encoding features of the multiple audio frames after downsampling processing to obtain the non-streaming encoding features of the audio stream; and obtain the speech recognition result of the audio stream based on the non-streaming encoding features of the audio stream.
[0097] Exemplarily, Figure 3 The schematic structural diagram of the electronic device 300 in Figure 4 can be as shown in Figure 4 This is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application.
[0098] Referring to Figure 4 , the electronic device may include a processor 410, an external memory interface 420, an internal memory 421, a universal serial bus (USB) interface 430, a charging management module 440, antenna 1, antenna 2, a mobile communication module 450, a wireless communication module 460, an audio module 470, a speaker 470A, a receiver 470B, a microphone 470C, a headphone interface 470D, a sensor module 480, a key 490, a display screen 491, etc. Among them, the sensor module 480 may include a pressure sensor 480A, a touch sensor 480B, etc.
[0099] It can be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0100] The processor 410 may include one or more processing units. For example, the processor 410 may include an application processor (AP), a modem processor, a graphics processor, an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0101] Among them, the controller may be the nerve center and command center of the electronic device. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0102] A memory may also be provided in the processor 410 for storing instructions and data. In some possible implementation manners, the memory in the processor 410 is a cache memory. This memory may save the instructions or data that the processor 410 has just used or recycled. If the processor 410 needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor 410, and thus improves the efficiency of the system.
[0103] In some possible implementation manners, the processor 410 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0104] Among them, the I2S interface can be used for audio communication. In some possible implementation manners, the processor 410 may include multiple groups of I2S buses. The processor 410 can be coupled to the audio module 470 through the I2S bus to implement communication between the processor 410 and the audio module 470. For example, in the embodiment of the present application, the processor 410 can send an instruction to the audio module 470 through the I2S bus to trigger the audio module 470 to receive the audio stream input by the user. Another example, in the embodiment of the present application, the audio module 470 can send the received audio stream to the processor 410 through the I2S bus to trigger the processor 410 to execute the speech recognition method provided in the embodiment of the present application.
[0105] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some possible implementation manners, the audio module 470 and the wireless communication module 460 can be coupled through the PCM bus interface.
[0106] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a two-way communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some possible implementation manners, the UART interface is usually used to connect the processor 410 and the wireless communication module 460.
[0107] The MIPI interface can be used to connect the processor 410 and peripheral devices such as the display screen 491. The MIPI interface can include a display serial interface (DSI), etc. In some possible implementation manners, the processor 410 and the display screen 491 communicate through the DSI interface to implement the display function of the electronic device. For example, in the embodiment of the present application, the processor 410 can send an instruction to the display screen 491 through the DSI interface to trigger the display screen 491 to display the speech recognition result obtained by speech recognition.
[0108] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or a data signal. In some possible implementation manners, the GPIO interface can be used to connect the processor 410 to the display screen 491, the wireless communication module 460, the audio module 470, the sensor module 480, etc.
[0109] The USB interface 430 is an interface that complies with the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc.
[0110] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device. In other embodiments of the present application, the electronic device may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0111] The charging management module 440 is configured to receive a charging input from a charger. Herein, the charger may be a wireless charger or a wired charger.
[0112] The wireless communication function of the electronic device may be implemented through antenna 1, antenna 2, the mobile communication module 450, the wireless communication module 460, the modulation and demodulation processor, and the baseband processor, etc.
[0113] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device can be used to cover a single or multiple communication frequency bands. The mobile communication module 450 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device. The wireless communication module 460 can provide solutions for wireless communications applied to the electronic device.
[0114] In some possible implementation manners, antenna 1 of the electronic device is coupled to the mobile communication module 450, and antenna 2 is coupled to the wireless communication module 460, so that the electronic device can communicate with the network and other devices through wireless communication technologies.
[0115] The electronic device realizes the display function through the GPU, the display screen 491, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 491 and the application processor.
[0116] The display screen 491 is used to display images, videos, etc. The display screen 491 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some possible implementation manners, the electronic device may include one or N display screens 491, where N is a positive integer greater than 1. For example, in the embodiments of the present application, the display screen 491 is used to display the speech recognition result obtained by speech recognition.
[0117] The NPU is a neural-network (NN) computing processor. By referring to the structure of a biological neural network, such as the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc. For example, in the embodiments of the present application, the speech recognition method provided in the embodiments of the present application is implemented through the speech recognition model provided by the NPU.
[0118] The external memory interface 420 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device. The external memory card communicates with the processor 410 through the external memory interface 420 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0119] The internal memory 421 can be used to store computer-executable program code, which includes instructions. The processor 410 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 421. The internal memory 421 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device (such as audio data, phone book, etc.). In addition, the internal memory 421 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0120] The electronic device can implement audio functions through the audio module 470, the speaker 470A, the receiver 470B, the microphone 470C, the headphone jack 470D, and the application processor, etc. Such as music playback, recording, etc.
[0121] The audio module 470 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into a digital audio signal. The audio module 470 can also be used to encode and decode audio signals. In some possible implementation manners, the audio module 470 can be disposed in the processor 410, or some functional modules of the audio module 470 can be disposed in the processor 410.
[0122] The speaker 470A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device can listen to music or a hands-free call through the speaker 470A. The receiver 470B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device answers a call or a voice message, it can be held close to the user's ear to receive the voice through the receiver 470B. The microphone 470C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak close to the microphone 470C to input the sound signal into the microphone 470C. The electronic device can be provided with at least one microphone 470C. In some other embodiments, the electronic device can be provided with two microphones 470C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device can also be provided with three, four or more microphones 470C, which can collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording. The headphone jack 470D is used to connect a wired headphone. The headphone jack 470D can be a USB interface 430, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0123] The pressure sensor 480A is used to sense a pressure signal and can convert the pressure signal into an electrical signal. In some possible implementation manners, the pressure sensor 480A can be disposed on the display screen 491. There are many types of pressure sensors 480A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc.
[0124] The touch sensor 480B, also known as the "touch panel". The touch sensor 480B can be disposed on the display screen 491, and the touch sensor 480B and the display screen 491 form a touch screen, also known as the "touch display screen". The touch sensor 480B is used to detect a touch operation acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 491. In some other embodiments, the touch sensor 480B can also be disposed on the surface of the electronic device, at a different position from the display screen 491.
[0125] It should be noted that Figure 4 the structure shown in Figure 4In addition to the components shown, the electronic device may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0126] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture or a cloud architecture. The embodiment of the present application takes the Android system of the layered architecture as an example to illustrate the software structure of the electronic device. Figure 5 A schematic diagram of the software structure of an electronic device provided in an embodiment of the present application.
[0127] The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some possible implementations, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime, the system library, and the kernel layer.
[0128] The application layer can include a series of application packages. Figure 5 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0129] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0130] like Figure 5 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0131] In the embodiment of the present application, the human-computer interaction between the user and the electronic device is realized through the above application layer and application framework layer. For example, the electronic device realizes the process of receiving the audio stream input by the user through the above application layer and application framework layer.
[0132] The system library may include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0133] The kernel layer is the layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, and sensor driver.
[0134] Figure 6 The flowchart of a speech recognition method provided by an embodiment of the present application. Refer to Figure 6 , the method includes the following S601 - S605:
[0135] S601. Receive an audio stream input by a user.
[0136] Among them, the audio stream can be a piece of speech instantaneously input by the user to the electronic device, or a piece of audio pre - recorded by the user. The embodiment of the present application does not limit the audio stream.
[0137] S602. Perform streaming encoding on multiple audio frames in the audio stream to obtain streaming encoding features of multiple audio frames in the audio stream.
[0138] Among them, the streaming encoding features are used to characterize audio features at the frame level, that is, the audio features of each audio frame. It should be understood that streaming encoding means encoding while inputting audio.
[0139] S603. Perform frame - reduction processing on the streaming encoding features of multiple audio frames in the audio stream to obtain streaming encoding features of multiple audio frames after frame - reduction processing.
[0140] Among them, the number of frames of multiple audio frames after frame - reduction processing is less than the number of frames of multiple audio frames in the audio stream. That is to say, through frame - reduction processing, the number of frames of the original audio frames in the audio stream is reduced.
[0141] S604. Perform non - streaming encoding on the streaming encoding features of multiple audio frames after frame - reduction processing to obtain non - streaming encoding features of the audio stream.
[0142] Among them, the non - streaming encoding features are used to characterize the audio features of the audio stream. It should be understood that since frame - reduction processing reduces the number of frames of the original audio frames in the audio stream, it also reduces the number of audio frames to be non - streamingly encoded, and can effectively shorten the length of the feature sequence to be non - streamingly encoded.
[0143] S605. Based on the non - streaming encoding features of the audio stream, obtain the speech recognition result of the audio stream.
[0144] The technical solution provided by the embodiment of the present application reduces the number of frames of the original audio frames in the audio stream by performing frame - reduction processing on the streaming encoding features of multiple audio frames in the audio stream, and thus reduces the number of audio frames to be non - streamingly encoded, can effectively shorten the length of the feature sequence to be non - streamingly encoded, further reduces the amount of non - streaming encoding operations, can effectively reduce the latency of speech recognition, and thus can effectively improve the efficiency of speech recognition and enhance the user's usage experience.
[0145] Figure 7A flowchart of a speech recognition method provided by an embodiment of the present application. Refer to Figure 7 , this method is applied to an electronic device, and the electronic device is associated with a speech recognition model. The method includes the following S701 - S707:
[0146] S701. The electronic device receives an audio stream input by the user.
[0147] Among them, the audio stream can be a piece of speech instantaneously input by the user to the electronic device, or a piece of audio pre - recorded by the user. The embodiments of the present application do not limit the audio stream. Exemplarily, the audio stream input by the user can be represented by [x1, x2, …, x T . It should be understood that T refers to the number of audio frames in the audio stream, and T is a positive integer greater than 0.
[0148] In some possible implementation manners, the electronic device may be provided with a speech recognition control, and this speech recognition control is used to trigger speech recognition of the audio stream input by the user. For example, the electronic device receives a triggering operation of the user on the speech recognition control, and in response to the triggering operation on the speech recognition control, receives the audio stream input by the user through the built - in microphone of the electronic device.
[0149] Or, in some other possible implementation manners, the electronic device may be set with a wake - up word for enabling the speech recognition function, and this wake - up word is used to trigger speech recognition of the audio stream input by the user. For example, the electronic device receives a wake - up instruction from the user, and if the wake - up word carried in the wake - up instruction is the same as the set wake - up word, it receives the audio stream input by the user through the built - in microphone of the electronic device.
[0150] It should be noted that the electronic device can also adopt other implementation manners to receive the audio stream input by the user. The embodiments of the present application do not limit this.
[0151] Exemplarily, Figure 8 A framework diagram of a speech recognition model provided by an embodiment of the present application. Refer to Figure 8 , which shows a framework of a speech recognition model. This speech recognition model adopts a cascaded encoder including a streaming encoder, a frame - downsampler, and a non - streaming encoder, a linear network (or referred to as a linear layer), a prediction network, and a fusion network.
[0152] Among them, the streaming encoder is used to process the feature sequence of real - time input audio frames to extract effective features from the feature sequence of audio frames. The output result of the streaming encoder can be used as the input of the linear network, and the linear network is used to perform normalization processing on the output result of the streaming encoder to obtain a frame - level probability distribution. In the subsequent embodiments of the present application, the character probability distribution is used to refer to this frame - level probability distribution.
[0153] The output result of the linear network and the output result of the streaming encoder can both be used as the input of the downsampler. The downsampler is used to perform downsampling processing on the output result of the streaming encoder based on the output result of the linear network. The output result of the downsampler can be used as the input of the non-streaming encoder. Correspondingly, the non-streaming encoder is used to process the output result of the downsampler, that is, to process the output result of the streaming encoder after downsampling processing, so as to further extract the effective features of the audio stream.
[0154] The prediction network is used to predict future text based on the historical speech recognition results (i.e., the recognized text) of the audio stream. The fusion network provides a function of feature fusion, which is used to perform fusion processing on the output results of the cascaded encoder and the prediction network, so as to obtain the final speech recognition result.
[0155] In the embodiments of the present application, by adding a downsampler to connect the streaming encoder and the non-streaming encoder, and then using the downsampler to perform downsampling processing on the output result of the streaming encoder, and then inputting the output result of the streaming encoder after downsampling processing into the non-streaming encoder, it can effectively shorten the feature sequence received by the non-streaming encoder, reduce the computational amount and power consumption of the non-streaming encoder, effectively reduce the latency of speech recognition, and thus improve the user experience. Moreover, by adding a linear network to perform normalization processing on the output result of the streaming encoder to obtain a frame-level probability distribution, and then inputting the frame-level probability distribution into the downsampler, the downsampler performs downsampling processing on the output result of the streaming encoder according to the frame-level probability distribution, increasing the amount of information referred to in the downsampling processing, and can improve the accuracy of the downsampling processing.
[0156] Next, based on Figure 8 the speech recognition model shown, the speech recognition process will be described.
[0157] S702. The electronic device performs streaming encoding on multiple audio frames in the audio stream based on the streaming encoder, and obtains the streaming encoding features of multiple audio frames in the audio stream.
[0158] As Figure 7 shown in S702, after receiving the audio stream input by the user, the audio stream can be input into the streaming encoder according to audio frames, and the streaming encoder performs streaming encoding on the input audio frames to obtain the streaming encoding features of multiple audio frames in the audio stream.
[0159] Among them, the streaming encoding feature is used to characterize the audio feature at the frame level, that is, the audio features of each audio frame. It should be understood that streaming encoding means encoding while inputting the audio. For example, in some possible implementation manners, during the process of receiving the audio stream input by the user, the audio stream is input into the streaming encoder frame by frame, that is, each received audio frame is input into the streaming encoder. Furthermore, the streaming encoder performs streaming encoding on the input audio frame to obtain the streaming encoding feature of the audio frame. Exemplarily, the streaming encoding features of multiple audio frames in the audio stream can be represented by H T s =[h1 s ,h2 s ,…,h T s .
[0160] S703. The electronic device normalizes the streaming encoding features of the multiple audio frames to obtain the character probability distribution of the multiple audio frames.
[0161] Among them, the normalization process refers to limiting the streaming encoding features of each audio frame to a value between (0, 1). For example, in some possible implementation manners, the normalization process can be implemented based on the Softmax function, and the Softmax function can normalize the original data value to be converted into a value between (0, 1).
[0162] The character probability distribution represents the probability values of the multiple audio frames on multiple candidate characters respectively. Exemplarily, taking the audio frame M as an example, the corresponding multiple candidate characters can be character 1 (probability value 0.9), character 2 (probability value 0.7), character 3 (probability value 0.3), and so on. It should be understood that if the probability value of an audio frame on a certain candidate character is large, it means that the audio frame is very likely to correspond to the candidate character. Exemplarily, the character probability distribution can be represented by P T s =[p1 s ,p2 s ,…,p T s .
[0163] In some possible implementation manners, the electronic device normalizes the streaming encoding features of the multiple audio frames based on a linear network to obtain the character probability distribution of the multiple audio frames.
[0164] Such as Figure 7As shown in S703, after obtaining the streaming encoding features of multiple audio frames in the audio stream, the streaming encoding features of multiple audio frames in the audio stream can be input into a linear network, and through this linear network, the streaming encoding features of the multiple audio frames are normalized to obtain the character probability distribution of the multiple audio frames.
[0165] In the above embodiment, by setting a linear network to normalize the streaming encoding features of multiple audio frames output by the streaming encoder, the character probability distribution of the multiple audio frames can be obtained, so as to perform the process of frame reduction using the character probability distribution of the multiple audio frames subsequently.
[0166] S704. The electronic device, based on a frame reducer, performs frame reduction on the streaming encoding features of multiple audio frames in the audio stream according to the character probability distribution of the multiple audio frames, to obtain the streaming encoding features of the multiple audio frames after frame reduction.
[0167] Among them, the number of frames of the multiple audio frames after frame reduction is less than the number of frames of the multiple audio frames in the audio stream. That is to say, through frame reduction, the number of original audio frames in the audio stream is reduced.
[0168] As Figure 7 shown in S704, after obtaining the character probability distribution of the multiple audio frames, the character probability distribution of the multiple audio frames and the streaming encoding features of multiple audio frames in the audio stream can be input into a frame reducer, and through this frame reducer, according to the character probability distribution of the multiple audio frames, frame reduction is performed on the streaming encoding features of multiple audio frames in the audio stream, to obtain the streaming encoding features of the multiple audio frames after frame reduction. An implementation manner of performing frame reduction based on the character probability distribution of each audio frame is provided. In this way, the amount of information referred to in frame reduction is increased, and the effect of frame reduction can be effectively improved.
[0169] In some possible implementation manners, the above frame reducer may include a convolution module and a frame selection module. Among them, the convolution module provides the function of feature extraction and has a strong local feature extraction ability. In the embodiment of the present application, the convolution module is used to perform convolution processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolution processing. The frame selection module provides the function of selecting audio frames to reduce the number of audio frames input to the non-streaming encoder. In the embodiment of the present application, the frame selection module is used to extract the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions from the streaming encoding features of the multiple audio frames after convolution processing, to obtain the streaming encoding features of the multiple audio frames after frame reduction.
[0170] Correspondingly, for the above-mentioned process of performing downsampling processing on the streaming encoding features of multiple audio frames in the audio stream according to the character probability distribution of the multiple audio frames based on the downsampler, the following (1-1) to (1-2) can be referred to:
[0171] (1-1) The electronic device performs convolution processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolution processing.
[0172] In some possible implementation manners, the electronic device performs convolution processing on the streaming encoding features of the multiple audio frames based on a convolution module to obtain the streaming encoding features of the multiple audio frames after convolution processing.
[0173] Exemplarily, the convolution module may include a depth wise convolution unit and a point wise convolution unit. Among them, the depth wise convolution unit is used to independently perform convolution operations on each channel of the input layer. In this way, a low-power feature extraction process can be implemented based on the depth wise convolution unit, thereby improving the efficiency of feature extraction. The point wise convolution unit is used to fuse information at the same spatial position of different channels. In this way, information between different channels can be fused, thereby increasing the amount of information referred to in speech recognition and thus improving the accuracy of speech recognition.
[0174] In some possible implementation manners, the number of depth wise convolution units may be one or more. In some other possible implementation manners, the number of point wise convolution units may be one or more.
[0175] Exemplarily, Figure 9 is a schematic framework diagram of a convolution module provided by an embodiment of the present application. Refer to Figure 9 , which shows a convolution module including one depth wise convolution unit and two point wise convolution units. Taking the timing feature H T s = [h1 s , h2 s , …, h T s output by the streaming encoder as an example, in the convolution module shown in Figure 9 , the timing feature H T s = [h1 s , h2 s , …, h T sAfter passing through a point-wise convolution layer, then a depth-wise convolution layer, and then another point-wise convolution layer, the local information-enhanced temporal feature H can be finally obtained. T s ' = [h1 s ', h2 s ', …, h T s ']。
[0176] (1-2) The electronic device extracts the streaming encoding features of the audio frames whose character probability distribution satisfies the preset condition from the streaming encoding features of the multiple audio frames after the convolution processing, and obtains the streaming encoding features of the multiple audio frames after the frame reduction processing.
[0177] In some possible implementation manners, the preset condition may be that the probability value of the preset character is less than the preset threshold. Among them, the preset character may be at least one of a blank character or an interval character. Exemplarily, taking the blank character as an example, the preset character may be a connectionist temporal classification (CTC) blank character.
[0178] The preset threshold may be a preset fixed threshold, such as 0.99, 0.999 or other thresholds. It should be noted that if the probability value of the preset character of an audio frame is less than the preset threshold, it means that the probability of this audio frame corresponding to a blank character or an interval character is small, that is, this audio frame is very likely to generate other meaningful characters.
[0179] Correspondingly, the process of the above-mentioned electronic device extracting the streaming encoding features of the audio frames whose character probability distribution satisfies the preset condition may be: determining the audio frames whose probability value of the preset character is less than the preset threshold according to the character probability distribution of the multiple audio frames. Then, extracting the streaming encoding features of the audio frames whose probability value of the preset character is less than the preset threshold from the streaming encoding features of the multiple audio frames, and obtaining the streaming encoding features of the audio frames whose character probability distribution satisfies the preset condition. That is to say, extracting the streaming encoding features of the audio frames that are very likely to generate other meaningful characters from the streaming encoding features of the multiple audio frames realizes the extraction of key audio frames or important audio frames, and thus realizes the frame reduction processing of the audio frames.
[0180] Exemplarily, Figure 10 is a schematic framework diagram of a frame reducer provided by an embodiment of the present application. Refer to Figure 10 , the frame reducer may include a convolution module 1001 and a frame selection module 1002. The process of performing frame reduction processing based on Figure 10 the shown frame reducer may be: First, the temporal feature H output by the streaming encoder Ts = [h1 s , h2 s , …, h T s Input the convolutional module 1001, and through the convolutional module 1001, perform convolution processing on the time series feature H T s = [h1 s , h2 s , …, h T s to obtain the time series feature H after local information enhancement T s ' = [h1 s ', h2 s ', …, h T s ']. Furthermore, input the time series feature H T s ' = [h1 s ', h2 s ', …, h T s '] and the frame-level probability distribution P T s = [p1 s , p2 s , …, p T s into the frame selection module 1002. Through the frame selection module 1002, extract the streaming coding features of the audio frames whose character probability distributions meet the preset conditions, and thus obtain the streaming coding features H T' s ' = [h1 s , h2 s , …, h T' s '] of multiple audio frames after frame reduction processing. In this example, the streaming coding features of multiple audio frames after frame reduction processing can be represented by H T' s ' = [h1 s , h2 s , …, h T' s ']. It should be understood that T' refers to the number of audio frames after frame reduction processing, and T' is a positive integer greater than 0 and less than T.
[0181] In the above embodiments, a frame rate reduction processing scheme based on convolution and frame selection is provided. By first performing convolution processing on the streaming encoding features of the multiple audio frames, time series features with enhanced local information can be obtained. Then, from the time series features with enhanced local information, the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions are extracted, so that the streaming encoding features of the extracted audio frames are also time series features with enhanced local information, which can effectively improve the expression ability of the features, thereby improving the accuracy of speech recognition.
[0182] In some other possible implementation manners, the above frame rate reducer may include a convolution module, a frame selection module, and an attention module. Correspondingly, based on the frame rate reducer, according to the character probability distribution of the multiple audio frames, the process of performing frame rate reduction processing on the streaming encoding features of the multiple audio frames in the audio stream to obtain the streaming encoding features of the multiple audio frames after frame rate reduction processing can be referred to the following (2-1) to (2-3):
[0183] (2-1) The electronic device performs convolution processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolution processing.
[0184] It should be noted that the relevant content of (2-1) can be referred to (1-1) and will not be elaborated here.
[0185] (2-2) The electronic device extracts the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions from the streaming encoding features of the multiple audio frames after convolution processing.
[0186] It should be noted that the relevant content of (2-2) can be referred to (1-2) and will not be elaborated here.
[0187] (2-3) The electronic device performs attention calculation on the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions based on the attention mechanism to obtain the streaming encoding features of the multiple audio frames after frame rate reduction processing.
[0188] In some possible implementation manners, the process of the above electronic device performing attention calculation may be: performing matrix multiplication processing on the streaming encoding features of the multiple audio frames after convolution processing and the streaming encoding features of the multiple audio frames after frame rate reduction processing to obtain a similarity matrix. Performing Scale (normalization) processing on the similarity matrix to obtain the similarity matrix after Scale processing. Performing normalization processing on the similarity matrix after Scale processing to obtain the normalized similarity matrix. Performing matrix multiplication processing on the streaming encoding features of the multiple audio frames after convolution processing and the normalized similarity matrix to obtain the streaming encoding features of the multiple audio frames after frame rate reduction processing.
[0189] Among them, the similarity matrix represents the similarity between the streaming encoding features of multiple audio frames after convolution processing and the streaming encoding features of multiple audio frames after downsampling processing. Scale processing refers to dividing each element in the similarity matrix by the dimension size of the global feature, such as the dimension size of the streaming encoding features of multiple audio frames after convolution processing. The value of each element in the normalized similarity matrix is a numerical value between (0, 1). At this time, the normalized similarity matrix can be used as a weight matrix to perform matrix multiplication with the streaming encoding features of multiple audio frames after convolution processing, thereby completing the attention calculation.
[0190] Exemplarily, Figure 11 is a schematic framework diagram of an attention module provided by an embodiment of the present application. Refer to Figure 11 , the attention module may include a first matrix multiplication unit, a normalization unit, a normalization unit, and a second matrix multiplication unit. In Figure 11 the shown attention module, Q is used to refer to the streaming encoding features of multiple audio frames after downsampling processing, and K and V are used to refer to the streaming encoding features of multiple audio frames after convolution processing. It should be understood that K and V are exactly the same and are used to refer to the temporal features including global information.
[0191] Among them, the first matrix multiplication unit is used to perform matrix multiplication on the streaming encoding features of multiple audio frames after convolution processing and the streaming encoding features of multiple audio frames after downsampling processing to obtain a similarity matrix. The normalization unit is used to perform Scale processing on the similarity matrix to obtain a Scale-processed similarity matrix. The normalization unit is used to perform normalization processing on the Scale-processed similarity matrix to obtain a normalized similarity matrix. The second matrix multiplication unit is used to perform matrix multiplication on the streaming encoding features of multiple audio frames after convolution processing and the normalized similarity matrix to obtain the streaming encoding features of multiple audio frames after downsampling processing. In this way, considering that the local characteristics of convolution are likely to limit the capture of global information, through attention calculation, the streaming encoding features of multiple audio frames including global information can be combined to comprehensively generate the final temporal features, which can make up for the information loss problem caused by frame loss.
[0192] Exemplarily, Figure 12 is a schematic framework diagram of another downsampler provided by an embodiment of the present application. Refer to Figure 12 , the downsampler may include a convolution module 1201, a frame selection module 1202, and an attention module 1203. Based on Figure 12 the shown downsampler to perform the downsampling process can be: First, the temporal feature H T s = [h1 s , h2s ,…,h T s Input the convolutional module 1201. Through the convolutional module 1201, convolve the sequential feature H T s =[h1 s ,h2 s ,…,h T s to obtain the sequential feature H with enhanced local information T s '=[h1 s ',h2 s ',…,h T s ']. At this time, the number of the sequential feature H T s '=[h1 s ',h2 s ',…,h T s ' can be two. One can be referred to as K, and the other can be referred to as V. Furthermore, input the sequential feature H T s '=[h1 s ',h2 s ',…,h T s ' and the character probability distribution P T s =[p1 s ,p2 s ,…,p T s output by the linear network into the frame selection module 1202. Through the frame selection module 1202, extract the streaming encoding feature H T' s '=[h1 s ,h2 s ,…,h T' s '] of the audio frames whose character probability distribution meets the preset conditions. At this time, the streaming encoding feature H T' s '=[h1 s ,h2 s ,…,h T' s ' of the audio frames whose character probability distribution meets the preset conditions can be referred to as Q. Furthermore, Q, K, and V can be input into the attention module. Through the attention calculation of this attention module, a sequential sequence H T' s ”=[h1 s ”,h2s ”, …, h T' s ”]. In this example, the streaming encoding features of multiple audio frames after downsampling processing can be represented by H T' s ” = [h1 s ”, h2 s ”, …, h T' s ”].
[0193] In the above embodiments, a downsampling processing scheme based on convolution, frame selection, and attention mechanism is provided. By first performing convolution processing on the streaming encoding features of the multiple audio frames, the temporal features with enhanced local information can be obtained. Then, from the temporal features with enhanced local information, the streaming encoding features of the audio frames whose character probability distribution satisfies a preset condition are extracted, so that the extracted streaming encoding features of the audio frames are also the temporal features with enhanced local information, which can effectively improve the expression ability of the features. Furthermore, by performing attention calculation on the streaming encoding features of the audio frames whose character probability distribution satisfies the preset condition to capture global information, the problem of information loss caused by frame loss can be compensated, so that lossless low-latency speech recognition can be achieved.
[0194] The above S704 corresponds to Figure 6 the process of performing downsampling processing on the streaming encoding features of multiple audio frames in the audio stream shown in S603 in
[0195] S705. The electronic device performs non-streaming encoding on the streaming encoding features of the multiple audio frames after downsampling processing based on the non-streaming encoder to obtain the non-streaming encoding features of the audio stream.
[0196] As Figure 7 shown in S705, after obtaining the streaming encoding features of the multiple audio frames after downsampling processing, the streaming encoding features of the multiple audio frames after downsampling processing can be input into the non-streaming encoder, and through the non-streaming encoder, non-streaming encoding is performed on the streaming encoding features of the multiple audio frames after downsampling processing to obtain the non-streaming encoding features of the audio stream.
[0197] Among them, the non-streaming encoding features are used to represent the audio features of the audio stream. It should be understood that non-streaming means encoding is performed after all inputs are completed. For example, in some possible implementation manners, after obtaining the streaming encoding features of all audio frames in the audio stream, the streaming encoding features of all audio frames in the audio stream can be input into the non-streaming encoder, and non-streaming encoding is performed on the streaming encoding features of all audio frames to obtain the non-streaming encoding features of the audio stream.
[0198] S706. The electronic device determines the output probability distribution of the audio stream based on the non-streaming coding features of the audio stream and the text prediction results of multiple audio frames in the audio stream.
[0199] Among them, the text prediction result is predicted based on the speech recognition result of the previous audio frame of the audio frame. In some possible implementation manners, the electronic device may input the speech recognition result of the previous audio frame of the audio frame into the prediction network of the speech recognition model, and the prediction network makes a prediction on the current audio frame based on the speech recognition result of the previous audio frame to obtain the text prediction result of the audio frame. In this way, based on the non-streaming coding features of the audio stream, the text prediction results of all audio frames in the audio stream are also combined to determine the output probability distribution of the audio stream, which can improve the accuracy of speech recognition. It should be understood that if the current audio frame is the first audio frame of the audio stream, the default text may be input into the prediction network of the speech recognition model to obtain the text prediction result of the audio frame. For example, the default text may be the text with a higher occurrence frequency in the speech recognition result, such as hot words.
[0200] The output probability distribution represents the probability values of the multiple audio frames on the multiple candidate texts corresponding to the text prediction results. For example, the output probability distribution may include the probability values of multiple candidate texts corresponding to all audio frames in the audio stream. Exemplarily, taking audio frame M as an example, the corresponding multiple candidate texts may be m1 (probability value 0.9), m2 (probability value 0.7), m3 (probability value 0.3), and so on. It should be understood that if the probability value of an audio frame on a certain candidate text is large, it means that the audio frame is very likely to correspond to the candidate text.
[0201] As Figure 7 shown in S706, after obtaining the non-streaming coding features of the audio stream and the text prediction results of multiple audio frames in the audio stream, the non-streaming coding features of the audio stream and the text prediction results of multiple audio frames in the audio stream may be input into the fusion network, and the fusion network performs a fusion process on the non-streaming coding features of the audio stream and the text prediction results of multiple audio frames in the audio stream to obtain the output probability distribution of the audio stream. Among them, the fusion process may be an accumulation process.
[0202] In this way, using the fusion network of the speech recognition model, the non-streaming coding features of the audio stream and the text prediction results of multiple audio frames in the audio stream are fused to comprehensively refer to the non-streaming coding features of the audio stream and the text prediction results of multiple audio frames in the audio stream, which can not only ensure the efficiency of speech recognition but also improve the accuracy of speech recognition.
[0203] It should be noted that the electronic device may first obtain the non-streaming coding features of the audio stream, and then obtain the text prediction result of the audio stream. Alternatively, the electronic device may first obtain the text prediction result of the audio stream, and then obtain the non-streaming coding features of the audio stream. Alternatively, the electronic device may also obtain the non-streaming coding features of the audio stream and the text prediction result of the audio stream synchronously. The embodiments of the present application do not limit the order of obtaining the non-streaming coding features of the audio stream and the text prediction result of the audio stream.
[0204] S707. The electronic device decodes the output probability distribution of the audio stream to obtain the speech recognition result of the audio stream.
[0205] Among them, the speech recognition result of the audio stream may be the text recognition result corresponding to the audio stream, such as a piece of text. Further, in some possible implementation manners, the speech recognition result of the audio stream may be displayed on the electronic device. In this way, the speech recognition result can be timely displayed for the user when the user input audio stream ends, improving the human-computer interaction efficiency.
[0206] The above S706 to S707 correspond to Figure 6 the process of obtaining the speech recognition result of the audio stream based on the non-streaming coding features of the audio stream shown in S605.
[0207] In one example, during the process that the electronic device receives the audio stream [x1, x2, …, x T input by the user, a frame of audio x1 received by the electronic device is input into the streaming encoder, and the streaming encoder performs streaming coding on the input frame of audio x1 to obtain the streaming coding feature h1 s of the frame of audio x1, and caches the streaming coding feature h1 s of the frame of audio x1. The above process of the streaming encoder is looped until the input of the audio stream ends, and the streaming coding features H T s = [h1 s , h2 s , …, h T s of multiple frames of audio in the audio stream can be cached. And the text prediction results of multiple frames of audio in the audio stream are obtained through the prediction network. Furthermore, all the cached streaming coding features H T s = [h1 s , h2 s , …, h T s are respectively input into the linear network and the downsampler, and the linear network processes the streaming coding features H T s = [h1 s, h2 s , …, h T s are normalized to obtain the character probability distribution P of the multiple audio frames T s = [p1 s , p2 s , …, p T s . Through the frame reducer, according to the character probability distribution P of the multiple audio frames T s = [p1 s , p2 s , …, p T s , the streaming encoding features H of the multiple audio frames in the audio stream T s = [h1 s , h2 s , …, h T s are frame-reduced to obtain the streaming encoding features H of the multiple audio frames after frame reduction T' s ” = [h1 s ”, h2 s ”, …, h T' s ”]. Furthermore, the streaming encoding features H of the multiple audio frames after frame reduction T' s ” = [h1 s ”, h2 s ”, …, h T' s ”] are input into the non-streaming encoder. Through the non-streaming encoder, the streaming encoding features H of the multiple audio frames after frame reduction T' s ” = [h1 s ”, h2 s ”, …, h T' s ”] are non-streaming encoded to obtain the non-streaming encoding features H of the audio stream T' ns ” = [h1 ns ”, h2 ns ”, …, h T' ns ”]. The non-streaming encoding features H of the audio stream T' ns ” = [h1 ns ”, h2 ns ”, …, h T' ns”]The text prediction results of multiple frames of audio in the audio stream are input into the fusion network frame by frame. The output probability distribution of each frame of audio is calculated frame by frame through the fusion network, so as to obtain the output probability distribution of the audio stream. Then, the output probability distribution of the audio stream is decoded to obtain the speech recognition result of the audio stream, and the final speech recognition result is displayed in the electronic device.
[0208] It should be noted that the above Figure 8 shown speech recognition model can also perform speech recognition on the audio stream received by the electronic device only based on the streaming encoder. For example, in one example, during the process when the electronic device receives the audio stream [x1, x2, …, x T , a frame of audio x1 received by the electronic device is input into the streaming encoder, and the input frame of audio x1 is stream-encoded through the streaming encoder to obtain the stream-encoded feature h1 of the frame of audio x1 s . And the text prediction result of the frame of audio x1 is obtained through the prediction network. Furthermore, the stream-encoded feature h1 of the frame of audio x1 s and the text prediction result of the frame of audio x1 are input into the fusion network, and the output probability distribution of the frame of audio x1 is calculated through the fusion network. Then, the output probability distribution of the frame of audio x1 is decoded to obtain the speech recognition result of the frame of audio x1, and the speech recognition result is displayed in the electronic device in real time. In this way, by repeatedly executing the above process of performing speech recognition on each audio frame based on the streaming encoder, the real-time display of the speech recognition result can be achieved.
[0209] It should also be noted that in some other possible implementation manners, speech recognition can also be performed on the audio stream received by the electronic device only based on the non-streaming encoder. For example, in one example, after the electronic device receives the user input audio stream [x1, x2, …, x T , the audio stream [x1, x2, …, x T is input into the linear network and the frame downsampler. The audio stream [x1, x2, …, x T is normalized through the linear network to obtain the character probability distribution of multiple audio frames in the audio stream. Through the frame downsampler, according to the character probability distribution of the multiple audio frames, the audio stream [x1, x2, …, x TPerform frame rate reduction processing to obtain the audio stream after frame rate reduction processing. Furthermore, input the audio stream after frame rate reduction processing into a non-streaming encoder. Through the non-streaming encoder, perform non-streaming encoding on the audio stream after frame rate reduction processing to obtain the non-streaming encoding features of the audio stream. Input the non-streaming encoding features of the audio stream and the text prediction results of multiple frames of audio in the audio stream into the fusion network frame by frame. Through the fusion network, calculate the output probability distribution of each frame of audio frame by frame, so as to obtain the output probability distribution of the audio stream. Then decode the output probability distribution of the audio stream to obtain the speech recognition result of the audio stream, and display the final speech recognition result in the electronic device.
[0210] The technical solution provided by the embodiments of the present application, after receiving the audio stream input by the user, can input the audio stream into the streaming encoder according to audio frames. Through the streaming encoder, encode the input audio frames to obtain the streaming encoding features of multiple audio frames in the audio stream. Then input the streaming encoding features of multiple audio frames in the audio stream into the frame rate reducer. Through the frame rate reducer, perform frame rate reduction processing on the streaming encoding features of multiple audio frames in the audio stream to obtain the streaming encoding features of multiple audio frames after frame rate reduction processing. Furthermore, input the streaming encoding features of multiple audio frames after frame rate reduction processing into the non-streaming encoder. Through the non-streaming encoder, encode the streaming encoding features of multiple audio frames after frame rate reduction processing to obtain the non-streaming encoding features of the audio stream. Finally, based on the non-streaming encoding features of the audio stream, obtain the speech recognition result of the audio stream. In this way, by setting the frame rate reducer to perform frame rate reduction processing on the streaming encoding features of multiple audio frames in the audio stream, the length of the feature sequence output by the streaming encoder can be effectively shortened, and thus the length of the feature sequence received by the non-streaming encoder can be effectively shortened. Furthermore, the computing amount of the non-streaming encoder is reduced, the speech recognition delay can be effectively reduced, and thus the speech recognition efficiency can be effectively improved, and the user experience is improved.
[0211] In view of the above Figure 7 For the speech recognition model adopted, before implementing this solution, it is also necessary to perform iterative training on the initial model based on audio training data to obtain the speech recognition model. Figure 13 It is a schematic flowchart of a training method for a speech recognition model provided by an embodiment of the present application. Refer to Figure 13 This method includes the following S1301-S1302:
[0212] S1301. Iteratively train the initial model based on audio training data. During any iterative training process, input the audio training data into the model obtained after the previous iterative training, perform streaming encoding on multiple audio frames in the audio sample to obtain the streaming encoding features of multiple audio frames in the audio sample; perform frame reduction processing on the streaming encoding features of multiple audio frames in the audio sample to obtain the streaming encoding features of multiple audio frames after frame reduction processing; perform non-streaming encoding on the streaming encoding features of multiple audio frames after frame reduction processing to obtain the non-streaming encoding features of the audio sample; obtain the speech recognition result of the audio sample based on the non-streaming encoding features of the audio sample; determine the model loss value based on the speech recognition result and the annotated text; and adjust the model parameters based on the model loss value.
[0213] Among them, the audio training data includes the audio sample and the annotated text of the audio sample. That is, perform iterative training of the initial model based on the audio sample and the annotated text of the audio sample to obtain a speech recognition model with a lower speech recognition latency.
[0214] The number of frames of multiple audio frames after frame reduction processing is less than the number of frames of multiple audio frames in the audio sample. That is, through frame reduction processing, the number of original audio frames in the audio sample is reduced. Furthermore, since frame reduction processing reduces the number of original audio frames in the audio sample, the number of audio frames to be non-streaming encoded is also reduced, which can effectively shorten the length of the feature sequence to be non-streaming encoded.
[0215] S1302. At the end of the iterative training, obtain the trained model as the speech recognition model.
[0216] The technical solution provided by the embodiments of the present application obtains a speech recognition model with a lower speech recognition latency through iterative training of the initial model. Among them, during any iterative training process, by performing frame reduction processing on the streaming encoding features of multiple audio frames in the audio sample, the number of original audio frames in the audio sample is reduced, and thus the number of audio frames to be non-streaming encoded is also reduced, which can effectively shorten the length of the feature sequence to be non-streaming encoded, and further reduce the amount of non-streaming encoding operations, effectively reducing the speech recognition latency, and thus effectively improving the speech recognition efficiency. Furthermore, adjusting the model parameters according to the model loss value can improve the learning ability of the model, thereby training a speech recognition model with better learning ability.
[0217] Figure 14 It is a schematic flowchart of a method for training a speech recognition model provided by an embodiment of the present application. Refer to Figure 14 , taking the process of any iterative training in model training as an example, the method includes the following S1401-S1410:
[0218] S1401. During any iteration of training, the electronic device inputs the audio training data into the model obtained after the previous iteration of training.
[0219] The audio training data refers to the training data of the initial model, including audio samples and the annotation text of the audio samples.
[0220] Exemplarily, Figure 15 is a schematic diagram of the model framework in the model training stage provided by an embodiment of the present application. Refer to Figure 15 In Figure 15 the model shown, a cascaded encoder including a streaming encoder, a frame rate reducer, and a non-streaming encoder, a linear network (or referred to as a linear layer), a prediction network, and a fusion network are adopted.
[0221] During the model training stage, the streaming encoder is used to process the feature sequence of audio frames in the audio sample to extract effective features from the feature sequence of audio frames. The output result of the streaming encoder can be used as the input of the linear network, and the linear network is used to normalize the output result of the streaming encoder to obtain the frame-level probability distribution. In the embodiments of the present application, the character probability distribution is subsequently used to refer to this frame-level probability distribution.
[0222] The output result of the linear network and the output result of the streaming encoder can both be used as the input of the frame rate reducer. The frame rate reducer is used to perform frame rate reduction processing on the output result of the streaming encoder based on the output result of the linear network. And during the model training stage, the linear network is also used to obtain the CTC loss value of the non-streaming encoder.
[0223] The output result of the frame rate reducer can be used as the input of the non-streaming encoder. Correspondingly, the non-streaming encoder is used to process the output result of the frame rate reducer, that is, to process the output result of the streaming encoder after frame rate reduction processing, to further extract the effective features of the audio sample.
[0224] The prediction network is used to predict future text based on the historical speech recognition result (i.e., the recognized text) of the audio sample. The fusion network provides a function of feature fusion and is used to perform fusion processing on the output result of the cascaded encoder and the output result of the prediction network to obtain the final speech recognition result. And during the model training stage, the fusion network is also used to obtain the recurrent neural networks-Transducer (RNNT) loss value of the non-streaming encoder.
[0225] Next, based on Figure 15 the model shown, the training process of the speech recognition model will be described.
[0226] S1402. The electronic device performs streaming encoding on multiple audio frames in the audio sample based on the streaming encoder, and obtains the streaming encoding features of the multiple audio frames in the audio sample.
[0227] As Figure 14 shown in S1402, after inputting the audio training data into the model obtained from the previous iterative training, the audio sample is input into the streaming encoder frame by frame. The streaming encoder performs streaming encoding on the input audio frames, and obtains the streaming encoding features of the multiple audio frames in the audio sample.
[0228] Among them, the streaming encoding features are used to represent the audio features at the frame level, that is, the audio features of each audio frame.
[0229] S1403. The electronic device performs normalization processing on the streaming encoding features of the multiple audio frames, and obtains the character probability distribution of the multiple audio frames.
[0230] Among them, the normalization processing refers to limiting the streaming encoding features of each audio frame to a value between (0, 1). The character probability distribution represents the probability values of the multiple audio frames on multiple candidate characters respectively.
[0231] In some possible implementation manners, the electronic device performs normalization processing on the streaming encoding features of the multiple audio frames based on a linear network, and obtains the character probability distribution of the multiple audio frames.
[0232] As Figure 14 shown in S1403, after obtaining the streaming encoding features of the multiple audio frames in the audio sample, the streaming encoding features of the multiple audio frames in the audio sample are input into the linear network. Through the linear network, the streaming encoding features of the multiple audio frames are normalized, and the character probability distribution of the multiple audio frames is obtained.
[0233] In the above embodiment, by setting up a linear network to normalize the streaming encoding features of the multiple audio frames output by the streaming encoder, the character probability distribution of the multiple audio frames can be obtained, so as to use the character probability distribution of the multiple audio frames to perform the process of frame reduction subsequently.
[0234] S1404. The electronic device performs frame reduction on the streaming encoding features of the multiple audio frames in the audio sample based on a frame reducer according to the character probability distribution of the multiple audio frames, and obtains the streaming encoding features of the multiple audio frames after frame reduction.
[0235] Among them, the number of frames of the multiple audio frames after frame reduction is less than the number of frames of the multiple audio frames in the audio sample. That is to say, through frame reduction, the number of original audio frames in the audio sample is reduced.
[0236] AsFigure 14 As shown in S1404, after obtaining the character probability distributions of the multiple audio frames, the character probability distributions of the multiple audio frames and the streaming encoding features of the multiple audio frames in the audio sample can be input into a downsampler. Through the downsampler, according to the character probability distributions of the multiple audio frames, the streaming encoding features of the multiple audio frames in the audio sample are downsampled to obtain the downsampled streaming encoding features of the multiple audio frames. An implementation method of performing downsampling based on the character probability distribution of each audio frame is provided. In this way, the amount of information referred to in the downsampling process is increased, and the effect of the downsampling process can be effectively improved.
[0237] In some possible implementation manners, the above downsampler may include a convolution module and a frame selection module. Correspondingly, the process of the above downsampler downsampling the streaming encoding features of the multiple audio frames in the audio sample according to the character probability distributions of the multiple audio frames to obtain the downsampled streaming encoding features of the multiple audio frames can be seen in the following (3-1) to (3-2):
[0238] (3-1) The electronic device performs convolution processing on the streaming encoding features of the multiple audio frames to obtain the convolution-processed streaming encoding features of the multiple audio frames.
[0239] It should be noted that the relevant content of (3-1) can be seen in (1-1) and will not be elaborated here.
[0240] (3-2) The electronic device extracts the streaming encoding features of the audio frames whose character probability distributions meet the preset conditions from the convolution-processed streaming encoding features of the multiple audio frames to obtain the downsampled streaming encoding features of the multiple audio frames.
[0241] It should be noted that the relevant content of (3-2) can be seen in (1-2) and will not be elaborated here.
[0242] In the above embodiment, a downsampling processing scheme based on convolution and frame selection is provided. By first performing convolution processing on the streaming encoding features of the multiple audio frames, the temporal features with enhanced local information can be obtained. Then, from the temporal features with enhanced local information, the streaming encoding features of the audio frames whose character probability distributions meet the preset conditions are extracted, so that the extracted streaming encoding features of the audio frames are also the temporal features with enhanced local information, which can effectively improve the expression ability of the features, and thus can improve the accuracy of speech recognition.
[0243] In some other possible implementation manners, the above downsampler may include a convolution module, a frame selection module, and an attention module. Correspondingly, based on the downsampler, according to the character probability distribution of the multiple audio frames, the process of performing downsampling on the streaming encoded features of the multiple audio frames in the audio sample to obtain the streaming encoded features of the multiple audio frames after downsampling can be referred to the following (4-1) to (4-3):
[0244] (4-1) Perform convolution processing on the streaming encoded features of the multiple audio frames to obtain the streaming encoded features of the multiple audio frames after convolution processing.
[0245] It should be noted that the relevant content of (4-1) can be referred to (1-1) and will not be elaborated here.
[0246] (4-2) Extract the streaming encoded features of the audio frames whose character probability distribution meets the preset conditions from the streaming encoded features of the multiple audio frames after convolution processing.
[0247] It should be noted that the relevant content of (4-2) can be referred to (1-2) and will not be elaborated here.
[0248] (4-3) Based on the attention mechanism, perform attention calculation on the streaming encoded features of the audio frames whose character probability distribution meets the preset conditions to obtain the streaming encoded features of the multiple audio frames after downsampling.
[0249] It should be noted that the relevant content of (4-3) can be referred to (2-3) and will not be elaborated here.
[0250] In the above embodiment, a downsampling processing scheme based on convolution, frame selection, and attention mechanism is provided. By first performing convolution processing on the streaming encoded features of the multiple audio frames, the temporal features with enhanced local information can be obtained. Then, from the temporal features with enhanced local information, the streaming encoded features of the audio frames whose character probability distribution meets the preset conditions are extracted, so that the extracted streaming encoded features of the audio frames are also the temporal features with enhanced local information, which can effectively improve the expression ability of the features. Furthermore, by performing attention calculation on the streaming encoded features of the audio frames whose character probability distribution meets the preset conditions to capture global information, the problem of information loss caused by frame dropping can be compensated, thereby enabling lossless performance low-latency speech recognition.
[0251] The above S1404 corresponds to Figure 13 the process of performing downsampling on the streaming encoded features of the multiple audio frames in the audio sample to obtain the streaming encoded features of the multiple audio frames after downsampling as shown in S1301.
[0252] S1405. The electronic device performs non-streaming encoding on the streaming encoding features of the multiple downsampled audio frames based on the non-streaming encoder to obtain the non-streaming encoding features of the audio sample.
[0253] As Figure 14 shown in S1405, after obtaining the streaming encoding features of the multiple downsampled audio frames, the streaming encoding features of the multiple downsampled audio frames can be input into the non-streaming encoder. Through the non-streaming encoder, non-streaming encoding is performed on the streaming encoding features of the multiple downsampled audio frames to obtain the non-streaming encoding features of the audio sample. Among them, the non-streaming encoding features are used to represent the audio features of the audio sample.
[0254] S1406. The electronic device determines the output probability distribution of the audio sample based on the non-streaming encoding features of the audio sample and the text prediction results of multiple audio frames in the audio sample.
[0255] Among them, the text prediction result is predicted based on the speech recognition result of the previous audio frame of the audio frame. In some possible implementation manners, the electronic device can input the speech recognition results of the previous audio frames of each audio frame into the prediction network of the model. Through the prediction network, the current audio frame is predicted based on the speech recognition result of the previous audio frame to obtain the text prediction results of each audio frame. In this way, by setting a prediction network in the model and using the prediction network to predict the text prediction results of the current audio frame, and then combining the text prediction results of the current audio frame, the output probability distribution of the audio sample is determined. The output probability distribution represents the probability values of the multiple audio frames on the multiple candidate texts corresponding to the text prediction results respectively.
[0256] As Figure 14 shown in S1406, after obtaining the non-streaming encoding features of the audio sample and the text prediction results of multiple audio frames in the audio sample, the non-streaming encoding features of the audio sample and the text prediction results of multiple audio frames in the audio sample are input into the fusion network. Through the fusion network, the non-streaming encoding features of the audio sample and the text prediction results of multiple audio frames in the audio sample are fused to obtain the output probability distribution of the audio sample. Among them, the fusion process can be an accumulation process.
[0257] In this way, by using the fusion network of the model, the non-streaming encoding features of the audio sample and the text prediction results of multiple audio frames in the audio sample are fused to comprehensively refer to the non-streaming encoding features of the audio sample and the text prediction results of multiple audio frames in the audio sample, which can not only ensure the efficiency of speech recognition but also improve the accuracy of speech recognition.
[0258] S1407. The electronic device decodes the output probability distribution of the audio sample to obtain the speech recognition result of the audio sample.
[0259] Among them, the speech recognition result of the audio sample may be the text recognition result corresponding to the audio sample, such as a piece of text.
[0260] The above S1406 to S1407 correspond to Figure 13 the process of obtaining the speech recognition result of the audio sample based on the non-streaming encoding features of the audio sample shown in S1301 in
[0261] S1408. The electronic device determines the model loss value based on the speech recognition result and the labeled text.
[0262] Among them, the labeled text may be a manually labeled text tag. It should be understood that the model loss value is used to represent the difference between the speech recognition result determined based on the model and the labeled text.
[0263] In some possible implementation manners, the electronic device determines the RNNT loss value of the non-streaming encoder based on the speech recognition result and the labeled text, and uses it as the model loss value.
[0264] Exemplarily, the electronic device may input the speech recognition result and the labeled text into the fusion network of the model, and determine the RNNT loss value of the non-streaming encoder through the fusion network. In this way, by determining the RNNT loss value of the non-streaming encoder, the loss value of the non-streaming encoder can be quickly determined.
[0265] Or, in some other possible implementation manners, the electronic device determines the CTC loss value of the non-streaming encoder based on the speech recognition result and the labeled text, and uses it as the model loss value.
[0266] Exemplarily, the electronic device may input the speech recognition result and the labeled text into the linear network of the model, and determine the CTC loss value of the non-streaming encoder through the linear network. In this way, by determining the CTC loss value of the non-streaming encoder, the loss value of the non-streaming encoder can be quickly determined.
[0267] Or, in some other possible implementation manners, after the electronic device determines the RNNT loss value of the non-streaming encoder and the CTC loss value of the non-streaming encoder, it performs a weighted summation process on the RNNT loss value of the non-streaming encoder and the CTC loss value of the non-streaming encoder, and uses the result obtained from the weighted summation process as the model loss value.
[0268] Thus, based on the RNNT loss value of the non-streaming encoder and combining it with the CTC loss value of the non-streaming encoder, the loss value of the non-streaming encoder is comprehensively determined, increasing the amount of information referred to in determining the loss value and improving the accuracy of determining the loss value.
[0269] It should be noted that, in the above embodiments, the loss value of the non-streaming encoder is taken as an example to illustrate the solution. In some other embodiments, the model loss value can also be comprehensively determined by combining the loss value of the streaming encoder and the loss value of the non-streaming encoder. The corresponding process can be referred to in the following (5-1) to (5-3). It should be understood that the loss value of the streaming encoder represents the difference between the speech recognition result output by the streaming encoder and the labeled text. The loss value of the non-streaming encoder represents the difference between the speech recognition result output by the non-streaming encoder and the labeled text.
[0270] (5-1) The electronic device determines the loss value of the streaming encoder based on the speech recognition result output by the streaming encoder and the labeled text.
[0271] In some possible implementation manners, the electronic device determines the RNNT loss value of the streaming encoder based on the speech recognition result output by the streaming encoder and the labeled text, and uses it as the loss value of the streaming encoder.
[0272] Exemplarily, the electronic device can input the speech recognition result output by the streaming encoder and the labeled text into the fusion network of the model, and determine the RNNT loss value of the streaming encoder through the fusion network. Thus, by determining the RNNT loss value of the streaming encoder, the loss value of the streaming encoder can be quickly determined.
[0273] Or, in some other possible implementation manners, the electronic device also determines the CTC loss value of the streaming encoder based on the speech recognition result output by the streaming encoder and the labeled text, and uses it as the loss value of the streaming encoder.
[0274] Exemplarily, the electronic device can input the speech recognition result output by the streaming encoder and the labeled text into the linear network of the model, and determine the CTC loss value of the streaming encoder through the linear network. Thus, by determining the CTC loss value of the streaming encoder, the loss value of the streaming encoder can be quickly determined.
[0275] Or, in some other possible implementation manners, after the electronic device determines the RNNT loss value of the streaming encoder and the CTC loss value of the streaming encoder, it performs a weighted summation process on the RNNT loss value of the non-streaming encoder and the CTC loss value of the non-streaming encoder, and uses the result obtained from the weighted summation process as the loss value of the streaming encoder.
[0276] In this way, based on the RNNT loss value of the streaming encoder and combining the CTC loss value of the streaming encoder, the loss value of the streaming encoder is comprehensively determined, which increases the amount of information referred to in determining the loss value and improves the accuracy of determining the loss value.
[0277] (5-2) The electronic device determines the loss value of the non-streaming encoder based on the speech recognition result output by the non-streaming encoder and the labeled text.
[0278] In some possible implementation manners, the electronic device determines the RNNT loss value of the non-streaming encoder based on the speech recognition result and the labeled text, and uses it as the loss value of the non-streaming encoder.
[0279] Alternatively, in some other possible implementation manners, the electronic device determines the CTC loss value of the non-streaming encoder based on the speech recognition result and the labeled text, and uses it as the loss value of the non-streaming encoder.
[0280] Alternatively, in some other possible implementation manners, after the electronic device determines the RNNT loss value of the non-streaming encoder and the CTC loss value of the non-streaming encoder, it performs a weighted summation process on the RNNT loss value of the non-streaming encoder and the CTC loss value of the non-streaming encoder, and uses the result obtained from the weighted summation process as the loss value of the non-streaming encoder.
[0281] (5-3) The electronic device determines the model loss value based on the loss value of the streaming encoder and the loss value of the non-streaming encoder.
[0282] In some possible implementation manners, the electronic device performs a weighted summation process on the loss value of the streaming encoder and the loss value of the non-streaming encoder, and uses the result obtained from the weighted summation process as the model loss value.
[0283] It should be noted that, in another possible implementation manner, the electronic device can also use other implementation manners to obtain the model loss value. The embodiments of this application do not limit this.
[0284] S1409. The electronic device adjusts the model parameters based on the model loss value.
[0285] In one example, during any iteration training, the audio sample Y T =[y1,y2,…,y T can be input into the streaming encoder to obtain the streaming encoded feature, which can be denoted as H T s =[h1 s ,h2 s ,…,h T s . HT s = [h1 s , h2 s , …, h T s is input into a linear network, and the streaming encoded feature H T s = [h1 s , h2 s , …, h T s is normalized to obtain the character probability distribution P of these multiple audio frames T s = [p1 s , p2 s , …, p T s . And the CTC loss value of the streaming encoder is calculated through the linear network, which can be denoted as L CTC s . Then, H T s = [h1 s , h2 s , …, h T s is input into the fusion network, and the output result of the prediction network is fused through the fusion network to calculate the RNNT loss value of the streaming encoder, which can be denoted as L RNNT s .
[0286] Furthermore, H T s = [h1 s , h2 s , …, h T s is input into the downsampler, and through the downsampler, according to the character probability distribution P of these multiple audio frames T s = [p1 s , p2 s , …, p T s , the streaming encoded feature H of multiple audio frames in this audio sample T s = [h1 s , h2 s , …, h T s is downsampled to obtain the streaming encoded feature H of multiple audio frames after downsampling T' s ” = [h1 s ”, h2 s ”, …, h T's ”].
[0287] Furthermore, the streaming encoding features H of multiple audio frames after downsampling processing T' s ” = [h1 s ”, h2 s ”, …, h T' s ”] are input into a non-streaming encoder. Through the non-streaming encoder, the streaming encoding features H of multiple audio frames after downsampling processing T' s ” = [h1 s ”, h2 s ”, …, h T' s ”] are non-streaming encoded to obtain the non-streaming encoding features H of the audio sample T' ns ” = [h1 ns ”, h2 ns ”, …, h T' ns ”].
[0288] The non-streaming encoding features H of the audio sample T' ns ” = [h1 ns ”, h2 ns ”, …, h T' ns ”] are input into a linear network. Through the linear network, the CTC loss value of the non-streaming encoder is calculated and can be denoted as L CTC ns . Then, the non-streaming encoding features H of the audio sample T' ns ” = [h1 ns ”, h2 ns ”, …, h T' ns ”] are input into a fusion network. Through the fusion network, the output results of the prediction network are fused, and the RNNT loss value of the non-streaming encoder is calculated and can be denoted as L RNNT ns .
[0289] After obtaining the CTC loss value of the streaming encoder, the RNNT loss value of the streaming encoder, the CTC loss value of the non-streaming encoder, and the RNNT loss value of the non-streaming encoder, the model loss value is determined based on the following weighted summation formula (1). Furthermore, based on this model loss value, the model parameters are adjusted to complete this iterative training.
[0290] L = α1 * L CTC s + α2 * LRNNT s +α3*L CTC ns +α4*L RNNT ns (1)
[0291] In the formula, L represents the model loss value; L CTC s represents the CTC loss value of the streaming encoder; α1 represents the weight coefficient corresponding to the CTC loss value of the streaming encoder, such as 0.2; L RNNT s represents the RNNT loss value of the streaming encoder; α2 represents the weight coefficient corresponding to the RNNT loss value of the streaming encoder, such as 1.0; L CTC ns represents the CTC loss value of the non-streaming encoder; α3 represents the weight coefficient corresponding to the CTC loss value of the non-streaming encoder, such as 0.2; L RNNT ns represents the RNNT loss value of the non-streaming encoder; α4 represents the weight coefficient corresponding to the RNNT loss value of the non-streaming encoder, such as 1.0.
[0292] In this way, the model loss value is comprehensively determined according to the loss value of the streaming encoder and the loss value of the non-streaming encoder, and then the model parameters are adjusted by using the model loss value, which can improve the learning ability of the model, so as to train a speech recognition model with better learning ability.
[0293] After adjusting the model parameters, the electronic device also determines whether the iterative training meets the target conditions. Furthermore, in the case where the iterative training does not meet the target conditions, S1410 is executed. In the case where the iterative training meets the target conditions, the model obtained in the current iterative process is obtained as the speech recognition model.
[0294] In some possible implementation manners, the target conditions satisfy at least one of the following conditions: the number of iterations of the iterative training reaches the target number; or, the model loss value is less than or equal to the target threshold. Among them, the target number is the preset number of training iterations, such as the number of iterations reaches 100. The embodiments of the present application do not limit the setting of the target number. The target threshold is a preset fixed threshold, such as the model loss value is less than 0.0001. The embodiments of the present application do not limit the setting of the target threshold.
[0295] S1410. When the model after adjusting the model parameters does not meet the target conditions, the electronic device performs the next iterative training based on the model after adjusting the model parameters until the model meets the target conditions.
[0296] The technical solution provided by the embodiments of the present application obtains a speech recognition model with a lower speech recognition latency through iterative training of an initial model. Among them, during any iterative training process, by performing downsampling on the streaming coding features of multiple audio frames in an audio sample, the number of audio frames in the original audio sample is reduced, that is, the number of audio frames to be non-streaming coded is reduced, which can effectively shorten the length of the feature sequence to be non-streaming coded, thereby reducing the amount of non-streaming coding operations, effectively reducing the speech recognition latency, and thus effectively improving the speech recognition efficiency. Furthermore, by adjusting the model parameters according to the model loss value, the learning ability of the model can be improved, and thus a speech recognition model with better learning ability can be trained.
[0297] It should be noted that the electronic device used to execute the model training process above Figure 14 can be the same as the electronic device used to execute the model application process above Figure 7 or different from the electronic device used to execute the model application process above Figure 7 For example, in some possible implementation manners, the electronic device used to execute the model application process above Figure 7 can be a terminal device, while the electronic device used to execute the model training process above Figure 14 can be a server.
[0298] It can be understood that in order to implement the above functions, the electronic device (such as a terminal) in the embodiments of the present application includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0299] Figure 16 It is a schematic framework diagram of a speech recognition device provided by the embodiments of the present application. Refer to Figure 16 , the speech recognition device includes a receiving module 1601, a streaming coding module 1602, a downsampling module 1603, a non-streaming coding module 1604, and an obtaining module 1605. Among them,
[0300] The receiving module 1601 is used to receive the audio stream input by the user;
[0301] A streaming encoding module 1602 for performing streaming encoding on multiple audio frames in the audio stream to obtain streaming encoding features of the multiple audio frames in the audio stream;
[0302] A frame reduction module 1603 for performing frame reduction processing on the streaming encoding features of the multiple audio frames in the audio stream to obtain the streaming encoding features of the multiple audio frames after frame reduction processing, where the number of frames of the multiple audio frames after frame reduction processing is less than the number of frames of the multiple audio frames in the audio stream;
[0303] A non-streaming encoding module 1604 for performing non-streaming encoding on the streaming encoding features of the multiple audio frames after frame reduction processing to obtain non-streaming encoding features of the audio stream;
[0304] An acquisition module 1605 for obtaining a speech recognition result of the audio stream based on the non-streaming encoding features of the audio stream.
[0305] The technical solution provided by the embodiments of the present application reduces the number of frames of the original audio frames in the audio stream by performing frame reduction processing on the streaming encoding features of the multiple audio frames in the audio stream, thereby reducing the number of audio frames to be non-streaming encoded, effectively shortening the length of the feature sequence to be non-streaming encoded, further reducing the amount of non-streaming encoding operations, effectively reducing the latency of speech recognition, and thus effectively improving the efficiency of speech recognition and enhancing the user experience.
[0306] In some possible implementation manners, the device further includes a normalization module for:
[0307] Performing normalization processing on the streaming encoding features of the multiple audio frames to obtain a character probability distribution of the multiple audio frames, where the character probability distribution represents probability values of the multiple audio frames on multiple candidate characters respectively;
[0308] The frame reduction module 1603 is specifically configured to:
[0309] Perform frame reduction processing on the streaming encoding features of the multiple audio frames in the audio stream according to the character probability distribution of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after frame reduction processing.
[0310] In some possible implementation manners, the frame reduction module 1603 is specifically configured to:
[0311] Perform convolution processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolution processing;
[0312] Extract the streaming encoding features of the audio frames whose character probability distribution meets a preset condition from the streaming encoding features of the multiple audio frames after convolution processing to obtain the streaming encoding features of the multiple audio frames after frame reduction processing.
[0313] In some possible implementation manners, the frame downsampling module 1603 is specifically configured to:
[0314] Perform convolution processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolution processing;
[0315] Extract the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions from the streaming encoding features of the multiple audio frames after convolution processing;
[0316] Based on the attention mechanism, perform attention calculation on the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions to obtain the streaming encoding features of the multiple audio frames after frame downsampling processing.
[0317] In some possible implementation manners, the frame downsampling module 1603 is specifically configured to:
[0318] Determine the audio frames whose probability value of the preset character is less than the preset threshold according to the character probability distribution of the multiple audio frames, where the preset character is at least one of a blank character or an interval character;
[0319] Extract the streaming encoding features of the audio frames whose probability value of the preset character is less than the preset threshold from the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions.
[0320] In some possible implementation manners, the frame downsampling module 1603 is specifically configured to:
[0321] Perform matrix multiplication processing on the streaming encoding features of the multiple audio frames after convolution processing and the streaming encoding features of the multiple audio frames after frame downsampling processing to obtain a similarity matrix, where the similarity matrix represents the similarity between the streaming encoding features of the multiple audio frames after convolution processing and the streaming encoding features of the multiple audio frames after frame downsampling processing;
[0322] Perform Scale processing on the similarity matrix to obtain the similarity matrix after Scale processing;
[0323] Perform normalization processing on the similarity matrix after Scale processing to obtain the normalized similarity matrix;
[0324] Perform matrix multiplication processing on the streaming encoding features of the multiple audio frames after convolution processing and the normalized similarity matrix to obtain the streaming encoding features of the multiple audio frames after frame downsampling processing.
[0325] In some possible implementation manners, the obtaining module 1605 is configured to:
[0326] Determine the output probability distribution of the audio stream based on the non-streaming coding features of the audio stream and the text prediction results of multiple audio frames in the audio stream. The text prediction results are predicted based on the speech recognition results of the previous audio frame of the audio frame. The output probability distribution represents the probability values of the multiple audio frames on the multiple candidate texts corresponding to the text prediction results respectively;
[0327] Decode the output probability distribution of the audio stream to obtain the speech recognition result of the audio stream.
[0328] Figure 17 It is a framework schematic diagram of a training device for a speech recognition model provided by an embodiment of the present application. See Figure 17 , the training device of the speech recognition model includes a training module 1701 and an acquisition module 1702. Among them,
[0329] The training module 1701 is used to iteratively train the initial model based on audio training data, and the audio training data includes audio samples and the annotated texts of the audio samples;
[0330] Among them, during any iterative training process, input the audio training data into the model obtained after the previous iterative training, perform streaming coding on multiple audio frames in the audio sample to obtain the streaming coding features of the multiple audio frames in the audio sample; perform downsampling processing on the streaming coding features of the multiple audio frames in the audio sample to obtain the streaming coding features of the multiple audio frames after downsampling processing. The number of frames of the multiple audio frames after downsampling processing is less than the number of frames of the multiple audio frames in the audio sample; perform non-streaming coding on the streaming coding features of the multiple audio frames after downsampling processing to obtain the non-streaming coding features of the audio sample; obtain the speech recognition result of the audio sample based on the non-streaming coding features of the audio sample; determine the model loss value based on the speech recognition result and the annotated text; adjust the model parameters based on the model loss value;
[0331] The acquisition module 1702 is used to obtain the trained model as the speech recognition model during the iterative training.
[0332] The technical solution provided by the embodiments of this application obtains a speech recognition model with a relatively low speech recognition delay through iterative training of an initial model. Among them, during any iterative training process, by performing downsampling on the streaming coding features of multiple audio frames in an audio sample, the number of original audio frames in the audio sample is reduced, that is, the number of audio frames to be non-streaming coded is reduced, which can effectively shorten the length of the feature sequence to be non-streaming coded, and further reduce the computational amount of non-streaming coding, effectively reducing the speech recognition delay, and thus effectively improving the speech recognition efficiency. Furthermore, by adjusting the model parameters according to the model loss value, the learning ability of the model can be improved, and thus a speech recognition model with better learning ability can be trained.
[0333] In some possible implementation manners, the device further includes a normalization module, configured to:
[0334] Perform normalization processing on the streaming coding features of the multiple audio frames to obtain the character probability distribution of the multiple audio frames, where the character probability distribution represents the probability values of the multiple audio frames on multiple candidate characters respectively;
[0335] The training module 1701 is specifically configured to:
[0336] According to the character probability distribution of the multiple audio frames, perform downsampling on the streaming coding features of the multiple audio frames in the audio sample to obtain the streaming coding features of the multiple audio frames after downsampling.
[0337] In some possible implementation manners, the training module 1701 is specifically configured to:
[0338] Perform convolution processing on the streaming coding features of the multiple audio frames to obtain the streaming coding features of the multiple audio frames after convolution processing;
[0339] Extract the streaming coding features of the audio frames whose character probability distribution meets the preset conditions from the streaming coding features of the multiple audio frames after convolution processing to obtain the streaming coding features of the multiple audio frames after downsampling.
[0340] In some possible implementation manners, the training module 1701 is specifically configured to:
[0341] Perform convolution processing on the streaming coding features of the multiple audio frames to obtain the streaming coding features of the multiple audio frames after convolution processing;
[0342] Extract the streaming coding features of the audio frames whose character probability distribution meets the preset conditions from the streaming coding features of the multiple audio frames after convolution processing;
[0343] Based on the attention mechanism, perform attention calculation on the streaming encoding features of the audio frames whose character probability distribution satisfies the preset conditions, and obtain the streaming encoding features of the multiple audio frames after the frame reduction process.
[0344] In some possible implementation manners, the training module 1701 is specifically configured to:
[0345] According to the character probability distribution of the multiple audio frames, determine the audio frames whose probability value of the preset character is less than the preset threshold, where the preset character is at least one of a blank character or an interval character;
[0346] Extract the streaming encoding features of the audio frames whose probability value of the preset character is less than the preset threshold from the streaming encoding features of the multiple audio frames, and obtain the streaming encoding features of the audio frames whose character probability distribution satisfies the preset conditions.
[0347] In some possible implementation manners, the training module 1701 is specifically configured to:
[0348] Perform matrix multiplication processing on the streaming encoding features of the multiple audio frames after the convolution process and the streaming encoding features of the multiple audio frames after the frame reduction process to obtain a similarity matrix, where the similarity matrix represents the similarity between the streaming encoding features of the multiple audio frames after the convolution process and the streaming encoding features of the multiple audio frames after the frame reduction process;
[0349] Perform Scale processing on the similarity matrix to obtain the similarity matrix after Scale processing;
[0350] Perform normalization processing on the similarity matrix after Scale processing to obtain the similarity matrix after normalization processing;
[0351] Perform matrix multiplication processing on the streaming encoding features of the multiple audio frames after the convolution process and the similarity matrix after normalization processing to obtain the streaming encoding features of the multiple audio frames after the frame reduction process.
[0352] In some possible implementation manners, the training module 1701 is specifically configured to:
[0353] Based on the non-streaming encoding features of the audio sample and the text prediction results of the multiple audio frames in the audio sample, determine the output probability distribution of the audio sample, where the text prediction results are predicted based on the speech recognition results of the previous audio frame of the audio frame, and the output probability distribution represents the probability values of the multiple audio frames on the multiple candidate texts corresponding to the text prediction results;
[0354] Decode the output probability distribution of the audio sample to obtain the speech recognition result of the audio sample.
[0355] An embodiment of the present application further provides an electronic device, including: a processor and a memory. The processor is connected to the memory, and the memory is used to store program code. The processor executes the program code stored in the memory, thereby implementing the speech recognition method and the training method of the speech recognition model provided by the embodiments of the present application.
[0356] An embodiment of the present application further provides a computer-readable storage medium, on which program code is stored. When the program code runs on an electronic device, the electronic device is caused to execute each function or step executed by the electronic device in the above method embodiment.
[0357] An embodiment of the present application further provides a computer program product, including program code. When the program code runs on an electronic device, the electronic device is caused to execute each function or step executed by the electronic device in the above method embodiment.
[0358] Among them, the electronic device, computer-readable storage medium or computer program product provided by the embodiments of the present application are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0359] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device (such as an electronic device) is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device (such as an electronic device) and unit described above can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated here.
[0360] In several embodiments provided by the present application, it should be understood that the disclosed system, device (such as an electronic device) and method can be implemented in other ways. For example, the device (such as an electronic device) embodiment described above is only illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.
[0361] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0362] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0363] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method in each embodiment of the present application. The aforementioned storage medium includes: flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disk and other various media that can store program codes.
[0364] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A speech recognition method, characterized in that, The method includes: Receiving an audio stream input by a user; Performing streaming encoding on multiple audio frames in the audio stream to obtain streaming encoding features of the multiple audio frames in the audio stream; Performing frame reduction processing on the streaming encoding features of the multiple audio frames in the audio stream to obtain streaming encoding features of the multiple audio frames after frame reduction processing, where the number of frames of the multiple audio frames after frame reduction processing is less than the number of frames of the multiple audio frames in the audio stream; Performing non-streaming encoding on the streaming encoding features of the multiple audio frames after frame reduction processing to obtain non-streaming encoding features of the audio stream; Based on the non-streaming encoding features of the audio stream, obtaining a speech recognition result of the audio stream.
2. The method according to claim 1, wherein After performing streaming encoding on multiple audio frames in the audio stream to obtain streaming encoding features of the multiple audio frames in the audio stream, the method further includes: Performing normalization processing on the streaming encoding features of the multiple audio frames to obtain a character probability distribution of the multiple audio frames, where the character probability distribution represents probability values of the multiple audio frames on multiple candidate characters respectively; The performing frame reduction processing on the streaming encoding features of the multiple audio frames in the audio stream to obtain streaming encoding features of the multiple audio frames after frame reduction processing includes: Performing frame reduction processing on the streaming encoding features of the multiple audio frames in the audio stream according to the character probability distribution of the multiple audio frames to obtain streaming encoding features of the multiple audio frames after frame reduction processing.
3. The method according to claim 2, characterized in that The performing frame reduction processing on the streaming encoding features of the multiple audio frames in the audio stream according to the character probability distribution of the multiple audio frames to obtain streaming encoding features of the multiple audio frames after frame reduction processing includes: Performing convolution processing on the streaming encoding features of the multiple audio frames to obtain streaming encoding features of the multiple audio frames after convolution processing; Extracting streaming encoding features of audio frames whose character probability distribution meets a preset condition from the streaming encoding features of the multiple audio frames after convolution processing to obtain streaming encoding features of the multiple audio frames after frame reduction processing.
4. The method according to claim 2, wherein The performing frame reduction processing on the streaming encoding features of the multiple audio frames in the audio stream according to the character probability distribution of the multiple audio frames to obtain streaming encoding features of the multiple audio frames after frame reduction processing includes: Performing convolution processing on the streaming encoding features of the multiple audio frames to obtain streaming encoding features of the multiple audio frames after convolution processing; Extracting streaming encoding features of audio frames whose character probability distribution meets a preset condition from the streaming encoding features of the multiple audio frames after convolution processing; Based on an attention mechanism, performing attention calculation on the streaming encoding features of the audio frames whose character probability distribution meets a preset condition to obtain streaming encoding features of the multiple audio frames after frame reduction processing.
5. The method according to claim 3 or 4, characterized in that, The extracting streaming encoding features of audio frames whose character probability distribution meets a preset condition from the streaming encoding features of the multiple audio frames after convolution processing includes: Determining, according to the character probability distribution of the multiple audio frames, audio frames whose probability value of a preset character is less than a preset threshold, where the preset character is at least one of a blank character or an interval character; Extract the streaming encoding features of the audio frames whose probability values of the preset characters are less than the preset threshold from the streaming encoding features of the multiple audio frames, so as to obtain the streaming encoding features of the audio frames whose character probability distributions meet the preset conditions.
6. The method according to claim 4, characterized in that, The performing attention calculation on the streaming encoding features of the audio frames whose character probability distributions meet the preset conditions based on the attention mechanism to obtain the streaming encoding features of the multiple audio frames after the frame reduction processing includes: Perform matrix multiplication on the streaming encoding features of the multiple audio frames after the convolution processing and the streaming encoding features of the multiple audio frames after the frame reduction processing to obtain a similarity matrix, where the similarity matrix represents the similarity between the streaming encoding features of the multiple audio frames after the convolution processing and the streaming encoding features of the multiple audio frames after the frame reduction processing; Perform a Scale processing on the similarity matrix to obtain a similarity matrix after the Scale processing; Perform a normalization processing on the similarity matrix after the Scale processing to obtain a similarity matrix after the normalization processing; Perform matrix multiplication on the streaming encoding features of the multiple audio frames after the convolution processing and the similarity matrix after the normalization processing to obtain the streaming encoding features of the multiple audio frames after the frame reduction processing.
7. The method according to claim 1, wherein The obtaining the speech recognition result of the audio stream based on the non-streaming encoding features of the audio stream includes: Based on the non-streaming encoding features of the audio stream and the text prediction results of the multiple audio frames in the audio stream, determine the output probability distribution of the audio stream, where the text prediction results are predicted based on the speech recognition results of the previous audio frames of the audio frames, and the output probability distribution represents the probability values of the multiple audio frames on the multiple candidate texts corresponding to the text prediction results respectively; Decode the output probability distribution of the audio stream to obtain the speech recognition result of the audio stream.
8. A training method for a speech recognition model, characterized in that, The method includes: Iteratively train an initial model based on audio training data, where the audio training data includes audio samples and the labeled texts of the audio samples; Wherein, in the process of any iterative training, input the audio training data into the model obtained after the previous iterative training, perform streaming encoding on the multiple audio frames in the audio sample to obtain the streaming encoding features of the multiple audio frames in the audio sample; perform frame reduction processing on the streaming encoding features of the multiple audio frames in the audio sample to obtain the streaming encoding features of the multiple audio frames after the frame reduction processing, and the number of frames of the multiple audio frames after the frame reduction processing is less than the number of frames of the multiple audio frames in the audio sample; perform non-streaming encoding on the streaming encoding features of the multiple audio frames after the frame reduction processing to obtain the non-streaming encoding features of the audio sample; obtain the speech recognition result of the audio sample based on the non-streaming encoding features of the audio sample; determine the model loss value based on the speech recognition result and the labeled text; adjust the model parameters based on the model loss value; At the end of the iterative training, obtain the trained model as the speech recognition model.
9. The method according to claim 8, wherein After streaming encoding multiple audio frames in the audio sample to obtain the streaming encoding features of the multiple audio frames in the audio sample, the method further includes: Performing normalization processing on the streaming encoding features of the multiple audio frames to obtain the character probability distribution of the multiple audio frames, where the character probability distribution represents the probability values of the multiple audio frames on multiple candidate characters respectively; The performing downsampling processing on the streaming encoding features of multiple audio frames in the audio sample to obtain the streaming encoding features of the multiple audio frames after downsampling processing includes: Performing downsampling processing on the streaming encoding features of multiple audio frames in the audio sample according to the character probability distribution of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after downsampling processing.
10. The method according to claim 9, wherein The performing downsampling processing on the streaming encoding features of multiple audio frames in the audio sample according to the character probability distribution of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after downsampling processing includes: Performing convolution processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolution processing; Extracting the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions from the streaming encoding features of the multiple audio frames after convolution processing to obtain the streaming encoding features of the multiple audio frames after downsampling processing.
11. The method according to claim 9, wherein The performing downsampling processing on the streaming encoding features of multiple audio frames in the audio sample according to the character probability distribution of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after downsampling processing includes: Performing convolution processing on the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the multiple audio frames after convolution processing; Extracting the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions from the streaming encoding features of the multiple audio frames after convolution processing; Performing attention calculation on the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions based on the attention mechanism to obtain the streaming encoding features of the multiple audio frames after downsampling processing.
12. The method according to claim 10 or 11, characterized in that The extracting the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions from the streaming encoding features of the multiple audio frames after convolution processing includes: Determining the audio frames whose probability value of the preset character is less than the preset threshold according to the character probability distribution of the multiple audio frames, where the preset character is at least one of a blank character or an interval character; Extracting the streaming encoding features of the audio frames whose probability value of the preset character is less than the preset threshold from the streaming encoding features of the multiple audio frames to obtain the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions.
13. The method according to claim 11, characterized in that, The performing attention calculation on the streaming encoding features of the audio frames whose character probability distribution meets the preset conditions based on the attention mechanism to obtain the streaming encoding features of the multiple audio frames after downsampling processing includes: Perform matrix multiplication on the streaming encoding features of the multiple audio frames after convolution processing and the streaming encoding features of the multiple audio frames after downsampling processing to obtain a similarity matrix, where the similarity matrix represents the similarity between the streaming encoding features of the multiple audio frames after convolution processing and the streaming encoding features of the multiple audio frames after downsampling processing; Perform a Scale process on the similarity matrix to obtain a similarity matrix after the Scale process; Perform a normalization process on the similarity matrix after the Scale process to obtain a similarity matrix after the normalization process; Perform matrix multiplication on the streaming encoding features of the multiple audio frames after convolution processing and the similarity matrix after the normalization process to obtain the streaming encoding features of the multiple audio frames after downsampling processing.
14. The method according to claim 8, wherein Obtaining the speech recognition result of the audio sample based on the non-streaming encoding features of the audio sample includes: Based on the non-streaming encoding features of the audio sample and the text prediction results of multiple audio frames in the audio sample, determine the output probability distribution of the audio sample, where the text prediction results are predicted based on the speech recognition results of the previous audio frame of the audio frame, and the output probability distribution represents the probability values of the multiple audio frames on the multiple candidate texts corresponding to the text prediction results respectively; Decode the output probability distribution of the audio sample to obtain the speech recognition result of the audio sample.
15. An electronic device, characterized in that, Comprising a memory and a processor; the memory is used for storing program codes; the processor is used for calling the program codes to execute the method according to any one of claims 1-7 or 8-14.
16. A computer-readable storage medium, characterized in that, Comprising program codes, when the program codes run on an electronic device, enabling the electronic device to execute the method according to any one of claims 1-7 or 8-14.
Citation Information
Patent Citations
Multimedia processing method and device
CN104639978A
Speech recognition method and device, electronic equipment and computer readable storage medium
CN113327603A
Multi-mode rejection method and system based on intelligent voice interaction
CN114267347A
Speech recognition method and device, storage medium and equipment
CN114333778A
Voice signal processing method and device, storage medium, electronic equipment and product
CN114550722A