Method, device and equipment for streaming speech recognition and model training

By blocking and coding the voice fragments, combined with local attention calculation within the block, the problem of low streaming speech recognition efficiency of autoregressive decoder is solved, efficient streaming speech recognition is achieved, and the recognition effect is improved.

CN115273830BActive Publication Date: 2025-05-23ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210870146.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-22
Publication Date
2025-05-23
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

The streaming speech recognition model based on autoregressive decoder has low recognition efficiency, resulting in low speech recognition efficiency.

Method used

By chunking the input voice segment, at least one chunk is generated, each chunk is encoded to generate an acoustic representation, and the number of words contained in each chunk and the timestamp and acoustic semantic characteristics of each word are predicted, and local attention calculations within the block are performed to decode text information.

Benefits of technology

Streaming speech recognition based on non-autoregressive decoder is realized. Streaming speech recognition can be achieved by calling the streaming speech recognition model at one time, improving the recognition efficiency, avoiding problems such as repeated word output and noisy word output, and improving the personalized customized information recognition effect in specific scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273830B_ABST
    Figure CN115273830B_ABST
Patent Text Reader

Abstract

The present application provides a method, device and equipment for streaming speech recognition and model training. The method of the present application, through a streaming speech recognition model based on a non-autoregressive decoder, blocks the speech acoustic features of the currently input speech segment, generates at least one block, encodes each block to generate an acoustic representation of each block, and predicts the number of words contained in each block and the timestamp and acoustic semantic features of each word, performs local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, thereby decoding the text information corresponding to each block, and can use the number of words contained in the block and the timestamp and acoustic semantic features of each word to guide the local attention learning within the block, realize streaming speech recognition based on a non-autoregressive decoder, and can realize streaming speech recognition by calling the streaming speech recognition model once, thereby improving the efficiency of streaming speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer technology, and in particular to a method, device and equipment for streaming speech recognition and model training. Background Art

[0002] Speech recognition technology is a technology that allows machines to convert speech signals into corresponding text or commands through the process of recognition and understanding. Among them, end-to-end speech recognition has received widespread attention from academia and industry, and has demonstrated better performance than traditional hybrid modeling solutions in most speech recognition tasks.

[0003] Speech recognition can use either a streaming speech recognition model or a non-streaming speech recognition model. In the process of processing speech streams, streaming speech recognition models support real-time return of recognition results, while non-streaming speech recognition models need to process complete sentences before returning recognition results. Currently, in many fields, such as online voice interaction and online speech recognition services, the efficiency of speech recognition is improved by using end-to-end streaming speech recognition models.

[0004] At present, end-to-end streaming speech recognition usually converts speech features into text based on an autoregressive decoder. It is necessary to recognize unrecognized characters in sequence based on recognized characters. The speech recognition model needs to be called once to recognize each character, which has low computational efficiency and leads to low speech recognition efficiency. Summary of the invention

[0005] The present application provides a method, apparatus and device for streaming speech recognition and model training, which are used to solve the problem of low recognition efficiency of streaming speech recognition models based on autoregressive decoders.

[0006] In a first aspect, the present application provides a method for streaming speech recognition, comprising:

[0007] Acquire the currently input speech segment in real time and extract the speech acoustic features of the speech segment;

[0008] Performing block processing on the speech acoustic features of the speech segment to generate at least one block, and encoding each block to generate an acoustic representation of each block;

[0009] According to the acoustic representation of each block, determine the number of words contained in each block and the timestamp and acoustic semantic features of each word;

[0010] According to the number of words contained in each block and the timestamp and acoustic semantic features of each word, local attention calculation is performed within the block to determine the text information corresponding to each block, and the text information corresponding to each block is spliced ​​to obtain the text information corresponding to the speech segment.

[0011] In a second aspect, the present application provides a method for training a streaming speech recognition model, comprising:

[0012] The speech acoustic features of the sample speech are divided into blocks by a block encoder of the streaming speech recognition model to generate a plurality of blocks, and each block is encoded to generate an acoustic representation of each block;

[0013] Determining the number of words contained in each block and the timestamp and acoustic semantic features of each word according to the acoustic representation of each block by the predictor of the streaming speech recognition model;

[0014] The block attention decoder of the streaming speech recognition model performs a local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, determines the first text information corresponding to each block, integrates the first text information corresponding to each block, obtains the first text corresponding to the sample speech, and uses the first text as the text recognition result;

[0015] The model parameters of the streaming speech recognition model are updated according to the text recognition result and the target text corresponding to the sample speech.

[0016] In a third aspect, the present application provides a streaming speech recognition device, comprising:

[0017] A preprocessing module, used to obtain the currently input speech segment in real time and extract the speech acoustic features of the speech segment;

[0018] A block encoding module, used for performing block processing on the speech acoustic features of the speech segment to generate at least one block, and encoding each block to generate an acoustic representation of each block;

[0019] A prediction module, used for determining the number of words contained in each block and the timestamp and acoustic semantic features of each word according to the acoustic representation of each block;

[0020] The block attention decoding module is used to perform local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, determine the text information corresponding to each block, and splice the text information corresponding to each block to obtain the text information corresponding to the speech segment.

[0021] In a fourth aspect, the present application provides a device for training a streaming speech recognition model, comprising:

[0022] A block encoding module, used for dividing the speech acoustic features of the sample speech into blocks through a block encoder of a streaming speech recognition model to generate a plurality of blocks, and encoding each block to generate an acoustic representation of each block;

[0023] A prediction module, used for determining the number of words contained in each block and the timestamp and acoustic semantic features of each word according to the acoustic representation of each block through the predictor of the streaming speech recognition model;

[0024] A block attention decoding module, used for performing local attention calculation within a block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word through the block attention decoder of the streaming speech recognition model, determining the first text information corresponding to each block, integrating the first text information corresponding to each block, obtaining the first text corresponding to the sample speech, and taking the first text as the text recognition result;

[0025] A parameter updating module is used to update the model parameters of the streaming speech recognition model according to the text recognition result and the target text corresponding to the sample speech.

[0026] In a fifth aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0027] The memory stores computer-executable instructions;

[0028] The processor executes the computer-executable instructions stored in the memory to implement the method described in the first aspect or the second aspect.

[0029] In a sixth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the first aspect or the second aspect.

[0030] In a seventh aspect, the present application provides a computer program product, including a computer program, which implements the method described in the first aspect or the second aspect when executed by a processor.

[0031] The method, device and equipment for streaming speech recognition and model training provided in the present application perform block processing on the speech acoustic features of the currently input speech segment to generate at least one block, encode each block to generate an acoustic representation of each block, and predict the number of words contained in each block and the timestamp and acoustic semantic features of each word, perform local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, thereby decoding the text information corresponding to each block, and can use the number of words contained in the block and the timestamp and acoustic semantic features of each word to guide local attention learning within the block, realize streaming speech recognition based on a non-autoregressive decoder, and realize streaming speech recognition by calling the streaming speech recognition model once, thereby improving the efficiency of streaming speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0033] Figure 1 An example diagram of a system architecture applicable to the method for streaming speech recognition provided in this application;

[0034] Figure 2 A flow chart of a method for streaming speech recognition provided in an exemplary embodiment of the present application;

[0035] Figure 3 A framework diagram of a streaming speech recognition model provided for an example embodiment of the present application;

[0036] Figure 4 A flow chart of a method for streaming speech recognition model training provided in an exemplary embodiment of the present application;

[0037] Figure 5 A system framework diagram for streaming speech recognition model training provided in an example embodiment of the present application;

[0038] Figure 6 A schematic diagram of the structure of a device for streaming speech recognition provided in an exemplary embodiment of the present application;

[0039] Figure 7 A schematic diagram of the structure of a device for training a streaming speech recognition model provided in an exemplary embodiment of the present application;

[0040] Figure 8 A schematic diagram of the structure of a device for streaming speech recognition model training provided in another exemplary embodiment of the present application;

[0041] Fig. 9 A schematic diagram of the structure of an electronic device provided in an exemplary embodiment of the present application.

[0042] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0043] Here, example embodiments are described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following example embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0044] First, the terms involved in this application are explained:

[0045] Automatic Speech Recognition (ASR) is a technology that converts speech into text.

[0046] End-to-End Speech Recognition (ASR) model: refers to a model that directly maps the input acoustic feature sequence into the text information of the word.

[0047] Speech recognition can use either a streaming speech recognition model or a non-streaming speech recognition model. In the process of processing speech streams, the streaming speech recognition model supports real-time return of recognition results, while the non-streaming speech recognition model needs to process the entire sentence before returning the recognition result. In many real-time speech recognition scenarios in actual applications, such as online voice interaction and online speech recognition services, the efficiency of speech recognition can be improved by adopting an end-to-end streaming speech recognition solution.

[0048] At present, end-to-end streaming speech recognition is usually based on an autoregressive decoder (AR Decoder) to convert speech features into text. It is necessary to recognize unrecognized characters in sequence based on recognized characters. The speech recognition model needs to be called once to recognize each character. The computational efficiency is low, and it is easy to have repeated words, noisy words, etc., and the recognition effect of personalized customized information in specific scenarios is poor.

[0049] The present application provides an end-to-end streaming speech recognition model based on a non-autoregressive decoder (NARDecoder), and provides a method for training a streaming speech recognition model. Based on a sample speech of a complete sentence of a historical input, the speech acoustic features of the sample speech and the annotated target text are obtained, where the target text is the text content corresponding to the sample speech. The speech acoustic features of the sample speech are divided into blocks by a block encoder of the streaming speech recognition model to generate a plurality of blocks, and each block is encoded to generate an acoustic representation of each block; and the predictor of the streaming speech recognition model is used to determine the words contained in each block according to the acoustic representation of each block. The number of words contained in each block and the timestamp and acoustic semantic features of each word are determined through the block attention decoder of the streaming speech recognition model, and the first text information corresponding to each block is determined, and the first text information corresponding to each block is integrated to obtain the first text corresponding to the sample speech, and the first text is used as the text recognition result; according to the text recognition result and the target text corresponding to the sample speech, the model parameters of the streaming speech recognition model are updated to realize the training of the streaming speech recognition model and obtain the trained streaming speech recognition model.

[0050] The present application also provides a method for streaming speech recognition. During online semantic recognition, the currently input speech segment is acquired in real time, and the speech acoustic features of the speech segment are extracted; the speech acoustic features of the speech segment are processed into blocks using a trained streaming speech recognition model to generate at least one block, and each block is encoded to generate an acoustic representation of each block; the number of words contained in each block and the timestamp and acoustic semantic features of each word are determined according to the acoustic representation of each block; local attention calculation is performed within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word to determine the text information corresponding to each block, and the text information corresponding to each block is obtained by decoding it once using a non-autoregressive decoder; the text information corresponding to each block is obtained by splicing the text information corresponding to each block, thereby improving the efficiency of streaming speech recognition.

[0051] For example, taking the real-time voice interaction scenario as an example, the streaming voice recognition method provided by the present application can be applied to Figure 1 The system architecture is shown in Figure 1. Figure 1 As shown, the system architecture includes: a terminal and a server.

[0052] The server may be a server that provides speech recognition services, a server of a speech interaction system, etc., and may be a server cluster deployed in the cloud. The server stores an end-to-end streaming speech recognition model based on a non-autoregressive decoder. Through the preset operation logic in the server, the server uses the streaming speech recognition model to perform speech recognition on the real-time input speech fragments, obtains the recognition result, and can feed back the recognition result to the terminal. In addition, the server may store the training data required for training, and implement the model training of the end-to-end streaming speech recognition model based on the non-autoregressive decoder based on the training data to obtain a trained streaming speech recognition model.

[0053] The terminal may specifically be a hardware device with network communication function, computing function and information display function, including but not limited to smart phones, tablet computers, desktop computers, Internet of Things devices, etc.

[0054] Through communication and interaction with the server, when the user uses the terminal to input a voice stream, the server can obtain the currently input voice segment in real time and extract the voice acoustic features of the voice segment; use the trained streaming speech recognition model to divide the voice acoustic features of the voice segment into blocks to generate at least one block, and encode each block to generate an acoustic representation of each block; according to the acoustic representation of each block, determine the number of words contained in each block and the timestamp and acoustic semantic features of each word; perform local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, determine the text information corresponding to each block, and splice the text information corresponding to each block to obtain the text information corresponding to the voice segment, and output the recognition result to the terminal according to the preset rules.

[0055] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0056] Figure 2 A flow chart of a method for streaming speech recognition provided in an exemplary embodiment of the present application. The execution subject of this embodiment may be the server mentioned above, such as Figure 2 As shown, the specific steps of this method are as follows:

[0057] Step S201: Acquire the currently input speech segment in real time and extract the speech acoustic features of the speech segment.

[0058] In this embodiment, according to the preset voice segment size, the real-time input voice stream is collected to obtain the currently input voice segment, and the voice acoustic features of the voice segment are extracted.

[0059] Exemplarily, the speech acoustic features of a speech segment may be Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Cepstral Coefficients (LPCC), short-time average energy, average amplitude change rate, Fbank features, etc., which are not specifically limited here.

[0060] After the speech acoustic features of the speech segment are extracted, the speech acoustic features of the speech segment are input into the trained streaming speech recognition model for speech recognition, and the processing flow of steps S202-S204 is implemented by the streaming speech recognition model to obtain a text recognition result, that is, text information corresponding to the speech segment. The streaming speech recognition model can be trained by the solution provided in the subsequent streaming speech model training method embodiment.

[0061] For example, Figure 3 The framework diagram of the streaming speech recognition model provided by an example embodiment of the present application. Figure 3 As shown, the trained streaming speech recognition model includes: a chunk encoder, a predictor, and a chunk attention decoder. The chunk encoder is used to chunk the speech acoustic features of a speech segment, generate multiple chunks, and encode each chunk to generate an acoustic representation of each chunk. The predictor is used to determine the number of words contained in each chunk, the timestamp of each word, and the acoustic semantic features based on the acoustic representation of each chunk. The chunk attention decoder is used to perform local attention calculation within the chunk based on the number of words contained in each chunk, the timestamp of each word, and the acoustic semantic features, determine the text information corresponding to each chunk, and integrate the text information corresponding to each chunk to obtain the text recognition result corresponding to the sample speech.

[0062] Step S202: block-processing the speech acoustic features of the speech segment to generate at least one block, and encode each block to generate an acoustic representation of each block.

[0063] In this step, the speech acoustic features of the speech segment are input into a block encoder, and the block encoder divides the speech acoustic features of the sample speech into blocks to generate multiple blocks, and encodes each block to generate an acoustic representation of each block. The acoustic representation of each block is input into a predictor.

[0064] Exemplarily, the block encoder can be a multi-layer neural network, and the block encoder can adopt any of the following neural networks: Deep-Feedforward Sequential Memory Networks (DFSMN), Convolutional Neural Network (CNN), Long Short-Term Memory Network (LSTM), Bi-directional Long Short-Term Memory Network (BLSTM), Transformer.

[0065] Specifically, according to a preset block size, the speech acoustic features of the speech segment are processed in blocks, and the speech acoustic features of the speech segment are divided into at least one block.

[0066] The preset block size can be expressed as a delay. The smaller the preset block size, the smaller the delay of the streaming speech recognition, but the lower the recognition accuracy. The larger the preset block size, the higher the accuracy of the streaming speech recognition, but the longer the delay of the streaming speech recognition. The preset block size can be set to 3 frames, 5 frames, 10 frames, 15 frames, etc., and can be set according to the needs of the actual application scenario when training the streaming speech recognition model, and is not specifically limited here.

[0067] The preset voice segment size can be set according to the preset block size and in combination with the specific application scenario. For example, the voice segment size can be equal to the preset block size. In this way, the voice segment collected in real time is a block, and streaming voice recognition can be performed in real time, thereby improving the efficiency and real-time performance of voice recognition.

[0068] In addition, the preset voice segment size may also be larger than the preset block size, which is not specifically limited here. When the preset voice segment is larger than the preset block size, the voice acoustic features of the voice segment are divided into blocks with repeated information, that is, each block contains historical and future information.

[0069] Furthermore, after the speech acoustic features of the speech segment are processed into blocks to generate at least one block, each block is dynamically encoded to convert the speech acoustic features of each block into a new discriminative high-level representation to obtain the acoustic representation of each block. The acoustic representation of each block is also called block memory.

[0070] Step S203: Determine the number of characters contained in each block and the timestamp and acoustic semantic features of each character according to the acoustic representation of each block.

[0071] After obtaining the acoustic representation of each block, the acoustic representation of each block is input into the predictor, and the predictor predicts the number of words contained in each block and the timestamp and acoustic semantic features of each word based on the acoustic representation of each block.

[0072] A character can be a character in Chinese, or a token or sub-token in the word segmentation result in English.

[0073] Exemplarily, the predictor may be a 2-layer neural network, and the neural network may be any of the following: deep neural networks (DNN), CNN, or LSTM.

[0074] Step S204: perform local attention calculation within the block according to the number of characters contained in each block and the timestamp and acoustic semantic features of each character, determine the text information corresponding to each block, and splice the text information corresponding to each block to obtain the text information corresponding to the voice segment.

[0075] In this step, the number of characters contained in each block, the timestamp of each character, and the acoustic semantic features are input into the block attention decoder, and the block attention decoder performs local attention calculation within the block according to the number of characters contained in each block, the timestamp of each character, and the acoustic semantic features, and determines the text information corresponding to each block; then, the text information corresponding to each block is spliced ​​to obtain the text information corresponding to the speech segment.

[0076] Exemplarily, the block attention decoder is a multi-layer neural network, and any of the following neural networks can be used: DFSMN, CNN, BLSTM, Transformer.

[0077] In this embodiment, the speech acoustic features of the currently input speech segment are processed into blocks to generate at least one block, each block is encoded to generate an acoustic representation of each block, and the number of words contained in each block and the timestamp and acoustic semantic features of each word are predicted. According to the number of words contained in each block and the timestamp and acoustic semantic features of each word, the local attention calculation within the block is performed, so as to decode the text information corresponding to each block, and the number of words contained in the block and the timestamp and acoustic semantic features of each word can be used to guide the local attention learning within the block, and realize streaming speech recognition based on non-autoregressive decoder. Streaming speech recognition can be realized by calling the streaming speech recognition model once, thereby improving the efficiency of streaming speech recognition. In addition, it can also avoid repeated words, noise words, etc., and the recognition effect of personalized customized information in specific scenarios is better.

[0078] In an optional embodiment, the predictor is used to determine the number of words contained in the block and the frame boundary of each word. Further, the timestamp of each word is calculated based on the frame boundary of each word, and the frame vector corresponding to each word is extracted from the acoustic representation of each block, and the acoustic semantic features of each word are determined based on the frame vector corresponding to each word.

[0079] Exemplarily, after obtaining the acoustic representation of each block, the acoustic representation of each block is input into the predictor, and the predictor can determine the prediction sequence corresponding to the acoustic representation of each block. The prediction sequence is used to indicate whether each of the multiple frames included in the acoustic features of each block corresponds to a word, so that the number of words included in each block can be determined according to the number of frames corresponding to the word in the prediction sequence, and the frame boundary of each word can be determined.

[0080] The prediction network used by the predictor to predict the number of words contained in the block and the frame boundary of each word can be trained in the following way: obtain a number of block samples, mark the number of words contained in each block sample and the frame boundary of each word, and train the prediction network through this marked information. The trained prediction network has the function of predicting the number of words contained in the block and the frame boundary of each word.

[0081] Optionally, according to the acoustic feature frames corresponding to each word contained in the block, a Continuous Integrate-and-Fire (CIF) model or a Connectionist Temporal Classification (CTC) model based on an attention mechanism is used to average multiple acoustic feature frames corresponding to each word to obtain the acoustic semantic features of each word.

[0082] In an optional embodiment, the acoustic representation of each block determined by the block encoder may not be input into the block attention decoder. In the above step S204, the block attention decoder performs a local self-attention calculation on the acoustic semantic features of the words contained in each block according to the number of words contained in each block and the timestamp of each word, and determines the text information corresponding to each block. The number of words contained in each block and the timestamp of each word are used to guide the local self-attention calculation of each block, so as to realize the function of the non-autoregressive decoder. The non-autoregressive decoder can decode multiple words contained in the block at the same time, so as to obtain the recognition result of the streaming speech recognition. The complete text information of the block can be recognized by calling the block attention decoder once, which can improve the efficiency of the streaming speech recognition. In addition, it can also avoid repeated words, noise words, etc., and has a better recognition effect for personalized customized information in specific scenarios.

[0083] In an optional embodiment, the acoustic representation of each block determined by the block encoder can be input into the block attention decoder. In the above step S204, the block attention decoder performs a local attention calculation on the acoustic representation of each block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, and determines the text information corresponding to each block.

[0084] Specifically, assuming that a plurality of characters are contained in a block m, the number of characters contained in each block and the timestamp and acoustic semantic features of each character are determined by a predictor and input into a block attention decoder, and the acoustic representation of each block determined by a block encoder is also input into a block attention decoder. The block attention decoder knows the acoustic semantic features of each character in the block. By calculating the correlation between the acoustic semantic features of the previous character Yt-1 and the acoustic representation of the block (represented by Cm), the weight corresponding to each element in the acoustic representation of the block can be obtained. Among them, each element in the acoustic representation Cm of the block corresponds to the encoding result corresponding to each frame contained in the block m. Based on the calculated weight, each element in the acoustic representation Cm of the block is weighted and summed, and the weighted sum result is decoded to obtain the next character Yt.

[0085] Since the block attention decoder knows the acoustic and semantic features of each word, it can decode each block at the same time and can decode each word contained in each block at the same time. Compared with the autoregressive decoder that iteratively decodes each word in sequence, it can improve the efficiency of streaming speech recognition. In addition, it can also avoid repeated words, noisy words, etc., and has better recognition effect on personalized customized information in specific scenarios.

[0086] Figure 4 A flow chart of a method for training a streaming speech recognition model provided in an exemplary embodiment of the present application. The execution subject of this embodiment may be the server mentioned above, such as Figure 4 As shown, the specific steps of this method are as follows:

[0087] Step S401: Obtain sample speech, speech acoustic features of the sample speech, and target text.

[0088] Before training the streaming speech recognition model, first obtain training data based on historical speech data. The training data includes sample speech, speech acoustic features of the sample speech, and target text.

[0089] The sample speech is usually a complete sentence of a speech input, and the speech acoustic features of the sample speech are the extracted acoustic features of the sample speech, and the specific implementation method is consistent with the implementation method adopted for extracting the speech acoustic features of the speech segment in the above step S201. The target text of the sample speech refers to the accurate recognition result of the annotated sample speech.

[0090] Exemplarily, the input duration of the sample speech is approximately between 10 seconds and 2 minutes. In some cases, there may be sample speech with an input duration of less than 10 seconds or an input duration of more than 2 minutes.

[0091] Since the size of the speech segments collected during streaming speech recognition is usually much smaller than the length of the sample speech input in one complete time, in this embodiment, in order to train a streaming speech recognition model that can accurately perform speech recognition based on speech segments with shorter input duration, when training the model, a preset block size can be set according to the interval duration of the collected speech segments. The streaming speech recognition model can block the sample speech based on the preset block size, and decode each block separately through local attention calculation within the block to obtain the text information corresponding to each block, so that the trained streaming speech recognition model can accurately recognize the text information corresponding to the speech segment for the speech segment with shorter input duration.

[0092] The preset block size is usually within the range of [300 milliseconds, 1 minute], and the preset block size can be set according to the needs of the actual application scenario.

[0093] After the training data is acquired, the streaming speech recognition model is trained based on the training data through the following steps S402-S405 to obtain a trained streaming speech recognition model.

[0094] Step S402: Block the speech acoustic features of the sample speech through the block encoder of the streaming speech recognition model to generate multiple blocks, and encode each block to generate an acoustic representation of each block.

[0095] When training the streaming speech recognition model, the speech acoustic features of the sample speech are input into the block encoder of the streaming speech recognition model, the speech acoustic features of the sample speech are divided into blocks by the block encoder to generate multiple blocks, and each block is encoded to generate an acoustic representation of each block. The acoustic representation of each block is input into the predictor.

[0096] Exemplarily, the block encoder may be a multi-layer neural network, and the block encoder may adopt any of the following neural networks: DFSMN, CNN, LSTM, BLSTM, Transformer.

[0097] Specifically, the block encoder performs block processing on the speech acoustic features of the speech segment according to a preset block size, and divides the speech acoustic features of the speech segment into at least one block.

[0098] Exemplarily, the preset block size can be expressed as a delay. The smaller the preset block size, the smaller the delay of the streaming speech recognition, but the lower the recognition accuracy. The larger the preset block size, the higher the accuracy of the streaming speech recognition, but the longer the delay of the streaming speech recognition. The preset block size can be set to 3 frames, 5 frames, 10 frames, 15 frames, etc., and can be set according to the needs of the actual application scenario when training the streaming speech recognition model, and is not specifically limited here.

[0099] Furthermore, after the speech acoustic features of the sample speech are processed into blocks to generate at least one block, each block is dynamically encoded to convert the speech acoustic features of each block into a new discriminative high-level representation to obtain the acoustic representation of each block. The acoustic representation of each block is also called block memory.

[0100] Step S403: Determine the number of characters contained in each block and the timestamp and acoustic semantic features of each character according to the acoustic representation of each block through the predictor of the streaming speech recognition model.

[0101] A character can be a character in Chinese, or a token or sub-token in the word segmentation result in English.

[0102] Exemplarily, the predictor may be a 2-layer neural network, and the neural network may be any of the following: deep neural networks (DNN), CNN, or LSTM.

[0103] Step S404: The block attention decoder of the streaming speech recognition model performs local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, determines the first text information corresponding to each block, and integrates the first text information corresponding to each block to obtain the first text corresponding to the sample speech, and uses the first text as the text recognition result.

[0104] In this embodiment, the block attention decoder is a non-autoregressive decoder, which uses the number of words contained in each block and the timestamp of each word to guide the local attention learning within the block, performs the local attention calculation within the block on the acoustic semantic features, and obtains the first text information corresponding to each block. Further, the first text information corresponding to each block is integrated to obtain the first text corresponding to the sample speech, and the first text is used as the text recognition result.

[0105] Exemplarily, the block attention decoder is a multi-layer neural network, and any of the following neural networks can be used: DFSMN, CNN, BLSTM, Transformer.

[0106] Step S405: Update the model parameters of the streaming speech recognition model according to the text recognition result and the target text corresponding to the sample speech.

[0107] After the text recognition result corresponding to the sample speech is recognized by the streaming speech recognition model, a loss value is calculated based on the difference between the text recognition result corresponding to the sample speech and the target text corresponding to the sample speech, and the model parameters of the streaming speech recognition model are updated based on the loss value.

[0108] Exemplarily, the cross entropy (CE) loss can be calculated based on the text recognition result output by the block attention decoder and the target text corresponding to the sample speech, and the model parameters of the streaming speech recognition model can be updated based on the cross entropy loss.

[0109] Exemplarily, the cross entropy loss and the minimum word error rate (MWER) loss can be calculated based on the text recognition results output by the block attention decoder and the target text corresponding to the sample speech, and the model parameters of the streaming speech recognition model can be updated based on the cross entropy loss and the minimum word error rate loss.

[0110] When the convergence condition is met, a better set of model parameters is obtained as the model parameters of the streaming speech recognition model, and a trained streaming speech recognition model is obtained.

[0111] In the present embodiment, during the training of the streaming speech recognition model, the speech acoustic features of the sample speech are divided into blocks by a block encoder to generate multiple blocks, and each block is encoded to generate an acoustic representation of each block; the number of words contained in each block, the timestamp and acoustic semantic features of each word are determined by a predictor according to the acoustic representation of each block; the first text information corresponding to each block is determined by a block attention decoder according to the number of words contained in each block, the timestamp and acoustic semantic features of each word, and the first text information corresponding to each block is integrated to obtain a first text corresponding to the sample speech, and the first text is used as a text recognition result; the model parameters of the streaming speech recognition model are updated according to the text recognition result and the target text corresponding to the sample speech, and a non-autoregressive decoder is used in the trained streaming speech recognition model, which can use the number of words contained in the block, the timestamp and acoustic semantic features of each word to guide the local attention learning in the block, and can decode multiple words contained in the block at the same time. The complete text information contained in the block can be recognized by calling the block attention decoder once, thereby improving the efficiency of streaming speech recognition.

[0112] Figure 5A system framework diagram for streaming speech recognition model training provided in an exemplary embodiment of the present application. In an optional embodiment, the streaming speech recognition model includes a block encoder, a predictor, and a block attention decoder, such as Figure 5 As shown, during the training process of the streaming speech recognition model, a sampler can be added. The sampler is a parameter-free computing module for sampling at least one word's text representation from the text representation of the target text according to the edit distance between the first text output by the block attention decoder and the target text, and using the sampling result to replace the acoustic semantic features of at least one word to obtain updated acoustic semantic features, the updated acoustic semantic features containing correct context information, and inputting the updated acoustic semantic features into the block attention decoder.

[0113] In this embodiment, after step S404, the block attention decoder inputs the first text obtained by the first decoding into the sampler. The sampler samples the text representation of at least one word from the text representation of the target text according to the edit distance between the first text output by the block attention decoder and the target text, and uses the sampling result to replace the acoustic semantic features of at least one word to obtain an updated acoustic semantic feature, the updated acoustic semantic feature contains correct context information to enhance the context information of the acoustic semantic feature output by the predictor, and the updated acoustic semantic feature is input into the block attention decoder, and the text recognition result obtained by decoding based on the acoustic semantic feature containing the correct context information is more accurate.

[0114] Furthermore, the block attention decoder performs local attention calculation within the block according to the number of words contained in each block, the timestamp of each word and the updated acoustic semantic features, determines the second text information corresponding to each block, and integrates the second text information corresponding to each block to obtain the second text corresponding to the sample speech, and uses the second text as the text recognition result.

[0115] By performing a second decoding according to the updated acoustic semantic features containing correct context information, a new text recognition result (ie, a second text corresponding to the sample speech) is obtained, which can improve the accuracy of the text recognition result.

[0116] Furthermore, in the scheme of this embodiment, in step S405, the first loss is determined based on the text recognition result (second text) output by the second pass of the block attention decoder and the target text corresponding to the sample speech, and the model parameters of the streaming speech recognition model are updated based on the first loss, which can improve the accuracy of speech recognition of the trained streaming speech recognition model.

[0117] Exemplarily, according to the text recognition result output by the second pass of the block attention decoder and the target text corresponding to the sample speech, the cross entropy (CE) can be calculated as the first loss.

[0118] Exemplarily, the cross entropy and the minimum word error rate can also be calculated based on the text recognition result output by the second pass of the block attention decoder and the target text corresponding to the sample speech, and the first loss can be determined based on the calculated cross entropy and the minimum word error rate.

[0119] Optionally, in the above step S405, the first loss can be determined based on the text recognition result (second text) output by the second pass of the block attention decoder and the target text corresponding to the sample speech, and the second loss can be determined based on the sum of the number of words contained in each block and the total number of words in the target text; based on the first loss and the second loss, the model parameters of the streaming speech recognition model are updated, and by increasing the calculation of the second loss to update the model parameters, the accuracy of the predictor of the trained streaming speech recognition model in predicting the number of words contained in the block and the timestamp of each word can be improved.

[0120] Exemplarily, according to the text recognition result output by the second pass of the block attention decoder and the target text corresponding to the sample speech, the cross entropy (CE) can be calculated as the first loss.

[0121] Exemplarily, the cross entropy and the minimum word error rate can also be calculated based on the text recognition result output by the second pass of the block attention decoder and the target text corresponding to the sample speech, and the first loss can be determined based on the calculated cross entropy and the minimum word error rate.

[0122] Exemplarily, according to the sum of the number of words contained in each block and the total number of words in the target text, the mean absolute error (MAE) can be calculated as the second loss.

[0123] Optionally, a weighted sum of the first loss and the second loss may be performed to determine a comprehensive loss, and the model parameters may be updated according to the comprehensive loss.

[0124] In this embodiment, the streaming speech recognition model adopts a non-autoregressive decoder structure. The decoder iterates twice during training. However, after the model is trained, only a single decoding is required in actual application decoding, so real-time streaming speech recognition can be achieved.

[0125] In an optional embodiment, the acoustic representation of each block determined by the block encoder may not be input into the block attention decoder. When the block attention decoder of the streaming speech recognition model performs the local attention calculation within the block according to the number of words contained in each block, the timestamp of each word, and the acoustic semantic features to determine the text information corresponding to each block, the block attention decoder of the streaming speech recognition model can perform the local self-attention calculation within the block for the acoustic semantic features of the words contained in each block according to the number of words contained in each block, the timestamp of each word, and the acoustic semantic features to determine the text information corresponding to each block, and the function of the non-autoregressive decoder is realized by using the number of words contained in each block and the timestamp of each word to guide the local self-attention calculation within the block for each block, so as to obtain the recognition result of the streaming speech recognition, and the complete text information of the block can be recognized by calling the block attention decoder once, which can improve the efficiency of the streaming speech recognition; in addition, it can also avoid the situation of repeated words, noise words, etc., and the recognition effect of personalized customized information in specific scenarios is better.

[0126] Exemplarily, when the block attention decoder performs the first decoding, the block attention decoder performs local self-attention calculation on the acoustic semantic features of the characters contained in each block according to the number of characters contained in each block and the timestamp and acoustic semantic features of each character, and determines the text information corresponding to each block.

[0127] Exemplarily, when the block attention decoder performs the second decoding, the block attention decoder performs local self-attention calculation within the block on the updated acoustic semantic features of the words contained in each block according to the number of words contained in each block, the timestamp of each word, and the updated acoustic semantic features, to determine the text information corresponding to each block.

[0128] In an alternative embodiment, the acoustic representation of each block determined by the block encoder can be input into the block attention decoder (e.g. Figure 5 In the example shown in FIG. 1 , the block attention decoder of the streaming speech recognition model performs a local attention calculation within the block according to the number of words contained in each block, the timestamp of each word, and the acoustic semantic features to determine the text information corresponding to each block. The block attention decoder of the streaming speech recognition model performs a local attention calculation within the block on the acoustic representation of each block according to the number of words contained in each block, the timestamp of each word, and the acoustic semantic features to determine the text information corresponding to each block.

[0129] Specifically, assuming that a plurality of characters are contained in a block m, the number of characters contained in each block and the timestamp and acoustic semantic features of each character are determined by a predictor and input into a block attention decoder, and the acoustic representation of each block determined by a block encoder is also input into a block attention decoder. The block attention decoder knows the acoustic semantic features of each character in the block. By calculating the correlation between the acoustic semantic features of the previous character Yt-1 and the acoustic representation of the block (represented by Cm), the weight corresponding to each element in the acoustic representation of the block can be obtained. Among them, each element in the acoustic representation Cm of the block corresponds to the encoding result corresponding to each frame contained in the block m. Based on the calculated weight, each element in the acoustic representation Cm of the block is weighted and summed, and the weighted sum result is decoded to obtain the next character Yt.

[0130] Since the block attention decoder knows the acoustic and semantic features of each word, it can decode each block at the same time and can decode each word contained in each block at the same time. Compared with the autoregressive decoder that iteratively decodes each word in sequence, it can improve the efficiency of streaming speech recognition. In addition, it can also avoid repeated words, noisy words, etc., and has better recognition effect on personalized customized information in specific scenarios.

[0131] Exemplarily, when the block attention decoder performs the first decoding, the block attention decoder performs intra-block local attention calculation on the acoustic representation of each block based on the number of words contained in each block and the timestamp and acoustic semantic features of each word, and determines the first text information corresponding to each block.

[0132] Exemplarily, when the block attention decoder performs the second decoding, the block attention decoder performs intra-block local attention calculation on the acoustic representation of each block according to the number of words contained in each block, the timestamp of each word, and the updated acoustic semantic features, to determine the second text information corresponding to each block.

[0133] Figure 6 This is a schematic diagram of the structure of a streaming speech recognition device provided by an exemplary embodiment of the present application. The device provided by this embodiment is applied to the above-mentioned server, such as Figure 6 As shown, the streaming speech recognition device 60 includes: a preprocessing module 61, a block encoding module 62, a prediction module 63 and a block attention decoding module 64.

[0134] Specifically, the preprocessing module 61 is used to obtain the currently input speech segment in real time and extract the speech acoustic features of the speech segment.

[0135] The block encoding module 62 is used to perform block processing on the speech acoustic features of the speech segment to generate at least one block, and encode each block to generate an acoustic representation of each block.

[0136] The prediction module 63 is used to determine the number of characters contained in each block and the timestamp and acoustic semantic features of each character according to the acoustic representation of each block.

[0137] The block attention decoding module 64 is used to perform local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, determine the text information corresponding to each block, and splice the text information corresponding to each block to obtain the text information corresponding to the speech segment.

[0138] The device provided in this embodiment can be specifically used to perform the above Figure 2 The solutions, specific functions and technical effects that can be achieved by the corresponding method embodiments are not described in detail here.

[0139] In an optional embodiment, when performing the local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word to determine the text information corresponding to each block, the prediction module 63 is specifically used to:

[0140] According to the number of characters contained in each block and the timestamp and acoustic semantic features of each character, local attention calculation is performed on the acoustic representation of each block to determine the text information corresponding to each block; or, according to the number of characters contained in each block and the timestamp of each character, local self-attention calculation is performed on the acoustic semantic features of the characters contained in each block to determine the text information corresponding to each block.

[0141] The device provided in this embodiment can be specifically used to execute the solution provided by any of the above-mentioned streaming speech recognition method embodiments, and the specific functions and technical effects that can be achieved are not described in detail here.

[0142] Figure 7 This is a schematic diagram of the structure of a device for training a streaming speech recognition model provided in an exemplary embodiment of the present application. The device provided in this embodiment is applied to the server mentioned above, such as Figure 7 As shown, the device 70 for streaming speech recognition model training includes: a block encoding module 71, a prediction module 72, a block attention decoding module 73 and a parameter updating module 74.

[0143] Specifically, the block encoding module 71 is used to block the speech acoustic features of the sample speech through the block encoder of the streaming speech recognition model to generate multiple blocks, and encode each block to generate an acoustic representation of each block.

[0144] The prediction module 72 is used to determine the number of words contained in each block and the timestamp and acoustic semantic features of each word according to the acoustic representation of each block through the predictor of the streaming speech recognition model.

[0145] The block attention decoding module 73 is used to perform local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word through the block attention decoder of the streaming speech recognition model, determine the first text information corresponding to each block, and integrate the first text information corresponding to each block to obtain the first text corresponding to the sample speech, and use the first text as the text recognition result.

[0146] The parameter updating module 74 is used to update the model parameters of the streaming speech recognition model according to the text recognition result and the target text corresponding to the sample speech.

[0147] The device provided in this embodiment can be specifically used to perform the above Figure 4 The solutions, specific functions and technical effects that can be achieved by the corresponding method embodiments are not described in detail here.

[0148] In an optional embodiment, Figure 8 As shown, the device 70 for streaming speech recognition model training also includes: a sampling module 75.

[0149] The sampling module 75 is used to: sample the text representation of at least one word from the text representation of the target text according to the edit distance between the first text and the target text, and use the sampling result to replace the acoustic semantic features of the at least one word to obtain updated acoustic semantic features.

[0150] The block attention decoding module 73 is also used to: perform local attention calculation within the block according to the number of words contained in each block, the timestamp of each word and the updated acoustic semantic features through the block attention decoder, determine the second text information corresponding to each block, and integrate the second text information corresponding to each block to obtain the second text corresponding to the sample speech, and use the second text as the text recognition result.

[0151] In an optional embodiment, when the block attention decoder of the streaming speech recognition model performs the local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word to determine the text information corresponding to each block, the block attention decoding module 73 is further used to:

[0152] The block attention decoder of the streaming speech recognition model performs intra-block local attention calculation on the acoustic representation of each block according to the number of words contained in each block, the timestamp of each word, and the acoustic semantic features, and determines the text information corresponding to each block; or, the block attention decoder of the streaming speech recognition model performs intra-block local self-attention calculation on the acoustic semantic features of the words contained in each block according to the number of words contained in each block, the timestamp of each word, and the acoustic semantic features, and determines the text information corresponding to each block.

[0153] In an optional embodiment, when updating the model parameters of the streaming speech recognition model according to the text recognition result and the target text corresponding to the sample speech, the parameter updating module 74 is further used to:

[0154] A first loss is determined according to the text recognition result and the target text corresponding to the sample speech, and a second loss is determined according to the sum of the number of words contained in each block and the total number of words in the target text; based on the first loss and the second loss, the model parameters of the streaming speech recognition model are updated.

[0155] In an optional embodiment, when updating the model parameters of the streaming speech recognition model according to the text recognition result and the target text corresponding to the sample speech, the parameter updating module 74 is further used to:

[0156] A first loss is determined according to the text recognition result and the target text corresponding to the sample speech, and a model parameter of the streaming speech recognition model is updated according to the first loss.

[0157] The device provided in this embodiment can be specifically used to execute the solution provided by any of the above-mentioned method embodiments for streaming speech recognition model training. The specific functions and technical effects that can be achieved will not be repeated here.

[0158] Fig. 9 This is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. Fig. 9 As shown, the electronic device 90 includes: a processor 901, and a memory 902 communicatively connected to the processor 901, and the memory 902 stores computer-executable instructions.

[0159] Among them, the processor executes the computer execution instructions stored in the memory to implement the solution provided by any of the above method embodiments, and the specific functions and technical effects that can be achieved are not repeated here.

[0160] An embodiment of the present application also provides a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed by a processor, they are used to implement the solution provided by any of the above method embodiments. The specific functions and technical effects that can be achieved are not repeated here.

[0161] An embodiment of the present application also provides a computer program product, which includes: a computer program, which is stored in a readable storage medium. At least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the solution provided by any of the above method embodiments. The specific functions and technical effects that can be achieved are not repeated here.

[0162] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel, and are only used to distinguish between different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit "first" and "second" to different types. The meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.

[0163] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0164] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for streaming speech recognition, It is characterized in that include: Acquire the currently input speech segment in real time and extract the speech acoustic features of the speech segment; Performing block processing on the speech acoustic features of the speech segment to generate at least one block, and encoding each block to generate an acoustic representation of each block; According to the acoustic representation of each block, determine the number of words contained in each block and the timestamp and acoustic semantic features of each word; According to the number of words contained in each block and the timestamp and acoustic semantic features of each word, local attention calculation is performed within the block to determine the text information corresponding to each block, and the text information corresponding to each block is spliced ​​to obtain the text information corresponding to the speech segment.

2. The method according to claim 1, It is characterized in that The method of performing local attention calculation within a block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word to determine the text information corresponding to each block includes: According to the number of words contained in each block and the timestamp and acoustic semantic features of each word, the acoustic representation of each block is subjected to local attention calculation within the block to determine the text information corresponding to each block; or, According to the number of characters contained in each block and the timestamp of each character, the acoustic semantic features of the characters contained in each block are calculated by local self-attention within the block to determine the text information corresponding to each block.

3. A method for training a streaming speech recognition model, It is characterized in that include: The speech acoustic features of the sample speech are divided into blocks by a block encoder of the streaming speech recognition model to generate a plurality of blocks, and each block is encoded to generate an acoustic representation of each block; Determining the number of words contained in each block and the timestamp and acoustic semantic features of each word according to the acoustic representation of each block by the predictor of the streaming speech recognition model; The block attention decoder of the streaming speech recognition model performs a local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, determines the first text information corresponding to each block, integrates the first text information corresponding to each block, obtains the first text corresponding to the sample speech, and uses the first text as the text recognition result; The model parameters of the streaming speech recognition model are updated according to the text recognition result and the target text corresponding to the sample speech.

4. The method according to claim 3, It is characterized in that Before updating the model parameters of the streaming speech recognition model according to the text recognition result and the target text, the method further includes: According to the edit distance between the first text and the target text, sampling the text representation of at least one word from the text representation of the target text, and using the sampling result to replace the acoustic semantic feature of the at least one word to obtain an updated acoustic semantic feature; The block attention decoder performs local attention calculation within the block according to the number of words contained in each block, the timestamp of each word and the updated acoustic semantic features, determines the second text information corresponding to each block, and integrates the second text information corresponding to each block to obtain the second text corresponding to the sample speech, and uses the second text as the text recognition result.

5. The method according to claim 3 or 4, It is characterized in that The block attention decoder of the streaming speech recognition model performs local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word to determine the text information corresponding to each block, including: The block attention decoder of the streaming speech recognition model performs a block local attention calculation on the acoustic representation of each block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, so as to determine the text information corresponding to each block; or, The block attention decoder of the streaming speech recognition model performs local self-attention calculation on the acoustic semantic features of the words contained in each block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, so as to determine the text information corresponding to each block.

6. The method according to claim 3 or 4, It is characterized in that The updating of the model parameters of the streaming speech recognition model according to the text recognition result and the target text corresponding to the sample speech includes: Determine a first loss according to the text recognition result and the target text corresponding to the sample speech, and determine a second loss according to the sum of the number of words contained in each block and the total number of words in the target text; According to the first loss and the second loss, the model parameters of the streaming speech recognition model are updated.

7. The method according to claim 3 or 4, It is characterized in that The updating of the model parameters of the streaming speech recognition model according to the text recognition result and the target text corresponding to the sample speech includes: A first loss is determined according to the text recognition result and the target text corresponding to the sample speech, and a model parameter of the streaming speech recognition model is updated according to the first loss.

8. A device for streaming speech recognition, It is characterized in that include: A preprocessing module, used to obtain the currently input speech segment in real time and extract the speech acoustic features of the speech segment; A block encoding module, used for performing block processing on the speech acoustic features of the speech segment to generate at least one block, and encoding each block to generate an acoustic representation of each block; A prediction module, used for determining the number of words contained in each block and the timestamp and acoustic semantic features of each word according to the acoustic representation of each block; The block attention decoding module is used to perform local attention calculation within the block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word, determine the text information corresponding to each block, and splice the text information corresponding to each block to obtain the text information corresponding to the speech segment.

9. A device for training a streaming speech recognition model, It is characterized in that include: A block encoding module, used for dividing the speech acoustic features of the sample speech into blocks through a block encoder of a streaming speech recognition model to generate a plurality of blocks, and encoding each block to generate an acoustic representation of each block; A prediction module, used for determining the number of words contained in each block and the timestamp and acoustic semantic features of each word according to the acoustic representation of each block through the predictor of the streaming speech recognition model; A block attention decoding module, used for performing local attention calculation within a block according to the number of words contained in each block and the timestamp and acoustic semantic features of each word through the block attention decoder of the streaming speech recognition model, determining the first text information corresponding to each block, integrating the first text information corresponding to each block, obtaining the first text corresponding to the sample speech, and taking the first text as the text recognition result; A parameter updating module is used to update the model parameters of the streaming speech recognition model according to the text recognition result and the target text corresponding to the sample speech.

10. An electronic device, It is characterized in that include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

11. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.

12. A computer program product, It is characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 7 when being executed by a processor.

Citation Information

Patent Citations

  • Speech recognition method and device, electronic equipment and computer readable storage medium

    CN113327603A

  • Streaming speech recognition system and method based on non-autoregression model

    CN114203170A