A speech recognition method, device, equipment and readable storage medium

By deploying the attention encoding layer in the speech feature encoder and utilizing the relationship between speech blocks, the problem of not being able to effectively utilize acoustic context information in the prior art is solved, and the accuracy of streaming speech recognition is improved.

CN116312480BActive Publication Date: 2025-05-23ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310126931.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2025-05-23
Estimated Expiration
2043-02-10

AI Technical Summary

Technical Problem

Existing end-to-end streaming speech recognition schemes fail to effectively utilize information from the acoustic context, resulting in low accuracy of the identified text.

Method used

The speech feature encoding subnet containing the attention coding layer is adopted. The speech characteristics of the speech block are determined by combining the feature extraction subnet and the feature encoding subnet, and the context information is used to improve the accuracy of text prediction.

Benefits of technology

Effectively utilize acoustic context information to improve the accuracy of text prediction of streaming speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312480B_ABST
    Figure CN116312480B_ABST
Patent Text Reader

Abstract

This specification discloses a speech recognition method, apparatus, device and readable storage medium. In response to a streaming speech recognition request, the continuously received audio data to be recognized is divided into speech blocks to be recognized according to a preset duration, and each speech block to be recognized is sequentially input into a pre-trained speech recognition model, and a first speech feature is obtained through a feature extraction subnet. The first speech feature of the speech block to be recognized and the first speech feature of a specified speech block are input into a feature encoding subnet, and a first attention score and a second attention score are obtained through an attention encoding layer, and then the second speech feature of the speech block to be recognized is determined, and the second speech feature is input into a decoder to determine the predicted text of the speech block to be recognized. It can be seen that the method of determining the first attention score and the second attention score through the attention encoding layer in the feature encoding subnet can effectively utilize the information of the acoustic context and improve the accuracy of text prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a speech recognition method, apparatus, device, and readable storage medium. Background Art

[0002] With the development of artificial intelligence, the field of human-computer interaction has attracted increasing attention. Among them, speech recognition technology, which converts speech signals into corresponding text, has been widely used in scenarios such as intelligent customer service, unmanned driving, and smart homes.

[0003] Currently, a streaming speech recognition solution can be used to return recognition results in real time during the processing of speech signals to meet the needs of obtaining real-time recognition results in scenarios such as real-time recording of meetings and real-time subtitles of live broadcasts.

[0004] However, existing end-to-end streaming speech recognition solutions cannot effectively utilize acoustic context information, resulting in low accuracy of recognized text. Summary of the Invention

[0005] This specification provides a speech recognition method, apparatus, device, and readable storage medium to partially solve the above-mentioned problems existing in the prior art.

[0006] This manual adopts the following technical solutions:

[0007] This specification provides a speech recognition method, wherein the speech recognition model includes a speech feature encoder and a decoder, the speech feature encoder includes a feature extraction subnet and a feature encoding subnet, and the method includes:

[0008] In response to a streaming speech recognition request, continuously receiving audio data to be recognized;

[0009] Dividing the continuously received audio data to be recognized into speech blocks to be recognized according to a preset duration;

[0010] According to the order in which the speech blocks to be recognized are divided, for each speech block to be recognized, the speech block to be recognized is input into a pre-trained speech recognition model, and the first speech feature of the speech block to be recognized is determined by the feature extraction subnet;

[0011] Determine a previously recognized speech block before the speech block to be recognized as a designated speech block;

[0012] The first speech feature of the speech block to be recognized and the first speech feature of the designated speech block are input into the feature encoding subnet, and a first attention score between the features of each dimension in the first speech feature of the speech block to be recognized and a second attention score between the first speech feature of the designated speech block and the first speech feature of the speech block to be recognized are determined through the attention encoding layer in the feature encoding subnet;

[0013] Determining a second speech feature of the speech block to be recognized based on the first attention score, the second attention score, the first speech feature of the designated speech block, and the first speech feature of the speech block to be recognized;

[0014] The second speech feature of the speech block to be recognized is input into the decoder to obtain a predicted text corresponding to the speech block to be recognized as a recognition result of the speech block to be recognized.

[0015] Optionally, the speech recognition model further includes a corrector;

[0016] Inputting the second speech feature of the speech block to be recognized into the decoder specifically includes:

[0017] Inputting the second speech feature of the speech block to be recognized into the decoder to obtain each predicted text corresponding to the speech block to be recognized and a first probability of each predicted text;

[0018] Determining candidate texts corresponding to the audio data to be recognized based on the predicted texts of the speech blocks to be recognized contained in the audio data to be recognized and the first probabilities of the predicted texts;

[0019] Inputting each candidate text corresponding to the audio data to be recognized and the second speech feature of each speech block to be recognized into the corrector, and obtaining a second probability of each candidate text corresponding to the audio data to be recognized output by the corrector;

[0020] According to the first probability and the second probability, the predicted text corresponding to the audio data to be recognized is selected from the candidate texts as the recognition result of the audio data to be recognized.

[0021] Optionally, before determining each candidate text corresponding to the audio data to be recognized, the method further includes:

[0022] selecting a target text from the predicted texts corresponding to the speech block to be recognized according to the first probabilities of the predicted texts as a recognition result of the speech block to be recognized;

[0023] Returning the recognition result of the speech block to be recognized to the user corresponding to the streaming speech recognition request.

[0024] Optionally, the method further includes:

[0025] The recognition results of the speech blocks to be recognized returned to the user are corrected according to the predicted text corresponding to the audio data to be recognized.

[0026] Optionally, pre-training the speech feature encoder specifically includes:

[0027] Acquire audio data without text annotations in advance, and divide the audio data into a number of speech blocks according to a preset duration;

[0028] For each speech block, the speech block is input into a speech feature encoder to be trained, and a first speech feature of the speech block is determined by a feature extraction subnet in the speech feature encoder;

[0029] Determining a reference speech feature corresponding to the speech block based on the first speech feature of the speech block and the first speech features of several speech blocks before the speech block;

[0030] Inputting a reference speech feature corresponding to the speech block into a feature coding subnet in the speech feature encoder to obtain a second speech feature of the speech block output by the feature coding subnet;

[0031] The speech feature encoder is trained with minimizing the difference between the reference speech feature of the speech block and the second speech feature of the speech block as a training goal.

[0032] Optionally, determining a reference speech feature corresponding to the speech block based on the first speech feature of the speech block and the first speech features of several speech blocks preceding the speech block specifically includes:

[0033] Masking several features of the first speech feature corresponding to the speech block;

[0034] The masked first speech feature of the speech block and the first speech features of several speech blocks before the speech block are fused to obtain a reference speech feature corresponding to the speech block.

[0035] Optionally, the speech feature encoder further includes a quantization subnet;

[0036] The speech feature encoder is trained with minimizing the difference between the first speech feature of the speech block and the second speech feature of the speech block as a training goal, specifically comprising:

[0037] Inputting the first speech feature of the speech block into the quantization subnet to obtain the quantized speech feature of the speech block;

[0038] The speech feature encoder is trained with minimizing the difference between the quantized speech feature of the speech block and the second speech feature of the speech block as a training goal.

[0039] Optionally, inputting the first speech feature of the speech block into the quantization subnet to obtain the quantized speech feature of the speech block specifically includes:

[0040] obtaining a plurality of predetermined codebooks;

[0041] Determining a target feature in each codebook corresponding to the first speech feature of the speech block;

[0042] The feature corresponding to the target feature in the first speech feature of the speech block is replaced with the target feature to obtain the quantized speech feature of the speech block.

[0043] Optionally, training the speech feature encoder with minimizing the difference between the quantized speech feature of the speech block and the second speech feature of the speech block as a training objective specifically includes:

[0044] determining a first loss of the speech block according to a similarity between the quantized speech feature of the speech block and the second speech feature of the speech block;

[0045] Mapping the first speech feature of the speech block to the codebooks to obtain interference quantization features of the speech block;

[0046] determining a second loss of the speech block according to a similarity between the second speech feature of the speech block and the interference quantization feature of the speech block, and a difference between the quantized speech feature of the speech block and the interference quantization feature of the speech block;

[0047] Obtaining a first weight of the first loss and a second weight of the second loss;

[0048] Weighting the first loss of each speech block and the second loss of each speech block according to the first weight and the second weight respectively to obtain a total loss;

[0049] The speech feature encoder is trained with minimization of the total loss as a training objective.

[0050] Optionally, pre-training the decoder specifically includes:

[0051] Acquire audio data with text annotations as training samples, and use the text annotations as annotations of the training samples;

[0052] Inputting the training sample into a trained speech feature encoder to obtain speech features of the training sample;

[0053] Inputting the speech features of the training sample into the decoder to obtain a first predicted text of the training sample;

[0054] The parameters of the decoder are adjusted with minimization of the difference between the first predicted text and the annotation of the training sample as an optimization goal.

[0055] Optionally, the speech recognition model further includes a corrector;

[0056] Adjusting the parameters of the decoder with minimizing the difference between the first predicted text and the annotation of the training sample as an optimization goal specifically includes:

[0057] Inputting the speech features of the training sample and the annotations of the training sample into the corrector, and obtaining a second predicted text of the training sample output by the corrector;

[0058] The parameters of the decoder and the corrector are adjusted with minimization of the difference between the first predicted text and the annotation of the training sample, and minimization of the difference between the second predicted text and the annotation of the training sample as optimization goals.

[0059] This specification provides a speech recognition device, wherein the speech recognition model includes a speech feature encoder and a decoder, the speech feature encoder includes a feature extraction subnet and a feature encoding subnet, and the device includes:

[0060] A receiving module, configured to continuously receive audio data to be recognized in response to a streaming speech recognition request;

[0061] A division module, configured to divide the continuously received audio data to be recognized into speech blocks to be recognized according to a preset duration;

[0062] a first speech feature determination module configured to input each speech block to be recognized into a pre-trained speech recognition model according to the order in which the speech blocks to be recognized are divided, and determine a first speech feature of the speech block to be recognized through the feature extraction subnet;

[0063] a designated speech block determining module, configured to determine a previously recognized speech block before the speech block to be recognized as the designated speech block;

[0064] an attention determination module, configured to take the first speech feature of the speech block to be recognized and the first speech feature of the designated speech block as input, input the first speech feature of the speech block to be recognized and the first speech feature of the designated speech block as input, input the first speech feature of the speech block to be recognized and the second attention score between the first speech feature of the speech block to be recognized and the first speech feature of the speech block to be recognized through the attention coding layer in the feature coding subnet;

[0065] a second speech feature determination module, configured to determine a second speech feature of the speech block to be recognized based on the first attention score, the second attention score, the first speech feature of the designated speech block, and the first speech feature of the speech block to be recognized;

[0066] The decoding module is used to input the second speech feature of the speech block to be recognized into the decoder to obtain the predicted text corresponding to the speech block to be recognized as the recognition result of the speech block to be recognized.

[0067] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned speech recognition method is implemented.

[0068] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned speech recognition method when executing the program.

[0069] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:

[0070] In the speech recognition method provided in this specification, in response to a streaming speech recognition request, the continuously received audio data to be recognized is divided into speech blocks to be recognized according to a preset duration, and the speech blocks to be recognized are sequentially input into a pre-trained speech recognition model in the order of division, a first speech feature is obtained through a feature extraction subnet, the first speech feature of the speech block to be recognized and the first speech feature of a specified speech block of the speech block to be recognized are input into a feature encoding subnet, a first attention score and a second attention score are obtained through an attention encoding layer, and then the second speech feature of the speech block to be recognized is determined, the second speech feature is input into a decoder, and the predicted text of the speech block to be recognized is determined. It can be seen that the method of determining the first attention score and the second attention score through the attention encoding layer in the feature encoding subnet can effectively utilize the information of the acoustic context and improve the accuracy of text prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] The accompanying drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification.

[0072] In the picture:

[0073] Figure 1 A flowchart of a speech recognition method in this specification;

[0074] Figure 2A schematic diagram of a speech recognition model in this specification;

[0075] Figure 3 A schematic diagram of a speech recognition model in this specification;

[0076] Figure 4 A flowchart of a speech recognition method in this specification;

[0077] Figure 5 A flowchart of a speech recognition method in this specification;

[0078] Figure 6 A flowchart of a speech recognition method in this specification;

[0079] Figure 7 A schematic diagram of a speech recognition device provided in this manual;

[0080] Figure 8 The corresponding Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION

[0081] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0082] In addition, it should be noted that all actions of acquiring signals, information or data in the present invention are performed in compliance with the corresponding data protection laws and policies of the country of location and with authorization from the owner of the corresponding device.

[0083] Speech recognition technology aims to automatically convert sound signals into corresponding text content. It is an important entry point in human-computer interaction and has been widely used in various scenarios such as intelligent customer service, unmanned driving, smart home, military communications, etc.

[0084] With the development of deep learning, various end-to-end speech recognition technologies have gradually been proposed, overcoming the modular design and independence assumptions in traditional methods, and becoming an increasingly popular research topic in academia and industry.

[0085] Based on application scenarios, speech recognition can be divided into streaming and non-streaming speech recognition. Streaming speech recognition returns recognition results in real time while processing the user's voice signal, while non-streaming speech recognition requires processing the entire sentence before returning a result. Streaming speech recognition offers low latency, meeting the need for real-time recognition results in scenarios such as real-time conference recording and live captioning of live broadcasts. It also improves the user experience during human-computer voice interaction. However, compared to non-streaming speech recognition, the limited acoustic context information in streaming speech recognition limits its recognition accuracy. Therefore, how to effectively utilize acoustic context information in streaming speech recognition and improve its accuracy has become an urgent issue.

[0086] Based on this, this specification provides a speech recognition method, which can effectively utilize the information of acoustic context and improve the accuracy of text prediction by deploying a feature coding subnet including an attention coding layer in a speech feature encoder, and determining the first attention score and the second attention score according to the first speech feature of the speech block to be recognized and the first speech feature of the specified speech block.

[0087] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0088] Figure 1 This is a flow chart of a speech recognition method provided in this manual.

[0089] S100: In response to a streaming speech recognition request, continuously receive audio data to be recognized.

[0090] The embodiments of this specification provide a speech recognition method, wherein the speech recognition model involved may be pre-trained. The speech recognition method may be executed by an electronic device, such as a server, that processes audio data to generate text. The electronic device that trains the speech recognition model and the electronic device that executes the speech recognition method may be the same or different, and this specification does not limit this.

[0091] Generally, speech recognition can be categorized as either streaming or non-streaming. Non-streaming speech recognition (offline recognition) involves the speech recognition model performing recognition only after receiving the complete audio data to be recognized. Streaming speech recognition, on the other hand, involves the model performing speech recognition simultaneously while continuously receiving the audio data to be recognized. Streaming speech recognition, due to its low latency, is widely used in industry for applications such as dictation and simultaneous interpretation.

[0092] In the embodiments of this specification, the specific technical solution is described in detail by taking the speech recognition model capable of performing streaming speech recognition as an example.

[0093] S102: Divide the continuously received audio data to be recognized into speech blocks to be recognized according to a preset duration.

[0094] This method divides the continuously received audio data into blocks, based on the preset duration and order of audio data reception. Speech recognition is then performed on each block. This method offers the advantages of rapid training and decoding, reducing speech recognition latency and making it well-suited for applications requiring rapid text generation.

[0095] However, a problem with streaming speech recognition based on speech blocks is that, because there is no connection between the speech blocks, the contextual information of the audio data cannot be properly utilized, resulting in poor speech recognition accuracy. To this end, the speech recognition model provided in the embodiments of this specification utilizes an attention encoding layer in the speech feature encoder to fully utilize the relationship between different speech blocks to obtain contextual information between the audio data, thereby improving the accuracy of streaming speech recognition.

[0096] S104: According to the order in which the speech blocks to be recognized are divided, for each speech block to be recognized, the speech block to be recognized is input into a pre-trained speech recognition model, and the first speech feature of the speech block to be recognized is determined through the feature extraction subnet.

[0097] Specifically, the model structure of the speech recognition model provided in the embodiments of this specification can be as follows: Figure 2 As shown, the speech recognition model includes at least a speech feature encoder and a decoder. The speech feature encoder includes a feature extraction subnet and a feature encoding subnet, with an attention encoding layer deployed in the feature encoding subnet. The speech feature encoder is used to extract speech features from the input speech block, and the decoder is used to predict the text corresponding to the speech block based on the speech features of the speech block output by the speech feature encoder.

[0098] Generally, the order in which the speech blocks are divided can be based on the order in which the multiple frames of audio data to be recognized are received. That is, the speech blocks containing the multiple frames of audio data to be recognized received earlier are divided earlier than the speech blocks containing the multiple frames of audio data to be recognized received later. To improve the recognition speed of streaming speech recognition, each time a preset duration of audio data to be recognized is received, it is input into the speech recognition model as a speech block to be recognized.

[0099] In the speech recognition model, the speech feature encoder and decoder are connected in series, meaning the encoder's output serves as the decoder's input. Within the speech feature encoder, the feature extraction subnet and the feature encoding subnet are connected in series, meaning the extraction subnet's output serves as the encoding subnet's input. Therefore, when a speech block to be recognized is input into the speech recognition model, it first passes through the feature extraction subnet in the speech feature encoding process, which then outputs the first speech feature of the block to be recognized.

[0100] Optionally, the feature extraction subnet can be composed of a convolutional network, specifically a two-dimensional convolutional network with a kernel size of 3×3 and a stride of 2, using the ReLU activation function, and a final fully connected layer with an output size equal to the model dimension. The feature extraction subnet is used to downsample the acoustic features of the input speech block to be recognized and model the acoustic features to obtain the first speech feature of the speech block to be recognized.

[0101] S106: Determine the previously recognized speech block before the speech block to be recognized as the designated speech block.

[0102] Furthermore, in order to utilize the acoustic context information in the audio data in the speech feature encoder and improve the accuracy of speech feature extraction, in an embodiment of the present specification, in addition to the current speech block to be recognized, the input to the feature encoding subnet may also include the recognized speech block before the speech block to be recognized, so as to obtain the relationship between the features of the speech block to be recognized and the features of the recognized speech block before the speech block to be recognized, thereby extracting the acoustic context information between the features of multiple consecutive speech blocks.

[0103] Therefore, in the embodiments of this specification, the previously recognized speech block before the speech block to be recognized is determined as the designated speech block for the speech block to be recognized. Of course, depending on the specific application scenario, several recognized speech blocks before the speech block to be recognized can also be determined as the designated speech blocks for the speech block to be recognized, such as two recognized speech blocks between the speech blocks to be recognized as the designated speech blocks for the speech block to be recognized, but this specification does not limit this.

[0104] In addition, it should be noted that the so-called previously recognized speech block before the unrecognized speech block refers to the speech block that was divided before the unrecognized speech block in the speech block division order. When the unrecognized speech block is input into the pre-trained speech recognition model, it will inevitably be input into the speech recognition model. Therefore, the speech blocks before the unrecognized speech block are all recognized speech blocks. For example, if the preset duration of the speech block division is 2 seconds, the speech block from 0 to 2 seconds is the previously recognized speech block before the speech block from 2 to 4 seconds.

[0105] Optionally, the intermediate results and predicted texts output by each subnet or each layer obtained by inputting the recognized speech block into the speech recognition model, such as the first speech feature output by the feature extraction subnet, the second speech feature output by the feature encoding subnet, the predicted text output by the decoder, etc., can all be stored in the database, and a correspondence between the speech block and the intermediate results and predicted texts can be established, so that when the intermediate results or predicted texts of the recognized speech block are needed, they can be directly extracted from the database without going through the speech recognition model output again, thereby reducing the delay of speech recognition.

[0106] S108: The first speech feature of the speech block to be recognized and the first speech feature of the designated speech block are input as input to the feature coding subnet, and the first attention score between the features of each dimension in the first speech feature of the speech block to be recognized and the second attention score between the first speech feature of the designated speech block and the first speech feature of the speech block to be recognized are determined through the attention coding layer in the feature coding subnet.

[0107] Specifically, the feature encoding subnet can be a neural network containing several layers of Conformer encoding layers. Each Conformer encoding layer is composed of a feedforward network, a block multi-head local self-attention mechanism and a causal convolution that does not pay attention to the right context. It is used to model the context dependency of the first speech feature output from the feature extraction subnet and output the second speech feature.

[0108] Specifically, the attention encoding layer deployed in the feature encoding subnet is the encoding layer of the above-mentioned block multi-head local self-attention mechanism, which introduces a block multi-head local self-attention mechanism with relative position encoding. The first speech feature of the speech block to be recognized and the first speech feature of the specified speech block are input into the attention encoding layer. The first attention score between the various dimensional features within the block in the first speech feature of the speech block to be recognized can be determined, and the attention score of the current block of the first speech feature of the speech block to be recognized can only be calculated with the previous block. Among them, the first attention score can represent the correlation between the various dimensional features in the first speech feature of the speech block to be recognized, and the second attention score can represent the correlation between the first speech feature of the speech block to be recognized and the first speech feature of the specified speech block.

[0109] Specifically, the first attention score used to characterize the correlation between the various dimensional features in the first speech feature of the speech block to be recognized can be determined by the similarity between the various dimensional features in the first speech feature of the speech block to be recognized. Similarly, the second attention score used to characterize the correlation between the first speech feature of the speech block to be recognized and the first speech feature of the specified speech block can be determined by the similarity between the first speech feature of the speech block to be recognized and the first speech feature of the specified speech block. Of course, other existing attention score determination methods can also be used, and this specification does not limit this.

[0110] S110: Determine the second speech feature of the speech block to be recognized based on the first attention score, the second attention score, the first speech feature of the designated speech block, and the first speech feature of the speech block to be recognized.

[0111] By introducing the attention coding layer into the feature coding subnet, the first attention score and the second attention score are determined. This not only makes full use of the acoustic information of the intra-block features of the speech block to be recognized, but also makes use of the contextual acoustic information of the features of the previous recognized speech block of the specified speech block to be recognized. This effectively utilizes the acoustic contextual information in the streaming audio data and improves the accuracy of speech recognition.

[0112] S112: Inputting the second speech feature of the speech block to be recognized into the decoder to obtain a predicted text corresponding to the speech block to be recognized as a recognition result of the speech block to be recognized.

[0113] In the speech recognition method provided in this specification, in response to a streaming speech recognition request, the continuously received audio data to be recognized is divided into speech blocks to be recognized according to a preset duration, and the speech blocks to be recognized are sequentially input into a pre-trained speech recognition model in the order of division, a first speech feature is obtained through a feature extraction subnet, the first speech feature of the speech block to be recognized and the first speech feature of the designated speech block of the speech block to be recognized are input into a feature encoding subnet, a first attention score and a second attention score are obtained through an attention encoding layer, and then the second speech feature of the speech block to be recognized is determined, the second speech feature is input into a decoder, and the predicted text of the speech block to be recognized is determined. It can be seen that by deploying a feature encoding subnet including an attention encoding layer in a speech feature encoder, and determining the first attention score and the second attention score according to the first speech feature of the speech block to be recognized and the first speech feature of the designated speech block, it is possible to effectively utilize the information of the acoustic context and improve the accuracy of text prediction.

[0114] In one or more embodiments of this specification, in order to further improve the utilization rate of acoustic context information, a corrector can also be deployed in the speech recognition model to re-score the predicted text of each speech block to be recognized by the corrector, and correct the predicted text of the speech block to be recognized to obtain a more accurate streaming speech recognition result. Figure 2 The speech recognition model shown in the figure can be used as follows after deploying the corrector. Figure 3 shown.

[0115] Therefore, in Figure 1 As shown in step S112, the second speech feature of the speech block to be recognized is input into the decoder. When the speech recognition model also includes a corrector, the specific steps are as follows: Figure 4 As shown:

[0116] S200: Inputting the second speech feature of the speech block to be recognized into the decoder to obtain each predicted text corresponding to the speech block to be recognized and a first probability of each predicted text.

[0117] Specifically, the second speech feature of the speech block to be recognized output by the speech feature encoder is input into the decoder, and the beam search algorithm is used in the decoder for streaming decoding to obtain each predicted text as a streaming decoding result, and obtain the first probability of each predicted text.

[0118] Alternatively, a vocabulary can be pre-set and classified in a decoder to determine the probability that the second speech feature of the speech block to be recognized corresponds to each word in the vocabulary. A beam search algorithm can then be used to determine multiple predicted texts, and a first probability for each predicted text can be determined based on the words contained in each predicted text and the probability of each word. The number of predicted texts can be pre-set and determined based on the specific application scenario, and this specification does not limit this.

[0119] S202: Determine candidate texts corresponding to the audio data to be recognized based on the predicted texts of the speech blocks to be recognized contained in the audio data to be recognized and the first probabilities of the predicted texts.

[0120] Since streaming speech recognition simultaneously receives the audio data to be recognized and performs speech recognition, after receiving the current audio data segment, the speech blocks corresponding to the current audio data segment are obtained. At this point, each speech block corresponding to the current audio data segment is input into the speech recognition model, and predicted texts are obtained for each speech block. Based on the first probabilities corresponding to the predicted texts for each speech block, the predicted texts with the highest first probabilities are selected as candidates for the speech block. This process then iterates through each speech block in the current audio data segment to obtain candidate texts corresponding to the current audio data segment.

[0121] S204: Inputting each candidate text corresponding to the audio data to be recognized and the second speech feature of each speech block to be recognized into the corrector, and obtaining a second probability of each candidate text corresponding to the audio data to be recognized output by the corrector.

[0122] Specifically, a start mark is added to the beginning of each candidate text corresponding to the audio data to be recognized, and then each candidate text with the start mark added and the second speech feature of each speech block to be recognized output by the aforementioned speech feature encoder are taken as input and input into the trained corrector. The corrector predicts each candidate text that contains the end mark but does not contain the start mark, and sums the conditional probability of each position of each candidate text to obtain the second probability of each candidate text.

[0123] S206: Selecting, from the candidate texts, a predicted text corresponding to the audio data to be recognized as a recognition result of the audio data to be recognized, based on the first probability and the second probability.

[0124] Furthermore, the second probability of each candidate text and the first probability of each predicted text corresponding to each speech block to be recognized contained in the audio data to be recognized are combined to determine the total probability of each candidate text, and the candidate text with the highest total probability is used as the predicted text corresponding to the audio data to be recognized.

[0125] Optionally, the weight of the first probability and the weight of the second probability can be determined separately, and the total probability can be obtained by weighted summing the weights of the first probability and the second probability. The specific weights can be predetermined based on the specific application scenario, and this specification does not limit this.

[0126] Based on Figure 4The illustrated speech recognition method deploys a corrector within the speech recognition model. The corrector inputs candidate texts corresponding to each speech block to be recognized contained in the audio data to be recognized, as well as the second speech features of each speech block to be recognized. The corrector then outputs a second probability for each candidate text corresponding to the audio data to be recognized. Furthermore, based on the first probability of each predicted text for each speech block to be recognized and the second probability of each candidate text, the predicted text of the audio data to be recognized is selected from the candidate texts as the recognition result for the audio data to be recognized. This indicates that by using the corrector to calculate the second probability of each candidate text corresponding to the audio data to be recognized, the predicted text of each speech block to be recognized is rescored, the predicted text of the speech block to be recognized is corrected, and a more accurate streaming speech recognition result is obtained.

[0127] In one or more embodiments of this specification, Figure 4 Before determining the candidate texts corresponding to the audio data to be recognized as shown in step S204, the recognition result of the speech block to be recognized can also be determined and returned to the user corresponding to the streaming speech recognition request to improve the efficiency and visualization of streaming speech recognition. Specifically, the determination is made through the following scheme:

[0128] First, according to the first probabilities of the predicted texts, a target text is selected from the predicted texts corresponding to the speech block to be recognized as a recognition result of the speech block to be recognized.

[0129] Secondly, the recognition result of the speech block to be recognized is returned to the user corresponding to the streaming speech recognition request.

[0130] Specifically, in the application scenario of streaming speech recognition, which requires streaming speech recognition to output real-time audio decoding text, in order to reduce the delay of speech recognition and improve the real-time display of speech recognition text, the predicted text with the highest first probability among the predicted texts of each speech block to be recognized can be returned to the user corresponding to the streaming speech recognition request as the recognition result of the speech block to be recognized. The recognition result of each speech block to be recognized can be displayed to the user by displaying text. By returning and displaying in real time, the user can observe the speech recognition result in real time, reducing the delay of speech recognition.

[0131] Further, in an optional embodiment of this specification, in the following Figure 4 After the recognition result of the audio data to be recognized is determined as shown in step S208, the recognition result of the audio data to be recognized may be returned to the user so as to correct the recognition result of each speech block to be recognized returned to the user in the above solution.

[0132] For example, the audio data to be recognized continuously received by the speech recognition model is divided into three speech blocks to be recognized. The predicted texts with the highest first probability among the predicted texts corresponding to these three speech blocks are: "today", "sky", and "sunny". Therefore, when the speech recognition model obtains these predicted texts, it returns these three predicted texts to the user in sequence. At this time, the predicted texts corresponding to this segment of the audio data to be recognized that the user can observe are "today", "sky", and "sunny" in sequence. Then, after completing the output of the predicted texts for all speech blocks to be recognized contained in this segment of the audio data to be recognized, the corrector re-scores each predicted text and finds that the recognition result corresponding to this segment of the audio data to be recognized is actually "today is sunny". At this time, the recognition result corresponding to the audio data to be recognized can also be returned to the user, and the recognition result corresponding to the audio data to be recognized can be used as the correct recognition result to correct the recognition results of the speech blocks to be recognized previously returned and displayed to the user. That is, "today is sunny" is used to correct "today", "sky", and "sunny".

[0133] In one or more embodiments of this specification, Figure 1 As shown in step S104, according to the order in which the speech blocks to be recognized are divided, for each speech block to be recognized, before the speech block to be recognized is input into the pre-trained speech recognition model, the speech recognition model needs to be pre-trained. The specific steps are as follows: Figure 5 As shown:

[0134] S300: pre-acquire audio data without text annotations, and divide the audio data into a plurality of speech blocks according to a preset duration.

[0135] With the development of deep learning, a variety of end-to-end speech recognition technologies have been proposed, overcoming the modular design and independence assumptions of traditional methods and becoming an increasingly popular research topic in academia and industry. However, deep learning-based speech recognition models primarily rely on data-driven optimization training, and recognition performance is largely dependent on the amount of labeled training data available. With limited training data, speech recognition often fails to achieve ideal performance.

[0136] In order to solve the problem of limited number of training samples with text annotations, the speech feature encoder in the embodiment of this specification can be trained using self-supervised learning. Specifically, the acquired audio data can be audio data without text annotations.

[0137] Since the speech feature encoder provided in this specification needs to be applied in the application scenario of streaming speech recognition, the audio data used as training samples is divided into several speech blocks according to the predicted duration, similar to the application of streaming speech recognition.

[0138] The preset duration here may be the same as or different from the division duration corresponding to the speech block to be recognized when the speech feature encoder is applied, and this specification does not limit this.

[0139] S302: For each speech block, input the speech block into a speech feature encoder to be trained, and determine a first speech feature of the speech block through a feature extraction subnet in the speech feature encoder.

[0140] Specifically, the feature extraction subnet can be composed of a convolutional network, specifically a two-dimensional convolutional network with a kernel size of 3×3 and a stride of 2. Its activation function uses the ReLU function, and it also includes a final fully connected layer with an output size equal to the model dimension. The feature extraction subnet is used to downsample the acoustic features of the input speech block and model the acoustic features to obtain the first speech feature of the speech block.

[0141] S304: Determine a reference speech feature corresponding to the speech block according to the first speech feature of the speech block and the first speech features of several speech blocks before the speech block.

[0142] Furthermore, in order to utilize the acoustic context information in the audio data during the training process of the speech feature encoder and improve the accuracy of speech feature extraction, in an embodiment of the present specification, in addition to the current speech block, the input to the feature encoding subnet can also include the speech block before the speech block, so as to obtain the relationship between the features of the speech block and the features of the speech block before the speech block, thereby extracting the acoustic context information between the features of multiple consecutive speech blocks.

[0143] The number of the several speech blocks preceding the speech block can be determined according to a specific application scenario and is at least one, and this specification does not impose any specific limitation on this.

[0144] S306: Input the reference speech feature corresponding to the speech block into the feature coding subnet in the speech feature encoder to obtain the second speech feature of the speech block output by the feature coding subnet.

[0145] Specifically, the feature encoding subnet can be a neural network containing several layers of Conformer encoding layers. Each Conformer encoding layer is composed of a feedforward network, a block multi-head local self-attention mechanism and a causal convolution that does not pay attention to the right context. It is used to model the context dependency of the first speech feature output from the feature extraction subnet and output the second speech feature.

[0146] S308: Training the speech feature encoder with minimizing the difference between the reference speech feature of the speech block and the second speech feature of the speech block as a training goal.

[0147] Further, in one or more embodiments of this specification, in Figure 5 Step S304 shows determining a reference speech feature corresponding to the speech block based on the first speech feature of the speech block and the first speech features of several speech blocks preceding the speech block, which is specifically implemented by the following scheme:

[0148] Several features of the first speech features corresponding to the speech block are masked, and the masked first speech features of the speech block are fused with the first speech features of several speech blocks before the speech block to obtain reference speech features corresponding to the speech block.

[0149] In the feature encoding subnet, several features of the first speech features corresponding to the speech block may be masked. These features may be continuous or discontinuous, and this specification does not limit this. The reference speech features corresponding to the speech block obtained by splicing and fusing the masked first speech features of the speech block with the speech features of several speech blocks between the speech blocks can be used to train the speech feature encoder based on the difference between the reference speech features corresponding to the speech block and the second speech features corresponding to the speech block, with the goal of predicting the features of the masked portion.

[0150] Further, in one or more embodiments of this specification, in Figure 5 As shown in step S308, the training goal is to minimize the difference between the reference speech feature of the speech block and the second speech feature of the speech block. In training the speech feature encoder, a quantization subnet can be deployed in the speech feature encoder. After determining the reference speech feature of the speech block, the quantized speech feature of the speech block is obtained based on the quantization subnet, and then the speech feature encoder is trained based on the quantized speech feature and the second speech feature. The specific scheme is as follows:

[0151] First, a plurality of predetermined codebooks are obtained.

[0152] Secondly, a target feature corresponding to the reference speech feature of the speech block in each codebook is determined.

[0153] Then, the feature corresponding to the target feature in the reference speech feature of the speech block is replaced with the target feature to obtain the quantized speech feature of the speech block.

[0154] Finally, the speech feature encoder is trained with minimizing the difference between the quantized speech feature of the speech block and the second speech feature of the speech block as a training goal.

[0155] Furthermore, based on the feature encoder deployed with the quantization subnetwork, minimizing the difference between the quantized speech feature of the speech block and the second speech feature of the speech block is used as a training goal. During the training of the speech feature encoder, a loss function can be determined based on the quantized speech feature and the second speech feature of the speech block. Then, minimizing the loss function is used as a training goal to train the speech feature encoder. The specific scheme is as follows:

[0156] Step 1: Determine a first loss of the speech block according to a similarity between the quantized speech feature of the speech block and the second speech feature of the speech block.

[0157] The first loss of the speech block represents that one of the training objectives of the speech feature encoder is to predict the quantized feature value of the masked part of the feature. The correlation between the similarity between the quantized speech feature of the speech block and the second speech feature of the speech block and the first loss can be a negative correlation, that is, the maximization of the similarity between the quantized speech feature of the speech block and the second speech feature of the speech block corresponds to the minimization of the first loss of the speech block.

[0158] Step 2: Map the first speech feature of the speech block to the codebooks to obtain interference quantization features of the speech block.

[0159] Step 3: Determine the second loss of the speech block according to the similarity between the second speech feature of the speech block and the interference quantization feature of the speech block, and the difference between the quantized speech feature of the speech block and the interference quantization feature of the speech block.

[0160] The second loss of the speech block indicates that one of the training objectives of the speech feature encoder is to minimize the similarity between the second speech feature output by the speech feature encoder and the interference quantized feature, and maximize the similarity between the second speech feature and the quantized speech feature. The second loss of the speech block can be used as an additional diversity loss penalty.

[0161] Step 4: Obtain a first weight of the first loss and a second weight of the second loss.

[0162] Step 5: Weight the first loss of each speech block and the second loss of each speech block according to the first weight and the second weight respectively to obtain the total loss.

[0163] Step 6: Taking minimization of the total loss as a training objective, train the speech feature encoder.

[0164] Furthermore, after the speech feature encoder is trained in a self-supervised learning manner, the decoder can be trained based on the trained speech feature encoder. The specific steps are as follows: Figure 6 As shown:

[0165] S400: Acquire audio data with text annotations as training samples, and use the text annotations as annotations of the training samples.

[0166] Specifically, after the speech feature encoder is trained using self-supervised learning, the decoder can be trained using supervised learning using the trained speech feature encoder. Because the speech feature encoder is already trained, it can effectively extract higher-level acoustic features from speech blocks, reducing the downstream task (i.e., the decoder)'s reliance on labeled training data. Therefore, the number of text-annotated training samples required for this step is significantly reduced compared to the audio data required for the aforementioned pre-training process, reducing the pressure on obtaining text-annotated audio data.

[0167] S402: Input the training sample into a trained speech feature encoder to obtain speech features of the training sample.

[0168] Generally, the trained speech feature encoder does not need to determine the speech quantization features. Therefore, in the process of training the encoder with the trained speech feature encoder and in the application process of the speech feature encoder, there is no need to remove the quantization subnet in the speech feature encoder.

[0169] S404: Inputting the speech features of the training sample into the decoder to obtain a first predicted text of the training sample.

[0170] S406: Adjusting the parameters of the decoder with minimization of the difference between the first predicted text and the annotation of the training sample as an optimization goal.

[0171] The speech features of the training samples obtained above are input into the decoder to obtain output vectors with the same number of features as the speech features of the training samples. The dimension of each output vector is the same as the size of the vocabulary. The text probability distribution vector is then calculated using the Softmax function to determine the first predicted text of the training sample. The loss is determined by the difference between the first predicted text and the annotation of the training sample. The parameters of the decoder are adjusted with the minimization of the loss as the optimization goal.

[0172] Furthermore, based on Figure 4 As shown in the description, a corrector can also be deployed in the speech recognition model composed of a speech feature encoder and a decoder, and the corrector also needs to be trained before speech recognition can be performed. Therefore, the corrector can be trained jointly with the encoder. Figure 6 In step S406, the optimization goal is to minimize the difference between the first predicted text and the annotation of the training sample, and the corrector is jointly trained based on adjusting the parameters of the decoder. The specific scheme is as follows:

[0173] First, the speech features of the training sample and the annotations of the training sample are input into the corrector, and a second predicted text of the training sample is obtained as output by the corrector.

[0174] The speech features of the training sample obtained by the trained speech feature encoder are input into the corrector, and the annotations of the training sample are also input into the corrector, so that the corrector outputs the second predicted text of the training sample in a Teacher Forcing manner.

[0175] Then, the parameters of the decoder and the corrector are adjusted with minimization of the difference between the first predicted text and the annotation of the training sample, and minimization of the difference between the second predicted text and the annotation of the training sample as optimization goals.

[0176] Specifically, the rescoring corrector is constructed using the Transformer decoder, which includes a word embedding calculation module, a position encoding calculation module, multiple layers of Transformer decoding layers, and a fully connected layer.

[0177] The Transformer decoding layer consists of a masked self-attention mechanism, a cross-attention mechanism, and a feedforward network. During the Teacher Forcing calculation process, a start marker is first added to the annotation of the input training sample. A word embedding module is then used to obtain a vector representation of the annotation. This vector representation is then added to its positional encoding vector and input into the Transformer decoding layer. Masked self-attention and cross-attention with the speech features of the training sample are then calculated in the decoding layer to effectively utilize the global acoustic context of the audio data used as the training sample. Finally, after passing through the fully connected layer, the resulting output vector has the same dimensions as the vocabulary and the same number of vectors as the annotation of the input training sample with the start marker added. The Softmax function is then used to calculate the text probability distribution vector of the output vector over each word in the vocabulary, thereby determining the second predicted text predicted by the corrector, which contains the end marker but not the start marker.

[0178] Furthermore, a first loss is determined based on the difference between the first predicted text and the annotation of the training sample, and a second loss is determined based on the difference between the second predicted text and the annotation of the training sample. The weight of the first loss and the weight of the second loss are obtained, and the first loss and the second loss are weighted respectively to obtain the total loss.

[0179] Then, the parameters of the decoder and corrector are adjusted with the goal of minimizing the total loss.

[0180] The weight of the first loss and the weight of the second loss in the total loss may be pre-set, or may be adjusted while adjusting the parameters of the decoder and the corrector, which is not limited in this specification.

[0181] Figure 7 This is a schematic diagram of a speech recognition device provided in this specification. The speech recognition model includes a speech feature encoder and a decoder. The speech feature encoder includes a feature extraction subnet and a feature encoding subnet, specifically including:

[0182] The receiving module 500 is configured to continuously receive audio data to be recognized in response to a streaming speech recognition request;

[0183] The division module 502 is configured to divide the continuously received audio data to be recognized into speech blocks to be recognized according to a preset duration;

[0184] A first speech feature determination module 504 is configured to input each speech block to be recognized into a pre-trained speech recognition model according to the order in which the speech blocks to be recognized are divided, and determine the first speech feature of the speech block to be recognized through the feature extraction subnet;

[0185] The designated speech block determining module 506 is configured to determine a previously recognized speech block before the speech block to be recognized as the designated speech block;

[0186] an attention determination module 508 for inputting the first speech feature of the speech block to be recognized and the first speech feature of the designated speech block as input to the feature encoding subnet, and determining, through the attention encoding layer in the feature encoding subnet, a first attention score between the various dimensional features in the first speech feature of the speech block to be recognized, and a second attention score between the first speech feature of the designated speech block and the first speech feature of the speech block to be recognized;

[0187] A second speech feature determination module 510 is configured to determine a second speech feature of the speech block to be recognized based on the first attention score, the second attention score, the first speech feature of the designated speech block, and the first speech feature of the speech block to be recognized;

[0188] The decoding module 512 is configured to input the second speech feature of the speech block to be recognized into the decoder, and obtain the predicted text corresponding to the speech block to be recognized as the recognition result of the speech block to be recognized.

[0189] Optionally, the speech recognition model further includes a corrector;

[0190] Optionally, the decoding module 512 is specifically used to input the second speech feature of the speech block to be recognized into the decoder to obtain each predicted text corresponding to the speech block to be recognized and the first probability of each predicted text; determine each candidate text corresponding to the audio data to be recognized based on the predicted text of each speech block to be recognized contained in the audio data to be recognized and the first probability of each predicted text; input each candidate text corresponding to the audio data to be recognized and the second speech feature of each speech block to be recognized into the corrector to obtain the second probability of each candidate text corresponding to the audio data to be recognized output by the corrector; and select the predicted text corresponding to the audio data to be recognized from the candidate texts as the recognition result of the audio data to be recognized based on the first probability and the second probability.

[0191] Optionally, the device further comprises:

[0192] The first return module 514 is specifically configured to select a target text from the predicted texts corresponding to the speech block to be recognized based on the first probabilities of the predicted texts as the recognition result of the speech block to be recognized; and return the recognition result of the speech block to the user corresponding to the streaming speech recognition request.

[0193] Optionally, the device further comprises:

[0194] The second returning module 516 is specifically configured to correct the recognition results of the speech blocks to be recognized that are returned to the user according to the predicted text corresponding to the audio data to be recognized.

[0195] Optionally, the device further comprises:

[0196] The first training module 518 is specifically used to pre-acquire audio data without text annotations, and divide the audio data into several speech blocks according to a preset time length; for each speech block, input the speech block into the speech feature encoder to be trained, and determine the first speech feature of the speech block through the feature extraction subnet in the speech feature encoder; determine the reference speech feature corresponding to the speech block based on the first speech feature of the speech block and the first speech features of several speech blocks before the speech block; input the reference speech feature corresponding to the speech block into the feature encoding subnet in the speech feature encoder, and obtain the second speech feature of the speech block output by the feature encoding subnet; train the speech feature encoder with the training goal of minimizing the difference between the reference speech feature of the speech block and the second speech feature of the speech block.

[0197] Optionally, the first training module 518 is specifically used to mask several features of the first speech features corresponding to the speech block; and fuse the masked first speech features of the speech block with the first speech features of several speech blocks before the speech block to obtain the reference speech features corresponding to the speech block.

[0198] Optionally, the speech feature encoder further includes a quantization subnet;

[0199] Optionally, the first training module 518 is specifically used to input the first speech feature of the speech block into the quantization subnet to obtain the quantized speech feature of the speech block; and train the speech feature encoder with the training goal of minimizing the difference between the quantized speech feature of the speech block and the second speech feature of the speech block.

[0200] Optionally, the first training module 518 is specifically used to obtain a plurality of predetermined codebooks; determine a target feature in each codebook corresponding to the first speech feature of the speech block; replace the feature corresponding to the target feature in the first speech feature of the speech block with the target feature to obtain the quantized speech feature of the speech block.

[0201] Optionally, the first training module 518 is specifically used to determine the first loss of the speech block based on the similarity between the quantized speech feature of the speech block and the second speech feature of the speech block; map the first speech feature of the speech block to the codebooks to obtain the interference quantization feature of the speech block; determine the second loss of the speech block based on the similarity between the second speech feature of the speech block and the interference quantization feature of the speech block, and the difference between the quantized speech feature of the speech block and the interference quantization feature of the speech block; obtain a first weight of the first loss and a second weight of the second loss; weight the first loss of each speech block and the second loss of each speech block according to the first weight and the second weight to obtain a total loss; and train the speech feature encoder with minimization of the total loss as the training goal.

[0202] Optionally, the device further comprises:

[0203] The second training module 520 is specifically used to obtain audio data with text annotations as training samples, and use the text annotations as annotations of the training samples; input the training samples into the trained speech feature encoder to obtain the speech features of the training samples; input the speech features of the training samples into the decoder to obtain the first predicted text of the training samples; and adjust the parameters of the decoder with minimizing the difference between the first predicted text and the annotation of the training sample as the optimization goal.

[0204] Optionally, the speech recognition model further includes a corrector;

[0205] Optionally, the second training module 520 is specifically used to input the speech features of the training sample and the annotations of the training sample into the corrector to obtain a second predicted text of the training sample output by the corrector; and adjust the parameters of the decoder and the corrector with the minimization of the difference between the first predicted text and the annotations of the training sample, and the minimization of the difference between the second predicted text and the annotations of the training sample as optimization goals.

[0206] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 The speech recognition method shown.

[0207] This manual also provides Figure 8 The schematic structure diagram of the electronic device shown in FIG. Figure 8 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0208] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0209] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing various functions included therein can also be considered as structures within the hardware component. Or even, the means for implementing various functions can be considered as both a software module implementing the method and a structure within the hardware component.

[0210] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0211] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0212] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0213] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0214] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0215] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0216] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0217] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0218] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0219] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0220] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0221] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0222] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0223] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A speech recognition method, It is characterized in that The speech recognition model includes a speech feature encoder and a decoder, the speech feature encoder includes a feature extraction subnet and a feature encoding subnet, and the method includes: In response to a streaming speech recognition request, continuously receiving audio data to be recognized; Dividing the continuously received audio data to be recognized into speech blocks to be recognized according to a preset duration; According to the order in which the speech blocks to be recognized are divided, for each speech block to be recognized, the speech block to be recognized is input into a pre-trained speech recognition model, and a first speech feature of the speech block to be recognized is determined through the feature extraction subnet; Determine a previously recognized speech block of the speech block to be recognized as a designated speech block; The first speech feature of the speech block to be recognized and the first speech feature of the designated speech block are input into the feature encoding subnet, and a first attention score between the features of each dimension in the first speech feature of the speech block to be recognized and a second attention score between the first speech feature of the designated speech block and the first speech feature of the speech block to be recognized are determined through the attention encoding layer in the feature encoding subnet; Determine the second speech feature of the speech block to be recognized according to the first attention score, the second attention score, the first speech feature of the designated speech block, and the first speech feature of the speech block to be recognized; The second speech feature of the speech block to be recognized is input into the decoder to obtain the predicted text corresponding to the speech block to be recognized as the recognition result of the speech block to be recognized.

2. The method according to claim 1, It is characterized in that The speech recognition model also includes a corrector; Inputting the second speech feature of the speech block to be recognized into the decoder specifically comprises: Inputting the second speech feature of the speech block to be recognized into the decoder to obtain each predicted text corresponding to the speech block to be recognized and a first probability of each predicted text; Determine candidate texts corresponding to the audio data to be recognized according to the predicted texts of the speech blocks to be recognized contained in the audio data to be recognized and the first probabilities of the predicted texts; Inputting each candidate text corresponding to the audio data to be recognized and the second speech feature of each speech block to be recognized into the corrector, and obtaining a second probability of each candidate text corresponding to the audio data to be recognized output by the corrector; According to the first probability and the second probability, the predicted text corresponding to the audio data to be recognized is selected from the candidate texts as the recognition result of the audio data to be recognized.

3. The method according to claim 2, It is characterized in that Before determining each candidate text corresponding to the audio data to be recognized, the method further includes: According to the first probabilities of the predicted texts, selecting a target text from the predicted texts corresponding to the speech block to be recognized as a recognition result of the speech block to be recognized; The recognition result of the speech block to be recognized is returned to the user corresponding to the streaming speech recognition request.

4. The method according to claim 3, It is characterized in that The method further comprises: According to the predicted text corresponding to the audio data to be recognized, the recognition results of the speech blocks to be recognized returned to the user are corrected.

5. The method according to claim 1, It is characterized in that Pre-training the speech feature encoder specifically includes: Acquire audio data without text annotations in advance, and divide the audio data into a plurality of speech blocks according to a preset duration; For each speech block, the speech block is input into a speech feature encoder to be trained, and a first speech feature of the speech block is determined by a feature extraction subnet in the speech feature encoder; Determine a reference speech feature corresponding to the speech block according to the first speech feature of the speech block and the first speech features of several speech blocks before the speech block; Inputting a reference speech feature corresponding to the speech block into a feature coding subnet in the speech feature encoder to obtain a second speech feature of the speech block output by the feature coding subnet; The speech feature encoder is trained with minimizing the difference between the reference speech feature of the speech block and the second speech feature of the speech block as a training goal.

6. The method according to claim 5, It is characterized in that Determining a reference speech feature corresponding to the speech block according to the first speech feature of the speech block and the first speech features of several speech blocks before the speech block specifically includes: Masking some features of the first speech features corresponding to the speech block; The masked first speech feature of the speech block and the first speech features of several speech blocks before the speech block are fused to obtain a reference speech feature corresponding to the speech block.

7. The method according to claim 5, It is characterized in that The speech feature encoder also includes a quantization subnet; The speech feature encoder is trained with minimizing the difference between the reference speech feature of the speech block and the second speech feature of the speech block as a training target, specifically comprising: Inputting the reference speech feature of the speech block into the quantization subnet to obtain the quantized speech feature of the speech block; The speech feature encoder is trained with minimizing the difference between the quantized speech feature of the speech block and the second speech feature of the speech block as a training goal.

8. The method according to claim 7, It is characterized in that Inputting the first speech feature of the speech block into the quantization subnet to obtain the quantized speech feature of the speech block specifically includes: Obtaining a plurality of predetermined codebooks; Determining a target feature in each codebook corresponding to the first speech feature of the speech block; The feature corresponding to the target feature in the first speech feature of the speech block is replaced with the target feature to obtain the quantized speech feature of the speech block.

9. The method according to claim 8, It is characterized in that Taking minimizing the difference between the quantized speech feature of the speech block and the second speech feature of the speech block as a training goal, training the speech feature encoder specifically includes: Determine a first loss of the speech block according to a similarity between the quantized speech feature of the speech block and the second speech feature of the speech block; Mapping the first speech feature of the speech block to the codebooks to obtain interference quantization features of the speech block; Determine a second loss of the speech block according to a similarity between the second speech feature of the speech block and the interference quantization feature of the speech block, and a difference between the quantized speech feature of the speech block and the interference quantization feature of the speech block; Obtaining a first weight of the first loss and a second weight of the second loss; According to the first weight and the second weight, the first loss of each speech block and the second loss of each speech block are weighted respectively to obtain a total loss; The speech feature encoder is trained with minimization of the total loss as a training objective.

10. The method according to claim 1, It is characterized in that Pre-training the decoder specifically includes: Acquire audio data with text annotations as training samples, and use the text annotations as annotations of the training samples; Inputting the training sample into a trained speech feature encoder to obtain speech features of the training sample; Inputting the speech features of the training sample into the decoder to obtain a first predicted text of the training sample; The parameters of the decoder are adjusted with minimization of the difference between the first predicted text and the annotation of the training sample as an optimization goal.

11. The method according to claim 10, It is characterized in that The speech recognition model also includes a corrector; Taking minimization of the difference between the first predicted text and the annotation of the training sample as an optimization goal, adjusting the parameters of the decoder specifically includes: Inputting the speech features of the training sample and the annotations of the training sample into the corrector to obtain a second predicted text of the training sample output by the corrector; The parameters of the decoder and the corrector are adjusted with minimization of the difference between the first predicted text and the annotation of the training sample and minimization of the difference between the second predicted text and the annotation of the training sample as optimization goals.

12. A speech recognition device, It is characterized in that The speech recognition model includes a speech feature encoder and a decoder, the speech feature encoder includes a feature extraction subnet and a feature encoding subnet, and the device includes: A receiving module, used for continuously receiving audio data to be recognized in response to a streaming speech recognition request; A division module, used for dividing the continuously received audio data to be recognized into speech blocks to be recognized according to a preset duration; A first speech feature determination module is used to input each speech block to be recognized into a pre-trained speech recognition model according to the order in which the speech blocks to be recognized are divided, and determine the first speech feature of the speech block to be recognized through the feature extraction subnet; A designated speech block determination module is used to determine a previously recognized speech block of the speech block to be recognized as the designated speech block; An attention determination module is used to take the first speech feature of the speech block to be recognized and the first speech feature of the designated speech block as input, input them into the feature encoding subnet, and determine the first attention score between the features of each dimension in the first speech feature of the speech block to be recognized, and the second attention score between the first speech feature of the designated speech block and the first speech feature of the speech block to be recognized through the attention encoding layer in the feature encoding subnet; A second speech feature determination module, configured to determine a second speech feature of the speech block to be recognized based on the first attention score, the second attention score, the first speech feature of the designated speech block, and the first speech feature of the speech block to be recognized; The decoding module is used to input the second speech feature of the speech block to be recognized into the decoder to obtain the predicted text corresponding to the speech block to be recognized as the recognition result of the speech block to be recognized.

13. A computer-readable storage medium, It is characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Speech recognition decoding method and device based on streaming attention model, equipment and computer readable storage medium

    CN112242144A

  • End-to-end online voice detection and recognition method and system, and equipment

    CN112951213A