Methods, apparatus, devices, and readable media for speech recognition

CN117351963BActive Publication Date: 2026-08-14JINGDONG CITY BEIJING DIGITS TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

[0007]应当理解,发明内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117351963B_ABST
    Figure CN117351963B_ABST
Patent Text Reader

Abstract

Embodiments of this disclosure relate to methods, apparatus, devices, and readable media for speech recognition. The method includes: adding at least one first node representing at least one first candidate text sequence to a prefix tree based on at least one first candidate text sequence identified from a first speech; adding at least one second node representing at least one second candidate text sequence to the prefix tree based on at least one second candidate text sequence identified from a second speech; determining scores corresponding to each of a plurality of text sequences; deleting the at least one first path from the prefix tree to delete at least one first node, obtaining an updated prefix tree; and determining a first target text sequence matching the first speech from at least one first candidate text sequence represented by the at least one first candidate text sequence that was not deleted, based at least on the updated prefix tree. This can improve the efficiency of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computer technology, and more specifically, to methods, apparatus, devices, and computer-readable storage media for speech recognition. Background Technology

[0002] With the development of internet technology, more and more applications and platforms offer natural language processing (NLP) capabilities, bringing numerous conveniences to users. Applications and platforms with NLP functions can provide NLP services to users based on trained machine learning models. Speech recognition is a crucial task within NLP. The goal is to improve the efficiency of speech recognition while ensuring its accuracy. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for speech recognition is provided. The method includes: adding at least one first node, representing at least one first candidate text sequence, to a prefix tree based on at least one first candidate text sequence identified from a first speech; adding at least one second node, representing at least one second candidate text sequence, to the prefix tree based on at least one second candidate text sequence identified from a second speech, wherein the at least one second node is connected to each first node, and the second speech is acquired immediately following the first speech; if the prefix tree includes multiple first nodes, determining a score corresponding to each of the multiple text sequences corresponding to multiple first paths obtained from the prefix tree, wherein each text sequence corresponding to a first path includes at least a combination of a first node and a candidate text sequence represented by a connected second node; if there exists at least one first path with a score less than a score threshold, deleting at least one first path from the prefix tree to delete at least one first node, obtaining an updated prefix tree; and determining a first target text sequence matching the first speech from at least one first candidate text sequence represented by the at least one first node that was not deleted, based at least on the updated prefix tree.

[0004] In a second aspect of this disclosure, an apparatus for speech recognition is provided. The apparatus includes: a first node adding module configured to add at least one first node representing at least one first candidate text sequence to a prefix tree based on at least one first candidate text sequence identified from a first speech; a second node adding module configured to add at least one second node representing at least one second candidate text sequence to the prefix tree based on at least one second candidate text sequence identified from a second speech, wherein the at least one second node is connected to each first node, and the second speech is acquired immediately following the first speech; a score determining module configured to determine a score corresponding to each of the multiple text sequences if the prefix tree includes multiple first nodes, based at least on the semantics of each of the multiple text sequences corresponding to multiple first paths obtained from the prefix tree, wherein each text sequence corresponding to a first path includes at least a combination of a first node and a candidate text sequence represented by a connected second node; a prefix tree updating module configured to delete at least one first path from the prefix tree to delete at least one first node if there is at least one first path with a score less than a score threshold, thereby obtaining an updated prefix tree; and a text sequence determining module configured to determine a first target text sequence matching the first speech from at least one first candidate text sequence represented by the at least one first node that has not been deleted, based at least on the updated prefix tree.

[0005] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method of the first aspect of this disclosure when executed by the at least one processing unit.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program that can be executed by a processor to perform the method according to a first aspect of this disclosure.

[0007] It should be understood that the description in the Summary of the Invention section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] The above and other features, advantages, and aspects of various implementations of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0010] Figure 2 A flowchart of a process for speech recognition according to some embodiments of the present disclosure is shown;

[0011] Figure 3 A schematic diagram of an example prefix tree according to some embodiments of the present disclosure is shown;

[0012] Figure 4 A schematic diagram illustrating the process of performing a speech recognition task using a machine learning model according to some embodiments of the present disclosure is shown;

[0013] Figure 5 A schematic diagram of key members of a prefix tree according to some embodiments of the present disclosure is shown;

[0014] Figure 6 A schematic diagram illustrating the prefix tree update process according to certain embodiments of the present disclosure is shown;

[0015] Figure 7 A schematic structural block diagram of an apparatus for speech recognition according to certain embodiments of the present disclosure is shown; and

[0016] Figure 8 A block diagram of a computing device in which one or more embodiments of the present disclosure may be implemented is shown. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0019] It should be noted that the acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0020] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0021] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.

[0022] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0023] It is understood that the above notification and user authorization acquisition process is merely illustrative and does not constitute a limitation on the embodiments of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the embodiments of this disclosure.

[0024] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0025] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0026] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating its parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual inputs based on the trained parameter values ​​to determine the corresponding output.

[0027] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, environment 100 may include electronic device 110.

[0028] Electronic device 110 can convert target speech 102 into a target text sequence 112 that matches target speech 102. That is, electronic device 110 can perform speech recognition on target speech 102 to generate the corresponding target text sequence 112. The target speech can be any appropriate language and of any duration. For example, electronic device 110 can perform speech recognition on target speech 102 to generate a text sequence in its corresponding language. For example, electronic device 110 can recognize speech in English to generate a text sequence in English. The target speech 102 can be local speech of electronic device 110 or speech acquired by electronic device 110 in real time.

[0029] Electronic device 110 may, for example, utilize a trained machine learning model 115 to perform a speech recognition task. Machine learning model 115 may be, for example, any suitable model including but not limited to, Transformer models, convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep neural networks (DNNs), etc. Machine learning model 115 may be a model native to electronic device 110 or a model installed on another electronic device 110 (e.g., installed on a remote device).

[0030] Electronic device 110 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, server devices, etc. Terminal devices may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. Server devices may be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server-side equipment may include computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0031] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0032] As briefly mentioned earlier, applications or platforms with natural language processing (NLP) capabilities can provide NLP services to users based on trained machine learning models. Speech recognition is a crucial task within NLP. With the continuous advancement and widespread adoption of artificial intelligence, expectations for speech recognition performance are rising. Beyond demanding the highest possible accuracy in recognizing a speech stream, users also require real-time processing of the audio stream and display of intermediate results. This necessitates streaming real-time speech recognition, rather than acquiring the entire audio file before displaying the results. Real-time recognition enhances the user experience. For example, in voice input and conference subtitle scenarios, users expect to see the speech transcription results (i.e., speech recognition results) simultaneously with their speech. If no display appears on the screen during speaking, users may feel the interaction is clunky and lose patience. Furthermore, in complex human-computer interaction scenarios, speech recognition is often a relatively early component of the system, and providing real-time recognition results helps improve the performance of downstream tasks. For instance, interruption and background noise filtering functions in intelligent customer service may utilize real-time speech recognition results as features.

[0033] Compared to traditional hybrid frameworks based on Hidden Markov Models (HMMs), end-to-end speech recognition achieves a significant improvement in accuracy for non-real-time recognition. For real-time speech recognition, while chunking and attention mechanisms can be introduced to ensure accuracy, the spike latency of Connectionist Temporal Classification (CTC, an algorithm used to address the problem of inconsistent input and output sequence lengths and lack of alignment) and chunking waiting issues in the model still result in a delay in the decoding of a word compared to its time in the original speech. This affects the real-time performance of the displayed results.

[0034] Traditionally, a set of cue-based schemes has been proposed to address the aforementioned latency issues. This cue-based scheme pads the currently acquired audio encoding with a virtual encoding of a chunk size (also called a virtual chunk or padding chunk). Leveraging the predictive capabilities of the end-to-end model encoder, which is similar to a masked language model (LM), the text corresponding to this padding chunk is predicted during decoding. This allows for the advance display of text corresponding to possible future speech during real-time recognition, reducing latency in displaying recognition results. Since the padding chunk itself does not exist, it is usually set to 0, called a zero-padding chunk (or simply zeropadding; in this paper, the two terms are used interchangeably). This scheme requires that it only affects intermediate recognition results, not the final recognition result. Therefore, the intermediate results of decoding the zero-padding chunk should be discarded when subsequent real audio arrives, and subsequent audio decoding should begin from the chunk preceding the zero-padding chunk (i.e., the already acquired chunk suitable for speech recognition, which can be called the current chunk). To meet this requirement, two solutions have traditionally been proposed. One approach is to re-decode from the initial moment of the audio each time it arrives, thus eliminating the need to consider the impact of previous predictions on the decoding of zero-filled blocks. The other approach is to back up the results of the current block that has already been decoded before decoding the zero-filled block. When the actual audio arrives later, decoding begins from the backed-up state, and the decoding results of the zero-filled block after the backup state are discarded.

[0035] The method described above, which re-decodes from the beginning of the audio, ensures the final result remains unaffected. However, it's easy to see that starting from scratch involves numerous redundant operations, resulting in significant computational and memory overhead. Furthermore, this overhead increases with audio length. While the model can pre-decode the text, the substantial computational cost slows down the decoding process, ultimately impacting the display speed of the speech recognition results. The second method backs up the decoding results of the actual speech before decoding the zero-padding blocks. This allows subsequent audio to be decoded from the backup without starting from scratch, avoiding substantial redundant computation. However, backup also incurs memory overhead. Additionally, after decoding the zero-padding blocks, all states corresponding to the decoded paths are no longer needed and must be deleted, which is also a time-consuming operation.

[0036] In view of the above, embodiments of this disclosure provide a method for speech recognition. The method includes: adding at least one first node representing at least one first candidate text sequence to a prefix tree based on at least one first candidate text sequence identified from a first speech; adding at least one second node representing at least one second candidate text sequence to the prefix tree based on at least one second candidate text sequence identified from a second speech, wherein the at least one second node is connected to each first node, and the second speech is acquired immediately following the first speech; if the prefix tree includes multiple first nodes, determining a score corresponding to each of the multiple text sequences corresponding to multiple first paths obtained from the prefix tree, wherein each text sequence corresponding to a first path includes at least a combination of candidate text sequences represented by a first node and a second node connected thereto; if there is at least one first path with a score less than a score threshold, deleting at least one first path from the prefix tree to delete at least one first node, thereby obtaining an updated prefix tree; and determining a first target text sequence matching the first speech from at least one first candidate text sequence represented by the at least one first node that was not deleted, based at least on the updated prefix tree. In this way, the embodiments of this disclosure can utilize a prefix tree structure to perform speech recognition by operating on each node in the prefix tree, which can improve the efficiency of speech recognition while ensuring the accuracy of the speech recognition results.

[0037] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0038] Figure 2 A flowchart of a process 200 for speech recognition according to some embodiments of the present disclosure is shown. Process 200 can be implemented at electronic device 110. For ease of discussion, reference will be made to... Figure 1 The environment 100 is used to describe the process 200.

[0039] In block 210, electronic device 110 adds at least one first node to the prefix tree, each representing at least one first candidate text sequence, based on at least one first candidate text sequence identified from the first speech. In block 220, electronic device 110 adds at least one second node to the prefix tree, each representing at least one second candidate text sequence, based on at least one second candidate text sequence identified from the second speech, with the at least one second node connected to each first node, the second speech being acquired immediately following the first speech.

[0040] The first and second voice recordings here can be voice recordings stored locally on the electronic device 110, or voice recordings collected in real time by the electronic device 110. For ease of description, the first and second voice recordings in the following text refer to the voice recordings collected in real time by the electronic device 110.

[0041] In some embodiments, the duration of the speech recognized by the electronic device 110 each time is fixed, for example, it can be fixed to a predetermined duration. This predetermined duration can be set by the user based on the input requirements of the machine learning model performing the speech recognition task, or it can be determined by the electronic device 110 itself. In this case, the first speech and the second speech can both be speech that meets the predetermined duration. For example, if both the first speech and the second speech are speech acquired in real time, the first speech can be, for example, speech acquired within a predetermined duration before the first moment, and the second speech can be, for example, speech acquired within a predetermined duration after the first moment.

[0042] It is understandable that the electronic device 110 can also collect third, fourth, and so on speech segments that meet a predetermined duration before and / or after the first speech segment. When using a machine learning model to perform a speech recognition task, such speech segments with the same duration (all meeting the predetermined duration) can also be referred to as segments to be provided to the machine learning model. For example, the first speech segment can also be referred to as the first segment, the second speech segment can also be referred to as the second segment, and so on.

[0043] Considering speech accuracy issues (e.g., pronunciation problems caused by a user's non-standard accent), electronic device 110 can perform speech recognition on speech that meets a predetermined duration to generate at least one first candidate text sequence that matches the speech. In some embodiments, electronic device 110 can provide speech that meets the predetermined duration to a trained machine learning model. The machine learning model performs speech recognition on the speech and outputs at least one candidate text sequence corresponding to the speech. Electronic device 110 obtains the model output from the machine learning model, i.e., obtains at least one candidate text sequence. Alternatively and / or additionally, in some embodiments, electronic device 110 can also utilize predetermined rules and / or algorithms to perform speech recognition on the speech to generate at least one text sequence. For example, electronic device 110 can recognize a first speech that meets the predetermined duration to generate at least one first candidate text sequence, and can recognize a second speech that meets the predetermined duration to generate at least one second candidate text sequence.

[0044] Electronic device 110 can determine the nodes to be added to the trie tree based on the identified candidate text sequences. A trie tree (also known as a word lookup tree, key tree, dictionary tree, etc.) is a tree structure, a variant of a hash tree, and a special form of an N-ary tree. In some embodiments, regarding the specific method of adding nodes to the trie tree, for the first node, if the first speech is the initially acquired speech, electronic device 110 can connect at least one first node after the root node of the trie tree (represented as a BoS node). The root node of the trie tree does not contain characters, i.e., it does not contain a text sequence. If the first speech is not the initially acquired speech, electronic device 110 can connect at least one node after each node in the trie tree corresponding to historical speech acquired before the first speech. For example, if the first speech is the speech immediately following the zeroth speech, electronic device 110 can include at least one zeroth node, which represents at least one zeroth candidate text sequence corresponding to the zeroth speech, and at least one first node is connected after each zeroth node. Similarly, for the second node, since the second speech is the speech acquired immediately after the first speech and is not the initially acquired speech, the electronic device 110 can directly connect at least one corresponding second node to each first node.

[0045] In some embodiments, the at least one first node described above may also be collectively referred to as a first node, and each of the at least one first node may also be referred to as an element included in the collectively referred to first node. Similarly, the at least one second node described above may also be collectively referred to as a second node, and each of the at least one second node may be referred to as an element included in the collectively referred to second node. In this case, the above process may also be described as follows: the electronic device 110 adds a first node representing at least one first candidate text sequence to the prefix tree based on at least one first candidate text sequence recognized from the first speech, where each element in the first node is used to represent one of the at least one first candidate text sequences. The electronic device 110 adds a second node representing at least one second candidate text sequence to the prefix tree based on at least one second candidate text sequence recognized from the second speech, where each element in the second node is used to represent one of the at least one second candidate text sequences. The second node is connected to the first node, and each element included in the second node is connected to each element of the first node (that is, all elements of the second node are connected to each element in the first node).

[0046] In some embodiments, since nodes in the prefix tree, excluding the root node, often correspond to only one character, the candidate text sequence (e.g., a first candidate text sequence, a second candidate text sequence) corresponding to the node can be, for example, a text unit. In this case, since the candidate text sequence corresponding to speech that meets the predetermined duration may include multiple candidate text units, for example, the predetermined duration includes N frames of speech, the electronic device 110 can identify and determine at least one candidate text unit corresponding to each frame of speech. For example, if the first speech is speech with a duration of 8 frames, for each of these 8 frames of speech, the electronic device 110 can determine at least one candidate text unit corresponding to that frame of speech. In this case, the predetermined duration of the speech can be determined as the first predetermined duration, and one frame can be determined as the second predetermined duration. After obtaining speech that meets the first predetermined duration, the electronic device 110 performs speech recognition on speech that meets the second predetermined duration (i.e., performs speech recognition on each frame of speech) to determine at least one candidate text unit corresponding to speech that meets the second predetermined duration, and adds at least one node representing the candidate text unit in the prefix tree. For example, after acquiring 8 frames of speech, the electronic device 110 can identify the first frame of speech as the first speech and add at least one first node representing at least one first candidate text unit of the first speech in the prefix tree. The electronic device 110 can identify the second frame of speech as the second speech and add at least one second node representing at least one second candidate text unit of the second frame of speech after each first node in the prefix tree, and so on.

[0047] Because some speech segments satisfying the first predetermined duration may lack corresponding candidate text units (e.g., candidate text units cannot be recognized in the 3rd frame of speech), this can lead to discrepancies in input and output sequence lengths and misalignment when using a machine learning model for speech recognition. To address this issue, electronic device 110 introduces the CTC criterion, which uses the special character "blank" to represent emptiness. The CTC criterion calculates the score of a candidate recognition result during decoding by summing the probabilities of all paths that can be merged into that candidate. For example, if the first predetermined duration is 5 frames, and electronic device 110 determines that the speech recognition result satisfying the first predetermined duration is ABC, then the paths AABC_, A_B_C, and AB__C (where '_' represents the special character "blank," and each path's length equals the speech duration of 5) in the prefix tree are all paths that can be merged into ABC, and their probabilities should be added to the score of ABC. The steps for merging paths into candidate results using the CTC criterion are: first, merge adjacent and identical non-special character "blank" characters into one character; then, remove all special character "blank" symbols.

[0048] Figure 3 A schematic diagram of an example prefix tree 300 according to some embodiments of the present disclosure is shown. Figure 3 As shown, the prefix tree 300 includes a root node 301 and several other nodes. If the first speech is the initially acquired speech, after the electronic device 110 recognizes the text units "I" and "lie down" from the first speech, it can connect the nodes 302-1 and 302-2 corresponding to these two text units after the root node. In this case, nodes 302-1 and 302-2 are the two first nodes corresponding to the first speech.

[0049] If the first speech is not the initially acquired speech, after the electronic device 110 recognizes the text units "love" and "sigh" from the first speech, it can connect the nodes 303-1 and 303-2 corresponding to these two text units to each node corresponding to the historical speech, that is, connect them to nodes 302-1 and 302-2. In this case, the prefix tree includes multiple nodes 303-1 and multiple nodes 303-2, which are the multiple first nodes corresponding to the first speech. In this case, node 302-1 is connected to subsequent nodes 303-1 and 303-2, and node 302-2 is also connected to nodes 303-1 and 303-2 (not shown in the figure).

[0050] Similarly, if the electronic device 110 recognizes the text units "kick" and "body" from the second speech, since these two text units are different from the text units corresponding to the previous nodes, the nodes 304-1 and 304-2 corresponding to these two text units can be connected to the nodes corresponding to each first speech (i.e., nodes 303-1 and 303-2) without merging. Each node 303-1 and each node 303-2 is connected to nodes 304-1 and 304-2. In this case, the prefix tree includes multiple nodes 304-1 and multiple nodes 304-2, which are the multiple second nodes corresponding to the second speech.

[0051] Return to reference Figure 2 In box 230, if the prefix tree includes multiple first nodes, the electronic device 110 determines the scores corresponding to each of the multiple text sequences based at least on the semantics of each of the multiple text sequences corresponding to the multiple first paths obtained from the prefix tree, wherein each text sequence corresponding to the first path includes at least a combination of the candidate text sequences represented by the first node and the second node connected to it.

[0052] If the first speech is the initially acquired speech, each first path includes only a root node, a first node, and a second node. In this case, the text sequence corresponding to each first path is obtained by concatenating a first text sequence and a second text sequence. If the first speech is not the initially acquired speech, each first path includes, in addition to the root node, a first node, and a second node, at least one node between the first node and the root node that corresponds to a historical speech. In this case, the text sequence corresponding to each first path is obtained by concatenating a first text sequence, a second text sequence, and a historical text sequence corresponding to the historical speech.

[0053] In some embodiments, the scores corresponding to each of the multiple text sequences determined by the electronic device 110 are the semantic scores of each of the multiple text sequences. The electronic device 110 can perform semantic analysis on these multiple text sequences to determine the semantic score corresponding to each text sequence. For example, if there are two first paths, and the text sequences corresponding to these two first paths are "I love" and "I sigh", the electronic device 110 can perform semantic analysis on these two text sequences and perform semantic scoring to determine the semantic scores corresponding to these two text sequences respectively. For example, the semantic score of the text sequence "I love" is 55, and the semantic score of the text sequence "I sigh" is 15.

[0054] In some embodiments, the scores corresponding to each of the plurality of text sequences determined by the electronic device 110 further include the matching degree scores between the candidate text sequence corresponding to the node and the corresponding speech. Specifically, the electronic device 110 can determine the plurality of first candidate text sequences corresponding to a plurality of first nodes and the plurality of first matching degree scores with the first speech, and the plurality of second candidate text sequences corresponding to at least one second node and the plurality of second matching degree scores with the second speech. The matching degree score can indicate the probability that the speech is recognized as the candidate text sequence. For example, if the electronic device 110 can recognize two text units "I" and "lie down" from the first speech, the electronic device 110 can determine the probability that the first speech is recognized as each of these two text units. For example, the probability that the first speech is recognized as the text unit "I" is 70%, and the probability that it is recognized as the text unit "lie down" is 30%. It is understood that the first speech can also be recognized as other text units, such as the text unit "nest". In this case, the electronic device 110 can determine that the probability that the first speech is recognized as the text unit "I" is 55%, the probability that it is recognized as the text unit "lie down" is 25%, the probability that it is recognized as the text unit "nest" is 10%, and so on. In some embodiments, if the number of identified candidate text sequences exceeds a threshold, the electronic device 110 may add only the nodes corresponding to the candidate text sequences whose probabilities (i.e., matching scores) are higher than the matching score threshold to the prefix tree. Alternatively and / or additionally, in some embodiments, the electronic device 110 may sort the multiple candidate text sequences corresponding to the speech based on the numerical value of the matching score. For example, the electronic device 110 may add the nodes corresponding to the few candidate text sequences with the highest matching scores to the prefix tree. For example, it may add only the nodes corresponding to the two candidate text sequences with the highest matching scores to the prefix tree.

[0055] In this scenario, for each of the multiple first paths, the electronic device 110 can determine the score of the first path based on the first matching score corresponding to the first node, the second matching score corresponding to the second node, and the semantic score of the first path. The score of each first path can, for example, be the sum of the matching scores and semantic scores of all nodes in the path except the root node. Taking a first path that only includes the root node, the first node, and the second node as an example, the score of this first path is the sum of the first matching score corresponding to the first node, the second matching score corresponding to the second node, and the semantic score of the first path.

[0056] In box 240, if there exists at least one first path with a score less than a score threshold, electronic device 110 deletes at least one first path from the prefix tree to delete at least one first node, obtaining an updated prefix tree. The score threshold here can be pre-configured by the user or determined by electronic device 110 itself.

[0057] If there exists at least one first path with a score less than a score threshold, the electronic device 110 can delete at least one first node by removing this first path from the prefix tree. For example, if the scores of the text sequences corresponding to two first paths are 55 and 15 respectively, and the score threshold is 50, then the electronic device 110 can delete the first path with a score of 15 from the prefix tree.

[0058] In some embodiments, if there is no first path with a score less than a score threshold, that is, when the scores of multiple text sequences corresponding to multiple first paths all reach the score threshold, the electronic device 110 may, for example, retain all of these multiple first paths, that is, not update the prefix tree. For example, if the scores of the text sequences corresponding to two first paths are 55 and 60 respectively, and the score threshold is 50, then the electronic device 110 may retain both of these first paths without deleting either of them.

[0059] In some embodiments, besides determining the score of the text sequence corresponding to the path based on semantics and deleting paths based on a comparison of the score and a score threshold, the electronic device 110 can also delete paths based on a predetermined number. Specifically, for a first path, if the electronic device 110 has pre-obtained a predetermined number of first paths, and the scores of the text sequences corresponding to multiple first paths included in the prefix tree all reach the score threshold, then the electronic device 110 can sort these multiple paths according to their scores and retain only at least one first path from the top of the sorted list, which is a predetermined number. That is, the electronic device 110 will delete at least one first path from the bottom of the sorted list. For example, if the prefix tree includes a total of 6 first paths, and the scores of the text sequences of these 6 first paths all reach the score threshold (for example, the score threshold is 5, and the scores corresponding to these 6 first paths are 5.5, 6, 6, 5.5, 7, and 6.5 respectively), and the number of remaining paths for the first path is 2, then the electronic device 110 can retain only the two first paths with the highest corresponding scores (that is, the two first paths with scores of 7 and 6.5) and delete the remaining first paths.

[0060] It should be noted that when the first speech is not the initially acquired speech, there is at least one node corresponding to the historical speech connected before the multiple first nodes of the prefix tree. When the electronic device 110 deletes at least one path, it only deletes the first node of this at least one first path and the nodes connected after the first node, and does not delete the nodes before the first node.

[0061] In some embodiments, if the electronic device 110 acquires speech within a predetermined time period prior to the current moment and identifies at least one candidate text sequence based on this speech, the electronic device 110 can add a node representing this at least one candidate text sequence to the prefix tree. Such a node added at the current moment can be referred to as the active node at the current moment. (Continue to refer to...) Figure 3 ,like Figure 3 As shown, if the current time is T4, then nodes 305-1 and 305-2 are the active nodes at the current time. Figure 3 The active nodes at the current moment are represented by a circular ring. In some embodiments, the active nodes may also include at least one node added in the previous moment. Specifically, considering the possibility that two consecutive speech segments correspond to the same text sequence, the electronic device 110 may merge two consecutive nodes with the same text sequence based on the CTC criterion. Figure 3 As shown, if at time T4, the text unit determined by electronic device 110 based on speech includes the text unit "kick", since there is a node 304-1 corresponding to the text unit "kick" at time T3, electronic device 110 does not need to add the node corresponding to the text unit "kick" obtained at time T4 after node 304-1. Electronic device 110 can directly add the matching score of the node of the text unit "kick" at time T4 to node 304-1. For example, if the matching score of node 304-1 is 55 at time T3, and the matching score of the node corresponding to the text unit "kick" is 35 at time T4, then after merging the nodes at time T4, electronic device 110 can make the matching score of node 304-1 added at time T3 90. The calculation method for adding the matching score here can also be, for example, multiplication. For instance, if the matching score of node 304-1 is 0.7 at time T3, and the matching score of the node corresponding to the text unit "kick" is 0.5 at time T4, then merging the nodes at time T4 can make the matching score of node 304-1 added at time T3 0.35. This disclosure does not limit the specific method of addition.

[0062] In some embodiments, for any node in the prefix tree other than the root node, the electronic device 110 may determine the node connected to the node before the node as the parent node of the node, and determine the node connected to the node after the node as the child node of the node. Exemplarily, both node 304-1 and node 304-2 are child nodes of node 303-1, and node 302-1 is the parent node of node 303-1.

[0063] For the prefix tree 300, at time T1, the electronic device 110 can recognize the text units "I" and "wo" from the voice obtained in real time. The electronic device 110 adds two child nodes corresponding to these two text units from the root node 301, that is, adds node 302-1 and node 302-2. At time T2, the electronic device 110 can recognize the text units "love" and "ai" from the voice obtained in real time. The electronic device 110 adds the nodes 303-1 and 303-2 corresponding to these two text units after each node added at time T1. That is, adds node 303-1 and node 303-2 after node 302-1 and node 302-2. The electronic device 110 can obtain 4 paths, and these 4 paths correspond to 4 text sequences "I love", "I ai", "wo love", "wo ai".

[0064] The electronic device 110 may, for example, determine the score of the path based on the matching degree score of the node and the semantic score of the path. If the scores of the paths corresponding to the text sequences "I love" and "I ai" are two paths among the 4 paths that reach the score threshold and meet the predetermined number (for example, 2), the electronic device 110 may delete the live nodes corresponding to the remaining two paths, that is, delete the nodes added at time T2 in this path (for example, delete node 303-1 and node 303-2 connected to node 302-2). For the convenience of description, hereinafter, when describing the deletion of the path, in the case of no special example, it is default that the predetermined number for the path is 2, and at least 2 paths have scores higher than the score threshold each time. The electronic device 110 deletes multiple paths of the prefix tree to only retain the two paths with the highest corresponding scores and scores higher than the score threshold. Figure 3 When the electronic device 1 deleting multiple paths of the prefix tree to only retain the two paths with the highest corresponding scores and scores higher than the score threshold.

[0065] Furthermore, since nodes 303-1 and 303-2 connected to node 302-2 are deleted, and node 302-2 is not an active node at the current time and has no child nodes, electronic device 110 also deletes node 302-2 from the prefix tree 300. Similarly, at time T3, if the text sequences "I love kicking" and "I love body" among the multiple paths obtained by electronic device 110 have the highest scores, electronic device 110 can delete the active nodes of other paths besides these two paths at time T3. Electronic device 110 deletes node 303-2 connected to node 302-1, as well as nodes 304-1 and 304-2 connected to node 303-2. At time T4, if the text sequences "I love kicking feet" and "I love kicking shuttlecock" among the multiple paths obtained by electronic device 110 have the highest scores, electronic device 110 can delete the active nodes of other paths besides these two paths at time T4. Electronic device 110 removes the connection to node 304-2 after node 303-1, as well as nodes 305-1 and 305-2 after node 304-2.

[0066] In box 250, electronic device 110 determines a first target text sequence that matches the first speech from at least one first candidate text sequence represented by at least one first node that has never been deleted, based at least on the updated prefix tree.

[0067] In some embodiments, if the updated prefix tree is deleted down to a single remaining first node, the electronic device 110 can determine the first candidate text sequence corresponding to that single first node as the first target text sequence. It should be noted that a single remaining first node does not mean that the number of paths containing that first node is 1. (Continue to refer to...) Figure 3 Regarding nodes 302-1 and 302-2 added to the prefix tree 300 by electronic device 110 at time T1, if these two nodes are designated as first nodes, and the node added at time T2 is designated as the second node, then at time T2, based on the comparison of paths and path scores and a predetermined number, electronic device 110 deletes node 302-2 and its connected nodes 303-1 and 303-2. That is, at time T2, electronic device 110 retains only two paths: node 301-node 302-1-node 303-1 and node 301-node 302-1-node 303-2. It can be observed that at this time, the prefix tree 300 only includes a single first node 302-1, but it includes two first paths.

[0068] In some embodiments, the electronic device 110 may perform multiple rounds of deletion operations on multiple first nodes until the prefix tree contains only a single first node. In some embodiments, if the prefix tree contains two or more first nodes that have not been deleted after deleting at least one first node from the prefix tree by deleting at least one first path, the electronic device 110 may also continue the deletion operation based on a third voice.

[0069] Specifically, the electronic device 110 may, for example, add at least one third node representing at least one third candidate text sequence to the updated prefix tree based on at least one third candidate text sequence recognized from the third speech. The at least one third node is connected to each second node in the path corresponding to the at least one first node that has not been deleted, with the third speech being acquired immediately following the second speech. If the updated prefix tree includes more than two first nodes, the electronic device 110 may determine the scores corresponding to the multiple text sequences corresponding to the multiple second paths based on the semantics of each of the multiple text sequences corresponding to the multiple second paths corresponding to the two or more first nodes. Each text sequence corresponding to a second path includes at least a combination of a first node, a second node connected to the first node, and a third node connected to the second node representing the candidate text sequence. The electronic device 110 may delete at least one second path with a score less than a score threshold from the updated prefix tree to delete at least one first node, obtaining a further updated prefix tree. The score threshold for the second path here may be the same as or different from the score threshold for the first path mentioned earlier. Electronic device 110 can determine a first target text sequence that matches the first speech from at least one first candidate text sequence represented by at least one first node that has never been deleted, based at least once the prefix tree has been updated again.

[0070] It is understandable that if the deletion operation on the first node is performed based on the second path including the third node and there are still at least two first nodes, the electronic device 110 can continue to perform the deletion operation on the first node based on the fourth voice, the fifth voice, etc., until the prefix tree only includes a single first node.

[0071] It is understandable that if the electronic device 110 only recognizes a first candidate text sequence based on the first speech, that is, the electronic device 110 only adds a first node to the prefix tree, the electronic device 110 can directly determine the first candidate text sequence corresponding to the first node as the first target text sequence.

[0072] In some embodiments, if no more speech is acquired within a predetermined time period, the electronic device 110 may, for example, select a first target text sequence from at least one first candidate text sequence based on a first matching score between each first candidate text sequence and the first speech. In this case, since the electronic device 110 no longer acquires speech within the predetermined time period, the electronic device 110 can determine that speech acquisition has ended, and the electronic device 110 then selects the first target text sequence from at least one first candidate text sequence based on the acquired speech. The electronic device 110 may, for example, directly select the first target text sequence from at least one first candidate text sequence based on the first matching score between each first candidate text sequence and the first speech. The electronic device 110 may also, for example, determine the first candidate text sequence corresponding to the first node included in the path with the highest score as the first target text sequence based on the score of each path in the plurality of paths included in the prefix tree, wherein each path includes at least a first node (e.g., it may also include a historical node corresponding to historical speech). The first target text sequence is the recognition result for the first speech.

[0073] Figure 4 A schematic diagram of a process 400 for performing a speech recognition task using a machine learning model according to some embodiments of the present disclosure is shown. The machine learning model includes an encoder 430 and a decoder 460. Both the encoder 430 and the decoder 460 can be encoders and decoders capable of performing CTC algorithms, and both the encoder 430 and the decoder 460 can, for example, include a CTC module (not shown) for performing the CTC algorithm.

[0074] Electronic device 110 can acquire speech stream 410, which can be, for example, a speech stream acquired in real time by electronic device 110. When acquiring speech stream 410, electronic device 110 can divide speech stream 410 into blocks according to a certain length and offset, such as extracting audio features (also called speech features) from segments of length 25ms. Audio features are one-dimensional vectors, thus speech stream 410 becomes a series of one-dimensional vectors arranged over time, which can be represented by a two-dimensional vector (T, D), where T represents time and D represents the feature dimension. To reduce computation and extract deeper audio features, electronic device 110 can use a convolutional neural network to downsample the extracted audio features, that is, to transform feature vectors from multiple time points into a single feature vector. For example, if the downsampling factor is 4, the downsampled audio features become (t, D), where t = T / / 4. The downsampled features are divided into different blocks according to a fixed time length, such as... Figure 4 The blocks shown are 401, 402, 403, etc. Figure 4The size of each block at the indicated location is 8, that is, the features of 8 consecutive moments form a block. The electronic device 110 provides the block to the encoder 430. The attention mechanism and convolutional network in the encoder 430 can encode the audio features and learn their relationship with the modeling unit (here it is the text unit or text sequence). Combining with the CTC module, it outputs the probability of all modeling units at each moment (that is, the matching degree score). That is, it outputs the matching degree scores of multiple text units corresponding to each moment respectively. The electronic device 110 can be represented by a two-dimensional vector (t, vocab_size + 1), where vocab_size refers to the size of the dictionary, and the words in the dictionary are represented by single characters. For example, a dictionary can be ['我', '爱', '踢', '足', '球', '毽', '子'], and at this time the size of the dictionary is 7. Since the CTC criterion also introduces a special character "blank" to represent emptiness, one more bit in the dimension is required to represent the probability of the special character "blank".

[0075] For block-based streaming real-time speech recognition, intermediate results need to be given in real time. That is, for each incoming block, the electronic device 110 needs to give the recognition result from the beginning of the speech to the current block. Assume Figure 4 that after downsampling the collected speech, 16 frames are obtained, that is, the current moment is t15. These frames form block 401 and block 402. Block 402 is the currently to-be-processed block, and block 401 is the cache for calculating block 402 in the attention mechanism. Assume that the content of the original speech corresponding to the moments t0 - t15 is all the pronunciations of "我爱踢" and part of the pronunciation of "足". Due to the CTC spike delay problem, after decoding block 402, the recognition result may be "我爱踢" or even "我爱", which is one or two text units less than the real content of the speech. The missing text units need subsequent speech to arrive and凑齐 a block size (that is, t23 and block 403 in the figure) to be determined. This waiting will directly affect the user experience.

[0076] It should be noted that there is an unclear expression "凑齐" in the original text which is directly translated as "凑齐" in the above translation. It may need to be further clarified according to the specific context for a more accurate translation.In the embodiments of this disclosure, zero-padding blocks can be used. When block 402 is acquired but block 403 is not acquired at the current time, 8 frames with a feature of 0 are directly added after block 402 to obtain block 420 (i.e., zero-padding block). Utilizing the prediction capabilities of the encoder 430 and decoder 460, which are similar to masked language models, block 420 is encoded and decoded to predict the content corresponding to block 403. The encoder 430 can output the encoding results 440 and 450 corresponding to blocks 402 and 420 respectively. The decoder 460 can decode these two encoding results to output the text recognition result corresponding to block 402 and the prediction result corresponding to block 403. Thus, even when only frame t15 is acquired (i.e., block 402 is acquired but block 403 is not acquired), the electronic device 110 may be able to recognize "I love playing football" or even "I love playing soccer". It should be noted that the encoding and decoding of block 420 here does not affect the encoding and decoding of block 403 after it is obtained.

[0077] For the encoding and decoding operations of blocks 420 and 403, the decoding based on zero-padding blocks is more complex. Each time, in addition to decoding the frames in the current block (i.e., block 402), it is also necessary to decode the frames in the zero-padding block (i.e., block 420). The electronic device 110 also needs to ensure that decoding block 420 does not affect the state of the already decoded frames in block 402; otherwise, the initial state of the decoded block 403 will change, leading to inconsistent final recognition results. The impact of zero-padding block decoding on the already decoded state of the current block includes two aspects: first, it may mistakenly modify the score of the active node corresponding to the last frame of the current block; second, it may mistakenly delete the active node of the current block. One possible scenario for the first aspect is that the candidate text unit obtained from decoding the first frame of the zero-padding block is the same as the candidate text unit decoded from the last frame of the current block. In this case, after merging according to the CTC criterion, no new node will be created, and the score originally saved in the current block will be overwritten by the new score. The second possibility mentioned above is that the active node of the current block is no longer an active node in the subsequent zero-padding block decoding process and will be pruned, resulting in accidental deletion.

[0078] To ensure that encoding and decoding of block 420 does not affect encoding and decoding of block 403, embodiments of this disclosure also propose prefix tree node members and decoding methods. Figure 5 A schematic diagram of a key member 500 of a prefix tree according to some embodiments of the present disclosure is shown. To address the problem of the first aspect described above, the electronic device 110 can use different variables (e.g., ...) to represent the scores of the current block decoding stage and the zero-padding block decoding stage. Figure 5The `real_chunk` and `zero_chunk` symbols shown in the diagram are used to avoid accidental modification issues. Since the CTC criterion involves the special character "blank", the electronic device 110 can calculate scores for paths ending with and without "blank" during decoding, and sum them as candidate scores. Because the calculation methods for both scores are the same, for simplicity and universality, they can be uniformly represented by "score". Each chunk is decoded frame-by-frame over time. The score of a path in the current frame can, for example, be equal to the sum of the semantic score of the corresponding path in the previous frame and the matching score of a modeling unit in this frame, for example (score('CBA') = score('CB') + prob('A')). Therefore, "prev" and "cur" are used to store the path scores of the previous and current frames, respectively. Regarding the second aspect of the problem mentioned above, the electronic device 110 can use `real_chunk_exists` and `zero_chunk_exists` to indicate whether the current node exists when decoding the current chunk and the zero-padding chunk, respectively. If a node does not exist, it does not need to be physically deleted from the prefix tree; simply setting these two flags is sufficient. A node only needs to be physically deleted if its flag is false and it has no child nodes. During zero-padding block decoding, it is stipulated that a node can only be physically deleted from the prefix tree if both `zero_chunk_exists` and `real_chunk_exists` are false. Since the `real_chunk_exists` of the currently active block node is true, even if the node needs to be deleted during the zero-padding block decoding stage, only `zero_chunk_exists` will be set to false, so it will not be physically deleted. Furthermore, in a node, "char" represents the characters (i.e., text units and / or text sequences) contained in the current node, "parent" points to the parent node of the current node. When obtaining real-time recognition results, the characters from all nodes traversed back to the root node from the current node are concatenated and reversed to obtain the real-time recognition result. "children" points to all child nodes of the current node, each child node represented by a mapping between its character value and address.

[0079] Figure 6 A schematic diagram of a prefix tree update process 600 according to certain embodiments of the present disclosure is shown. The audio of the current block is subjected to feature extraction and downsampling to obtain T frames. The electronic device 110 adds a zero-padding block of length T frames and value 0 after the current block. After encoding and CTC, the probability 601 of each modeling unit in the 2T-frame modeling unit is obtained, which is the matching score of each modeling unit.

[0080] In box 610, electronic device 110 acquires the current frame modeling unit probability.

[0081] In box 620, electronic device 110 traverses the tree to obtain the active node (S_active) after the previous frame has been decoded.

[0082] In box 630, electronic device 110 takes the N1 modeling units with the highest probability in the current frame to obtain a set (top_ele).

[0083] In box 640, electronic device 110 determines each element (ele) in the set, the active node and its merged tree node, and updates the corresponding score.

[0084] In box 650, electronic device 110 traverses the tree and sets the N2 tree nodes with the highest scores to true, and the others to false, and deletes redundant nodes.

[0085] In frame 660, electronic device 110 determines whether the current frame is the last frame of a zero-padding block.

[0086] In box 670, if the current frame is the last frame of a zero-padding block, real-time decoding is complete, the node with the highest score in the tree is found, and the real-time recognition result is obtained by backtracking. Otherwise, in box 680, the electronic device 110 treats the next frame as the current frame and continues processing the next frame. The above is an abstract summary of the decoding process, because real-time recognition and decoding of zero-padding blocks must consider the previously mentioned issues of erroneous modification of the current block score and erroneous deletion of nodes. Therefore... Figure 6 The above steps also differ in their processing logic during the decoding stages of the current block and the zero-padding block. Specifically (for ease of description, decoding the current block is referred to as Stage 1, and decoding the zero-padding block is referred to as Stage 2):

[0087] First, in Phase 1, nodes with real_chunk_exists set to true are selected as active nodes; in Phase 2, nodes with zero_chunk_exists set to true are selected as active nodes.

[0088] Secondly, for the elements (decoded paths) in the active nodes of Phase 1 and the corresponding nodes in the tree (either existing or newly created) of the merged new path, the values ​​of `real_chunk_score_cur` and `zero_chunk_score_cur` are both set to the sum of the `real_chunk_score_prev` value of the elements in the active nodes and the element's probability. `real_chunk_exists` and `zero_chunk_exists` are both set to true. The values ​​related to zero-padding blocks in the Phase 1 nodes are kept consistent with the values ​​related to the current block so that the values ​​related to zero-padding blocks can be used for decoding in Phase 2. Phase 2 stores the intermediate decoding results in variables related to zero-padding blocks. If the node corresponding to the merged new path already exists in the tree, only `zero_chunk_score_pre` is used to update the value of `zero_chunk_score_cur`, while `zero_chunk_exists` is set to true. If the merged node does not exist in the tree, a new tree node is created, and the method for modifying the values ​​related to zero-padding blocks in the new node is the same as when the node already exists. The scores related to the current chunk are all set to 0, and real_chunk_exists is set to false, indicating that the node was not created during the decoding of the current chunk, which facilitates subsequent deletion processing. It is easy to see that because the scores of stage one and stage two decoding are represented by zero_chunk_score and real_chunk_score respectively, the situation of real_chunk_score being mistakenly modified will not occur.

[0089] Third, after merging and updating the `real_chunk_score_cur` and `zero_chunk_score_cur` of the nodes, a better decoding path after decoding the current frame needs to be determined through any appropriate pruning method (such as beam pruning) as the initial candidate path for decoding the next frame. Pruning here refers to the deletion operation performed on the nodes and paths in the prefix tree mentioned above. In Phase 1, the `real_chunk_exists` and `zero_chunk_exists` of the nodes retained after pruning are set to true, while these two variables are set to false for other nodes. If a node's `real_chunk_exists` is false and it has no child nodes, then that node is deleted from the tree. In Phase Two, the `zero_chunk_exitsts` property of the retained nodes after pruning is set to true, while the `zero_chunk_exitsts` property of other nodes is set to false. `real_chunk_exists` remains unchanged. A node can only be deleted from the tree if both `zero_chunk_exists` and `real_chunk_exists` are false and it has no child nodes. This prevents candidate nodes retained in the current block decoding phase from being mistakenly deleted during the zero-padding block decoding phase (because `real_chunk_exists` is true). Furthermore, to improve decoding efficiency, this invention introduces threshold-based pruning in addition to the histogram-based pruning method. A score threshold `s_threshold` is set. If a node's `score_cur` is less than the difference between the largest `score_cur` of the active nodes in the tree and `s_threshold`, then that node cannot be considered an active node after decoding.

[0090] Fourth, after decoding in stages one and two, the tree node with the highest zero_chunk_score_cur is found. Tracing back to the root node, the character fields of all nodes along the backtracking path are concatenated and reversed to obtain the real-time result of the prompt-based speech recognition. The node with the highest zero_chunk_score_cur is chosen because it records the optimal candidate generated from decoding all frames in stages one and two.

[0091] Therefore, the embodiments of this disclosure solve the problem of erroneous modification and deletion of the currently active node in the tree structure decoding by special design of the score in the tree node and whether the node has members, as well as different update strategies for the member variable values ​​at different decoding stages.

[0092] The foregoing described in detail the encoding and decoding of the speech stream and the updating of the prefix tree by the electronic device 110. The following describes an example scenario of such speech recognition. In some embodiments, such a speech recognition scheme can be applied to a conference scenario. For example, if participants in a conference expect the electronic device 110 to capture their speech in real time and present the corresponding text sequence, the electronic device 110 can employ such a speech recognition scheme to perform the speech recognition task.

[0093] In such a scenario, electronic device 110 can display a text sequence corresponding to the user's speech in real time on the user interface; this text sequence can also be referred to as subtitles, for example. Such a user interface could be, for example, a user interface for a conferencing application.

[0094] In some embodiments, before determining the first target text sequence (i.e., after acquiring the first speech and recognizing at least one first candidate text sequence based on the first speech), the electronic device 110 can select a first text sequence to be presented in the user interface from at least one first candidate text sequences based on the first matching score of each of the at least one first candidate text sequences with the first speech. For example, the electronic device 110 can determine the first candidate text sequence with the highest first matching score among the at least one first candidate text sequences as the first text sequence to be presented in the user interface. In some embodiments, in addition to the matching score, the electronic device 110 can also perform semantic analysis on the at least one first candidate text sequence to determine the semantic score corresponding to each of the at least one first candidate text sequences. The electronic device 110 can then accumulate the semantic score and the first matching score corresponding to each of the at least one first candidate text sequences, and determine the first candidate text sequence with the highest accumulated score (i.e., the first matching score of the first candidate text sequence + the semantic score of the first candidate text sequence) as the first text sequence. It is understood that subsequent determinations, such as the determination of the second text sequence, can also incorporate the semantic scores corresponding to the candidate text sequences.

[0095] In some embodiments, if no second speech of a predetermined duration is captured after the first speech, the electronic device 110 generates a predicted text sequence corresponding to the second speech based at least on the first text sequence. Specifically, if the first speech is the initially captured speech, the electronic device 110 can generate the predicted text sequence based on the first text sequence. The specific prediction method here can be the method described above, which predicts the block corresponding to the zero-padding block based on the current block and the zero-padding block. If the first speech is not the initially captured speech, the electronic device 110 can generate the predicted text sequence based on the first text sequence and historical text sequences. The historical text sequence here can be, for example, the historical text sequence corresponding to the path with the highest score in the prefix tree. The electronic device 110 then presents the first text sequence and the predicted text sequence on the user interface.

[0096] In some embodiments, after acquiring the second speech and recognizing at least one second candidate text sequence, the electronic device 110 can further select a second text sequence to be presented in the user interface based on a second matching score between each of the at least one second candidate text sequence and the second speech. For example, the electronic device 110 can determine the second candidate text sequence with the highest second matching score among the at least one second candidate text sequences as the second text sequence. If the predicted text sequence differs from the second text sequence, the electronic device 110 can switch the predicted text sequence to the second text sequence in the user interface.

[0097] Furthermore, if the determined first target text sequence is different from the first text sequence, the electronic device 110 can also switch the first text sequence presented in the user interface to the first target text sequence after determining the first target text sequence. For example, if the electronic device 110 previously presented the first text sequence and the second text sequence, in response to the first target text sequence being different from the first text sequence, the electronic device 110 can switch to presenting the first target text sequence and the second text sequence.

[0098] Therefore, it is possible to display subtitles corresponding to speech in real time, and it can automatically switch incorrect text sequences in the subtitles to correct text sequences, which helps to improve the user experience.

[0099] In summary, the embodiments of this disclosure can utilize a prefix tree structure to perform speech recognition by operating on each node in the prefix tree, which can improve the efficiency of speech recognition while ensuring the accuracy of the speech recognition results.

[0100] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 7A schematic structural block diagram of a speech recognition device 700 according to certain embodiments of the present disclosure is shown. The device 700 may be implemented as or included in an electronic device 110. Various modules / components in the device 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0101] like Figure 7 As shown, the device 700 includes a first node adding module 710, configured to add at least one first node representing at least one first candidate text sequence to a prefix tree based on at least one first candidate text sequence identified from a first speech. The device 700 also includes a second node adding module 720, configured to add at least one second node representing at least one second candidate text sequence to the prefix tree based on at least one second candidate text sequence identified from a second speech, wherein the at least one second node is connected to each first node, and the second speech is acquired immediately following the first speech. The device 700 also includes a score determining module 730, configured to determine a score corresponding to each of the multiple text sequences if the prefix tree includes multiple first nodes, based at least on the semantics of each of the multiple text sequences corresponding to multiple first paths obtained from the prefix tree, wherein each text sequence corresponding to a first path includes at least a combination of a first node and a candidate text sequence represented by a connected second node. The device 700 also includes a prefix tree updating module 740, configured to delete at least one first path from the prefix tree to delete at least one first node if there exists at least one first path with a score less than a score threshold, thereby obtaining an updated prefix tree. The apparatus 700 also includes a text sequence determination module 750, configured to determine a first target text sequence that matches the first speech from at least a first candidate text sequence represented by at least a first node that has never been deleted, based at least based on the updated prefix tree.

[0102] In some embodiments, the apparatus 700 further includes a first text sequence determination module, configured to determine a first candidate text sequence corresponding to a single first node as a first target text sequence that matches the first speech if the prefix tree includes a single first node.

[0103] In some embodiments, the score determination module 730 includes: a matching degree score determination module, configured to determine multiple first matching degree scores between multiple first candidate text sequences corresponding to multiple first nodes and the first speech, and at least one second matching degree score between at least one second candidate text sequence corresponding to at least one second node and the second speech; a semantic score determination module, configured to determine the semantic score of each of the multiple text sequences based on the semantics of each of the multiple text sequences corresponding to multiple first paths; and a determination module, configured to determine the score of each of the multiple first paths based on the first matching degree score corresponding to the first node included in the first path, the second matching degree score corresponding to the second node, and the semantic score of the first path.

[0104] In some embodiments, the text sequence determination module 750 includes: a third node adding module, configured to add at least one third node representing at least one third candidate text sequence to an updated prefix tree based on at least one third candidate text sequence identified from a third speech, wherein the at least one third node is connected to each second node in a path corresponding to at least one first node that has not been deleted, and the third speech is acquired immediately following the second speech; a second score determination module, configured to determine a score corresponding to each of the multiple text sequences corresponding to multiple second paths based on the semantics of the multiple text sequences corresponding to the multiple second paths corresponding to the multiple first nodes if the updated prefix tree includes two or more first nodes, wherein each text sequence corresponding to a second path includes at least a combination of a first node, a second node connected to the first node, and a third node connected to the second node representing a candidate text sequence; a node deletion module, configured to delete at least one first node from the updated prefix tree if there is at least one second path with a score less than a score threshold, thereby obtaining a re-updated prefix tree; and a target text sequence determination module, configured to determine a first target text sequence matching the first speech from at least one first candidate text sequence represented by at least one first node that has not been deleted, based at least on the re-updated prefix tree.

[0105] In some embodiments, the text sequence determination module 750 includes: a first determination module configured to determine the first candidate text sequence corresponding to the single first node as the first target text sequence if the updated prefix tree is deleted down to the remaining single first node; and a second determination module configured to select the first target text sequence from at least one first candidate text sequence based on the first matching score between each first candidate text sequence and the first speech if no more speech is acquired in the following predetermined time period.

[0106] In some embodiments, both the first voice and the second voice are voices that meet a predetermined duration.

[0107] In some embodiments, both the first speech and the second speech are speech collected in real time. The first speech is speech collected within a predetermined time period before the first moment, and the second speech is speech collected within a predetermined time period after the first moment.

[0108] In some embodiments, the first node adding module 710 includes: a first connection module configured to connect at least one first node after the root node of the prefix tree if the first speech is the initially acquired speech; and a second connection module configured to connect at least one node after each node corresponding to historical speech acquired before the first speech in the prefix tree if the first speech is not the initially acquired speech.

[0109] In some embodiments, the apparatus 700 further includes a first selection module configured to select a first text sequence to be presented in a user interface based on a first matching score between each of the at least one first candidate text sequences and the first speech, before determining a first target text sequence.

[0110] In some embodiments, the apparatus 700 further includes: a predicted text sequence generation module configured to generate a predicted text sequence corresponding to the second speech if no second speech of a predetermined duration is collected after the first speech; and a first presentation module configured to present the first text sequence and the predicted text sequence on a user interface.

[0111] In some embodiments, the apparatus 700 further includes: a second selection module configured to, after acquiring the second speech and recognizing at least one second candidate text sequence, select a second text sequence to be presented in the user interface based on a second matching score between each of the at least one second candidate text sequence and the second speech; and a first switching module configured to switch the predicted text sequence to the second text sequence in the user interface if the predicted text sequence is different from the second text sequence.

[0112] In some embodiments, the apparatus 700 further includes a second switching module configured to switch the first text sequence presented in the user interface to the first target text sequence after determining the first target text sequence if the determined first target text sequence is different from the first text sequence.

[0113] In some embodiments, the user interface includes the user interface of a conferencing application.

[0114] The units and / or modules included in device 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 700 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0115] It should be understood that one or more steps in the above methods can be performed by suitable electronic devices or combinations of electronic devices. Such electronic devices or combinations of electronic devices may include, for example, […]. Figure 1 Electronic device 110.

[0116] Figure 8 A block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 8 The electronic device 800 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 8 The electronic device 800 shown can be used to achieve Figure 1 Electronic devices 110 or Figure 7 Device 700.

[0117] like Figure 8 As shown, electronic device 800 is in the form of a general-purpose electronic device. Components of electronic device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 800.

[0118] Electronic device 800 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data and accessible within electronic device 800.

[0119] Electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 8 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 820 may include computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0120] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 800 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0121] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 800 can also communicate with one or more external devices (not shown) via communication unit 840 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 800, or with any device that enables electronic device 800 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0122] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0123] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0124] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0125] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0127] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for speech recognition, comprising: Based on at least one first candidate text sequence identified from the first speech, at least one first node representing the at least one first candidate text sequence is added to the prefix tree; Based on at least one second candidate text sequence identified from the second speech, at least one second node representing the at least one second candidate text sequence is added to the prefix tree, the at least one second node being connected after each first node, and the second speech being acquired immediately after the first speech; If the prefix tree includes multiple first nodes, the semantic score of each of the multiple text sequences corresponding to the multiple first paths obtained from the prefix tree is determined at least based on the semantics of each of the multiple text sequences. Each text sequence corresponding to the first path includes at least a combination of the candidate text sequences represented by the first node and the second node connected to it. Based on the semantic scores of each of the multiple text sequences and the matching scores between the corresponding speech and the corresponding candidate text sequences of each node in the multiple first paths, the scores corresponding to each of the multiple text sequences are determined. If there exists at least one first path with a score less than a score threshold, the at least one first path is deleted from the prefix tree to delete at least one first node, resulting in an updated prefix tree. The nodes of the prefix tree include different score variables for the current block decoding stage and the zero-padding block decoding stage, respectively, and each has a first flag indicating whether the node exists in the current block decoding stage and a second flag indicating whether the node exists in the zero-padding block decoding stage. In the zero-padding block decoding stage, when the first flag and the second flag indicate that the node does not exist, the node is deleted from the prefix tree. as well as Based at least on the updated prefix tree, a first target text sequence matching the first speech is determined from at least one first candidate text sequence represented by at least one first node that has never been deleted.

2. The method according to claim 1, further comprising: If the prefix tree includes a single first node, the first candidate text sequence corresponding to the single first node is determined as the first target text sequence that matches the first speech.

3. The method according to claim 1, wherein determining the score corresponding to each of the plurality of text sequences includes: Determine the first matching degree scores of the multiple first candidate text sequences corresponding to the multiple first nodes and the first speech, and the at least one second candidate text sequence corresponding to the at least one second node and the at least one second matching degree score of the second speech; as well as For each of the plurality of first paths, the score of the first path is determined based on the first matching score corresponding to the first node included in the first path, the second matching score corresponding to the second node, and the semantic score of the first path.

4. The method of claim 1, wherein determining a first target text sequence matching the first speech from at least one first candidate text sequence represented by at least one first node that has never been deleted comprises: Based on at least one third candidate text sequence identified from the third speech, at least one third node representing the at least one third candidate text sequence is added to the updated prefix tree. The at least one third node is connected to each second node in the path corresponding to the at least one first node that has not been deleted. The third speech is collected immediately after the second speech. If the updated prefix tree includes two or more first nodes, the scores corresponding to the multiple text sequences corresponding to the multiple second paths are determined based on the semantics of each of the multiple text sequences corresponding to the multiple second paths corresponding to the two or more first nodes. Each text sequence corresponding to a second path includes at least a combination of a first node, a second node connected to the first node, and a third node connected to the second node representing a candidate text sequence. If there exists at least one second path whose score is less than the score threshold, delete the at least one second path from the updated prefix tree to delete at least one first node, and obtain a further updated prefix tree; as well as Based at least on the updated prefix tree, a first target text sequence matching the first speech is determined from at least one first candidate text sequence represented by at least one first node that has never been deleted.

5. The method of claim 1, wherein determining a first target text sequence matching the first speech from at least one first candidate text sequence represented by at least one first node that has never been deleted comprises: If the updated prefix tree is deleted down to the remaining single first node, the first candidate text sequence corresponding to the single first node is determined as the first target text sequence. as well as If no more speech is acquired within the next predetermined time period, the first target text sequence is selected from the at least one first candidate text sequence based on the first matching score between each first candidate text sequence and the first speech.

6. The method according to claim 1, wherein both the first voice and the second voice are voices that satisfy a predetermined duration.

7. The method according to claim 6, wherein the first speech and the second speech are both real-time acquired speech, the first speech is speech acquired within the predetermined duration before the first moment, and the second speech is speech acquired within the predetermined duration after the first moment.

8. The method of claim 1, wherein adding at least one first node representing the at least one first candidate text sequence to the prefix tree comprises: If the first speech is the initially acquired speech, connect the at least one first node after the root node of the prefix tree; as well as If the first speech is not the initially acquired speech, the at least one node is connected to each node in the prefix tree corresponding to the historical speech acquired before the first speech.

9. The method according to claim 1, further comprising: Before the first target text sequence is determined, a first text sequence to be presented in the user interface is selected from the at least one first candidate text sequences based on the first matching degree score of each of the at least one first candidate text sequences with the first speech.

10. The method of claim 9, further comprising: If no second speech of predetermined duration is captured after the first speech, a predicted text sequence corresponding to the second speech is generated at least based on the first text sequence; and The user interface presents the first text sequence and the predicted text sequence.

11. The method of claim 10, further comprising: After acquiring the second speech and recognizing the at least one second candidate text sequence, based on the second matching score between each of the at least one second candidate text sequence and the second speech, a second text sequence to be presented in the user interface is selected from the at least one second candidate text sequence; and If the predicted text sequence is different from the second text sequence, the user interface switches the predicted text sequence to the second text sequence.

12. The method according to claim 9, further comprising: If the determined first target text sequence is different from the first text sequence, the first text sequence presented in the user interface is switched to the first target text sequence after the first target text sequence is determined.

13. The method of claim 9, wherein the user interface comprises a user interface of a conferencing application.

14. An apparatus for speech recognition, comprising: The first node adding module is configured to add at least one first node to the prefix tree, representing the at least one first candidate text sequence, based on at least one first candidate text sequence identified from the first speech. The second node adding module is configured to add at least one second node representing the at least one second candidate text sequence to the prefix tree based on at least one second candidate text sequence identified from the second speech, wherein the at least one second node is connected to each first node and the second speech is acquired immediately after the first speech. The scoring module is configured to, if the prefix tree includes multiple first nodes, determine the semantic score of each of the multiple text sequences corresponding to the multiple first paths obtained from the prefix tree, wherein each text sequence corresponding to the first path includes at least a combination of a first node and a candidate text sequence represented by a connected second node; and determine the score corresponding to each of the multiple text sequences based on the semantic score of each of the multiple text sequences and the matching score between the corresponding speech of each node in the multiple first paths and the corresponding candidate text sequence. A prefix tree update module is configured to delete at least one first path from the prefix tree to remove at least one first node if there exists at least one first path with a score less than a score threshold, thereby obtaining an updated prefix tree. The prefix tree update module is further configured to: nodes in the prefix tree include different score variables for the current block decoding stage and the zero-padding block decoding stage, and each node has a first flag indicating whether the node exists in the current block decoding stage and a second flag indicating whether the node exists in the zero-padding block decoding stage; in the zero-padding block decoding stage, when the first flag and the second flag indicate that the node does not exist, the node is deleted from the prefix tree. as well as The text sequence determination module is configured to determine a first target text sequence that matches the first speech from at least one first candidate text sequence represented by at least one first node that has never been deleted, based at least on the updated prefix tree.

15. An electronic device comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 13 when executed by the at least one processing unit.

16. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Word graph rescoring method and system for deep learning language models

    CN108415898A

  • System for using silence in speech recognition

    CN1307715A