Method and apparatus for speech recognition, device, and readable medium
By adopting the prefix tree structure and dynamic node operation method in the speech recognition technology, the problem of low speech recognition efficiency in the prior art is solved, streaming real-time speech recognition is realized, and the accuracy and efficiency of recognition results are improved.
Patent Information
- Application Number
- PCT/CN2024/133647
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-30
Smart Images

Figure CN2024133647_30052025_PF_FP_ABST
Abstract
Description
Method, apparatus, device and readable medium for speech recognition
[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, apparatus and readable media for speech recognition” and application number 202311560513.3, filed on November 21, 2023. The entire contents of that application are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to methods, apparatuses, devices, and computer-readable storage media for speech recognition. Background Art
[0003] With the development of internet technology, more and more applications and platforms are offering natural language processing (NLP) capabilities, bringing significant benefits to users. These applications and platforms can provide NLP services based on trained machine learning models. Speech recognition is a key task within NLP. It is hoped that this capability can be improved in efficiency while ensuring the accuracy of speech recognition results. Summary of the Invention
[0004] In a first aspect of the present disclosure, a method for speech recognition is provided. The method includes: adding at least one first node representing at least one first candidate text sequence to a prefix tree based on at least one first candidate text sequence recognized from a first speech; adding at least one second node representing at least one second candidate text sequence to the prefix tree based on at least one second candidate text sequence recognized from a second speech, wherein at least one second node is connected to each first node, and the second speech is collected immediately after the first speech; if the prefix tree includes multiple first nodes, determining scores corresponding to each of the multiple text sequences based on at least one semantic relationship of multiple text sequences corresponding to multiple first paths obtained from the prefix tree, wherein the text sequence corresponding to each first path includes at least a combination of candidate text sequences represented by a first node and a second node connected thereto; if there is at least one first path with a score less than a score threshold, deleting at least one first path from the prefix tree to delete at least one first node, thereby obtaining an updated prefix tree; and determining, based at least on the updated prefix tree, a first target text sequence matching the first speech from at least one first candidate text sequence represented by the at least one first node that has not been deleted.
[0005] In other aspects of the present disclosure, a device for speech recognition, an electronic device, and a computer-readable storage medium are also provided. The device, electronic device, and computer-readable storage medium can implement the method according to the first aspect of the present disclosure.
[0006] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0008] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0009] FIG2 shows a flow chart of a process for speech recognition according to some embodiments of the present disclosure;
[0010] FIG3 shows a schematic diagram of an example prefix tree according to some embodiments of the present disclosure;
[0011] FIG4 is a schematic diagram illustrating a process of performing a speech recognition task using a machine learning model according to some embodiments of the present disclosure;
[0012] FIG5 shows a schematic diagram of key members of a prefix tree according to some embodiments of the present disclosure;
[0013] FIG6 is a schematic diagram showing a process of updating a prefix tree according to certain embodiments of the present disclosure;
[0014] FIG7 shows a schematic structural block diagram of an apparatus for speech recognition according to certain embodiments of the present disclosure; and
[0015] FIG8 illustrates a block diagram of a computing device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0016] It should be noted that the acquisition, storage and application of user personal information involved in the technical solution of this disclosure are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0017] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0018] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG1 , the environment 100 may include an electronic device 110 .
[0019] The electronic device 110 can convert the target speech 102 into a target text sequence 112 that matches the target speech 102. That is, the electronic device 110 can perform a speech recognition task on the target speech 102 to generate a corresponding target text sequence 112. The target speech here can be speech of any appropriate language and any duration. For example, the electronic device 110 can perform speech recognition on the target speech 102 to generate a text sequence in its corresponding language. For example, the electronic device 110 can recognize speech in English to generate a text sequence in English. The target speech 102 here can be a local speech of the electronic device 110, or it can be a speech collected by the electronic device 110 in real time.
[0020] For example, the electronic device 110 can use a trained machine learning model 115 to perform speech recognition tasks. The machine learning model 115 can include, but is not limited to, any appropriate model such as a Transformer model, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), etc. The machine learning model 115 can be a model local to the electronic device 110 or a model installed on another electronic device 110 (e.g., installed on a remote device).
[0021] The electronic device 110 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, server devices, etc. The terminal device may be any type of mobile terminal, fixed terminal, or portable terminal. The server device may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc.
[0022] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0023] As briefly mentioned above, applications or platforms with natural language processing capabilities can provide users with natural language processing services based on trained machine learning models. Speech recognition is a key task within natural language processing. With the continuous advancement and widespread adoption of artificial intelligence technology, expectations for speech recognition performance are rising. In addition to demanding the highest possible accuracy in recognizing a speech segment, people also demand real-time processing of audio streams and display of intermediate results. This is known as streaming real-time speech recognition, rather than waiting until the entire audio is captured before displaying the results. Real-time recognition can enhance the user experience. For example, in voice input and conference captioning scenarios, people expect to see the speech transcription results (i.e., speech recognition results) simultaneously as they speak. If nothing is displayed on the screen during speech, users may experience a clunky interaction and lose patience. On the other hand, in complex human-computer interaction scenarios, speech recognition is often a relatively front-end component of the system, and providing real-time recognition results helps improve the performance of downstream tasks. For example, interruption and background noise filtering features in intelligent customer service may use real-time speech recognition results as features.
[0024] Compared to traditional hybrid frameworks based on Hidden Markov Models (HMMs), end-to-end speech recognition has achieved significant improvements in non-real-time recognition accuracy. For real-time speech recognition, while chunking and attention mechanisms can be used to maintain accuracy, the spike latency and chunking latency of the Connectionist Temporal Classification (CTC) algorithm (an algorithm used to address the problem of misaligned input and output sequence lengths) in the model still result in a delay between the time a word is inferred and decoded and its presence in the original speech, affecting the real-time display of the results.
[0025] Traditionally, a series of hint-based solutions have been proposed to address the aforementioned latency issue. This hint-based solution adds a chunk-sized dummy code (also called a dummy chunk or padding chunk) to the currently acquired audio code. Leveraging the predictive capabilities of the encoder in the end-to-end model, similar to those of a masked language model (LM), the text corresponding to this padding chunk is predicted during the decoding process. This allows for the display of text corresponding to possible future speech during real-time recognition, reducing the latency of displaying recognition results. Because the padding chunk doesn't exist, its value is typically set to 0, resulting in a zero-padding chunk (also referred to interchangeably in this article). This solution only affects intermediate recognition results, not the final one. Therefore, the intermediate results of decoding the zero-padding chunk should be completely discarded when the real audio arrives. Subsequent audio decoding should begin from the chunk immediately preceding the zero-padding chunk (i.e., the chunk already acquired and capable of speech recognition, referred to as the current chunk). To meet this requirement, two solutions have traditionally been proposed. One approach is to re-decode from the initial audio moment each time audio arrives, eliminating the need to consider the impact of previous predictions on the decoding of zero-padded blocks. The other approach is to back up the decoded results of the current block before decoding the zero-padded block. When real audio arrives, decoding begins from the backed-up state, and the decoding results of the zero-padded blocks after the backup state are discarded.
[0026] While the method described above of re-decoding from the beginning of the audio can ensure that the final result is not affected, it is easy to see that starting over again will result in many redundant operations, resulting in high computational and memory overhead. Furthermore, this overhead increases with the length of the audio. Although the model allows the text to be decoded in advance, the high computational overhead will slow down the decoding process, ultimately affecting the speed at which the speech recognition results are displayed. The second method backs up the decoding results of the real speech before decoding the zero-padded blocks. This allows subsequent audio to be decoded from the backup without having to start over, thus avoiding a large amount of redundant computation. However, backups also require memory overhead. Furthermore, after decoding the zero-padded blocks, all states corresponding to the decoding path are no longer needed and need to be deleted, which is also a time-consuming operation.
[0027] In view of this, an embodiment of the present disclosure provides a method for speech recognition. The method includes: based on at least one first candidate text sequence recognized from a first speech, adding at least one first node representing the at least one first candidate text sequence to a prefix tree; based on at least one second candidate text sequence recognized from a second speech, adding at least one second node representing the at least one second candidate text sequence to the prefix tree, at least one second node being connected to each first node, and the second speech is collected immediately after the first speech; if the prefix tree includes multiple first nodes, at least based on the semantics of multiple text sequences corresponding to multiple first paths obtained from the prefix tree, determining the scores corresponding to each of the multiple text sequences, the text sequence corresponding to each first path at least including a combination of candidate text sequences represented by the first node and the second node connected thereto; if there is at least one first path with a score less than a score threshold, deleting at least one first path from the prefix tree to delete at least one first node, thereby obtaining an updated prefix tree; and based at least on the updated prefix tree, determining a first target text sequence matching the first speech from at least one first candidate text sequence represented by the at least one first node that has not been deleted. In this way, the embodiments of the present disclosure can utilize the prefix tree structure and complete speech recognition by operating each node in the prefix tree, thereby improving the efficiency of speech recognition while ensuring the accuracy of the speech recognition result.
[0028] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0029] 2 shows a flow chart of a process 200 for speech recognition according to some embodiments of the present disclosure. The process 200 may be implemented at the electronic device 110. For ease of discussion, the process 200 will be described with reference to the environment 100 of FIG. 1 .
[0030] In block 210, the electronic device 110 adds at least one first node representing at least one first candidate text sequence to the prefix tree based on at least one first candidate text sequence recognized from the first speech. In block 220, the electronic device 110 adds at least one second node representing at least one second candidate text sequence to the prefix tree based on at least one second candidate text sequence recognized from the second speech, wherein the at least one second node is connected to each first node, and the second speech is collected immediately after the first speech.
[0031] The first voice and the second voice here can be voices stored locally in the electronic device 110, or can be voices collected in real time by the electronic device 110. For the sake of convenience, the first voice and the second voice below refer to the voices collected in real time by the electronic device 110.
[0032] In some embodiments, the duration of the speech recognized by the electronic device 110 each time is fixed, for example, it can be fixed to a speech of a predetermined duration. The predetermined duration here can be set by the user based on the input requirements of the machine learning model that performs the speech recognition task, or it can be determined by the electronic device 110 itself. In this case, the first speech and the second speech can, for example, both be speech that meets the predetermined duration. Exemplarily, if the first speech and the second speech are both speech collected in real time, the first speech can, for example, be speech collected within a predetermined duration before the first moment, and the second speech can, for example, be speech collected within a predetermined duration after the first moment.
[0033] It is understandable that the electronic device 110 may further collect a third voice, a fourth voice, and so on that meet a predetermined duration before the first voice and / or after the second voice. When a machine learning model is used to perform a voice recognition task, such voices of the same duration (all meeting the predetermined duration) may also be referred to as blocks to be provided to the machine learning model. For example, the first voice may also be referred to as a first block, and the second voice may also be referred to as a second block, and so on.
[0034] Taking into account the pitch problem of the speech (for example, the pitch problem caused by the non-standard accent of the user who makes the speech), the electronic device 110 can perform speech recognition on the speech that meets the predetermined duration to generate at least one first candidate text sequence that meets the speech. In some embodiments, the electronic device 110 can provide the speech that meets the predetermined duration to a trained machine learning model. The machine learning model performs speech recognition on the speech and outputs at least one candidate text sequence corresponding to the speech. The electronic device 110 obtains the model output from the machine learning model, that is, obtains at least one candidate text sequence. Alternatively and / or additionally, in some embodiments, the electronic device 110 can also use predetermined rules and / or algorithms to perform speech recognition on the speech to generate at least one text sequence. Exemplarily, the electronic device 110 can recognize the first speech that meets the predetermined duration to generate at least one first candidate text sequence, and can recognize the second speech that meets the predetermined duration to generate at least one second candidate text sequence.
[0035] The electronic device 110 can determine the nodes to be added to the prefix tree based on the identified candidate text sequence. The prefix tree (Trie tree, also known as a word search tree, key tree, dictionary tree, etc.) is a tree structure, a variant of a hash tree, and a special form of an N-ary tree. In some embodiments, regarding the specific method of adding nodes to the prefix tree, for the first node, if the first voice is the initially collected voice, the electronic device 110 can connect at least one first node to the root node of the prefix tree (represented as a BoS node). The root node of the prefix tree does not contain characters, that is, it does not contain a text sequence. If the first voice is not the initially collected voice, the electronic device 110 can connect at least one node to each node corresponding to the historical voice collected before the first voice in the prefix tree. Exemplarily, if the first voice is the voice collected immediately after the zeroth voice, the electronic device 110 may include at least one zeroth node, at least one zeroth node is used to represent at least one zeroth candidate text sequence corresponding to the zeroth voice, and at least one first node is connected to each zeroth node. Similarly, for the second node, since the second voice is collected immediately after the first voice and is not the initially collected voice, the electronic device 110 can directly connect the corresponding at least one second node to each first node.
[0036] In some embodiments, the at least one first node mentioned above may also be collectively referred to as a first node, and each first node in the at least one first node may also be referred to as an element included in the collectively referred to first node. Similarly, the at least one second node mentioned above may also be collectively referred to as a second node, and each second node in the at least one second node may be referred to as an element included in the collectively referred to second node. In this case, the above process may also be described as the electronic device 110 adding a first node representing at least one first candidate text sequence to a prefix tree based on at least one first candidate text sequence recognized from the first speech, and each element in the first node is used to represent a first candidate text sequence in the at least one first candidate text sequence. The electronic device 110 adds a second node representing at least one second candidate text sequence to a prefix tree based on at least one second candidate text sequence recognized from the second speech, and each element in the second node is used to represent a second candidate text sequence in the at least one second candidate text sequence, the second node is connected to the first node, and each element included in the second node is connected to each element of the first node (that is, all elements of the second node are connected to each element in the first node).
[0037] In some embodiments, since nodes other than the root node in the prefix tree often correspond to only one character, the candidate text sequences corresponding to the nodes (e.g., the first candidate text sequence, the second candidate text sequence) can be, for example, text units. In this case, since the candidate text sequence corresponding to the speech that meets the predetermined duration may include multiple candidate text units, for example, the predetermined duration includes N frames of speech, the electronic device 110 can recognize each frame of speech and determine at least one candidate text unit corresponding to each frame of speech. For example, if the first speech is a speech with a duration of 8 frames, for each frame of speech in these 8 frames, the electronic device 110 can determine at least one candidate text unit corresponding to the frame of speech. In this case, the predetermined duration of the speech can be determined as the first predetermined duration, and one frame can be determined as the second predetermined duration. After the electronic device 110 obtains the speech that meets the first predetermined duration, it performs speech recognition on the speech that meets the second predetermined duration (i.e., performs speech recognition on each frame of speech) to determine at least one candidate text unit corresponding to the speech that meets the second predetermined duration, and adds at least one node representing the candidate text unit to the prefix tree. For example, after electronic device 110 acquires a speech with a duration of 8 frames, it can determine the first frame of speech as the first speech and add at least one first node representing at least one first candidate text unit of the first speech in the prefix tree. Electronic device 110 can determine the second frame of speech as the second speech and add at least one second node representing at least one second candidate text unit of the second frame of speech after each first node in the prefix tree, and so on.
[0038] Since there may be speech that does not have a corresponding candidate text unit in the speech that meets the first predetermined duration (for example, the third frame of speech cannot be recognized as a candidate text unit), when using a machine learning model to perform a speech recognition task, this situation may cause the input and output sequences to have different lengths and be unable to align. To solve this problem, the electronic device 110 introduced the CTC criterion and introduced the special character "blank" to represent blank. The method used by the CTC criterion to calculate the score of a candidate recognition result during decoding is to find the sum of the probabilities of all paths that can be merged into the candidate. If the first predetermined duration is 5 frames, and the electronic device 110 determines that the speech recognition result of the speech that meets the first predetermined duration is ABC, then the paths AABC_, A_B_C, AB__C ('_' represents the special character "blank", and the length of each path is equal to the speech duration 5) in the prefix tree are all paths that can be merged into ABC, and their probabilities should be added to the score of ABC. The steps for merging paths into candidate results according to the CTC criterion are to first merge adjacent and identical non-special characters "blank" into one character, and then remove all special character "blank" symbols.
[0039] FIG. 3 shows a schematic diagram of an example prefix tree 300 according to some embodiments of the present disclosure. As shown in FIG. 3, the prefix tree 300 includes a root node 301 and a plurality of other nodes. If the first voice is the initially collected voice, after the electronic device 110 recognizes the text units "I" and "lie" from the first voice, the nodes 302-1 and 302-2 corresponding to these two text units can be connected after the root node. In this case, the nodes 302-1 and 302-2 are the two first nodes corresponding to the first voice.
[0040] If the first voice is not the initially collected voice, for example, after the electronic device 110 recognizes the text units "love" and "alas" from the first voice, the nodes 303-1 and 303-2 corresponding to these two text units can be connected after each node corresponding to the historical voice, that is, after the nodes 302-1 and 302-2. In this case, the prefix tree includes a plurality of nodes 303-1 and a plurality of nodes 303-2, and the plurality of nodes 303-1 and the plurality of nodes 303-2 are the plurality of first nodes corresponding to the first voice. In this case, the node 302-1 is connected to the subsequent nodes 303-1 and 303-2, and the node 302-2 is also connected to the nodes 303-1 and 303-2 (not shown in the figure).
[0041] Similarly, if the electronic device 110 recognizes the text units "kick" and "body" from the second voice, since these two text units are different from the text units corresponding to the previous nodes, the nodes 304-1 and 304-2 corresponding to these two text units can be connected after each node corresponding to the first voice (that is, the nodes 303-1 and 303-2), and no merging is required. Each of the nodes 303-1 and each of the nodes 303-2 is connected to the nodes 304-1 and 304-2. In this case, the prefix tree includes a plurality of nodes 304-1 and a plurality of nodes 304-2, and the plurality of nodes 304-1 and the plurality of nodes 304-2 are the plurality of second nodes corresponding to the second voice.
[0042] Returning to FIG. 2, in block 230, if the prefix tree includes a plurality of first nodes, the electronic device 110 determines the scores corresponding to each of the plurality of text sequences at least based on the semantics of each of the plurality of text sequences corresponding to a plurality of first paths obtained from the prefix tree. Each text sequence corresponding to a first path includes at least a combination of candidate text sequences represented by the first node and the second node connected thereto.
[0043] If the first speech is the initially collected speech, each first path only includes a root node, a first node, and a second node. In this case, the text sequence corresponding to each first path is obtained by connecting a first text sequence and a second text sequence. If the first speech is not the initially collected speech, each first path, in addition to including a root node, a first node, and a second node, also includes at least one node corresponding to the historical speech between the first node and the root node. In this case, the text sequence corresponding to each first path is obtained by connecting a first text sequence, a second text sequence, and a historical text sequence corresponding to the historical speech.
[0044] In some embodiments, the scores corresponding to the multiple text sequences determined by the electronic device 110 are the semantic scores of the multiple text sequences. The electronic device 110 can perform semantic analysis on the multiple text sequences to determine the semantic score corresponding to each text sequence. For example, if two first paths are included, and the text sequences corresponding to the two first paths are "I love" and "I sigh", the electronic device 110 can perform semantic analysis on the two text sequences and perform semantic scoring to determine the semantic scores corresponding to the two text sequences. For example, the semantic score of the text sequence "I love" is 55, and the semantic score of the text sequence "I sigh" is 15.
[0045] In some embodiments, the scores corresponding to the multiple text sequences determined by the electronic device 110 further include the matching degree scores between the candidate text sequences corresponding to the nodes and the corresponding speech. Specifically, the electronic device 110 can determine the multiple first matching degree scores between the multiple first candidate text sequences corresponding to the multiple first nodes and the first speech, and the at least one second matching degree scores between the at least one second candidate text sequences corresponding to the at least one second node and the second speech. The matching degree score can indicate the probability that the speech is recognized as the candidate text sequence. Exemplarily, if the electronic device 110 can recognize two text units, "我 (I)" and "卧 (lie)", from the first speech, the electronic device 110 can determine the probabilities that the first speech is recognized as each of these two text units. For example, the probability that the first speech is recognized as the text unit "我 (I)" is 70%, and the probability that it is recognized as the text unit "卧 (lie)" is 30%. It can be understood that the first speech can also be recognized as other text units, such as the text unit "窝 (nest)". In this case, the electronic device 110 can determine that the probability that the first speech is recognized as the text unit "我 (I)" is 55%, the probability that it is recognized as the text unit "卧 (lie)" is 25%, and the probability that it is recognized as the text unit "窝 (nest)" is 10%, and so on. In some embodiments, when the number of recognized candidate text sequences is greater than a threshold number, the electronic device 110 can also add only the nodes corresponding to the candidate text sequences with corresponding probabilities (i.e., matching degree scores) higher than the matching degree threshold to the prefix tree. Alternatively and / or additionally, in some embodiments, the electronic device 110 can also sort the multiple candidate text sequences corresponding to the speech based on the numerical magnitudes of the matching degree scores. For example, the electronic device 110 can add the nodes corresponding to the several candidate text sequences with the largest matching degree scores to the prefix tree. For example, only the nodes corresponding to the two candidate text sequences with the largest matching degree scores can be added to the prefix tree.
[0046] In this case, for each first path among the multiple first paths, the electronic device 110 can determine the score of the first path based on the first matching degree score corresponding to the first node included in the first path, the second matching degree score corresponding to the second node, and the semantic score of the first path. The score of each first path can be, for example, the sum of the matching degree scores and the semantic score of all nodes except the root node in the path. Taking an example that a certain first path only includes the root node, the first node, and the second node, the score of this first path is the sum of the first matching degree score corresponding to the first node, the second matching degree score corresponding to the second node, and the semantic score of the first path.
[0047] At block 240 , if there is at least one first path with a score less than a score threshold, electronic device 110 deletes the at least one first path from the prefix tree to delete at least one first node, thereby obtaining an updated prefix tree. The score threshold may be preconfigured by the user or determined by electronic device 110.
[0048] If there is at least one first path with a score less than the score threshold, electronic device 110 may delete at least one first node by deleting the at least one first path from the prefix tree. For example, if the scores of the text sequences corresponding to the two first paths are 55 and 15, respectively, and the score threshold is 50, electronic device 110 may delete the first path with a score of 15 from the prefix tree.
[0049] In some embodiments, if there is no at least one first path with a score less than a score threshold, that is, if the scores of the multiple text sequences corresponding to the multiple first paths all meet the score threshold, the electronic device 110 may, for example, retain all of the multiple first paths, that is, not update the prefix tree. For example, if the scores of the text sequences corresponding to two first paths are 55 and 60, respectively, and the score threshold is 50, the electronic device 110 may retain both first paths and not delete any of the first paths.
[0050] In some embodiments, in addition to determining the scores of text sequences corresponding to paths based on semantics and deleting paths based on a comparison of the scores with a score threshold, electronic device 110 may also delete paths based on a predetermined number. Specifically, for a first path, if electronic device 110 has pre-obtained a predetermined number of first paths, and the scores of the text sequences corresponding to multiple first paths included in the prefix tree all meet the score threshold, electronic device 110 may sort these multiple paths by score and retain only the predetermined number of at least one first path with the highest score. In other words, electronic device 110 will delete at least one first path with the lowest score. For example, if the prefix tree includes six first paths, and the scores of the text sequences of these six first paths all meet the score threshold (for example, if the score threshold is 5, the scores corresponding to these six first paths are 5.5, 6, 6, 5.5, 7, and 6.5, respectively), and the number of remaining first paths is 2, electronic device 110 may retain only the two first paths with the highest scores (i.e., the two first paths with scores of 7 and 6.5) among the multiple first paths and delete the remaining first paths.
[0051] It should be noted that when the first speech is not the initially collected speech, there is at least one node corresponding to the historical speech connected before the multiple first nodes of the prefix tree. When the electronic device 110 deletes at least one path, it only deletes the first node of the at least one first path and the nodes connected after the first node, but does not delete the nodes before the first node.
[0052] In some embodiments, if the electronic device 110 obtains speech within a predetermined time period before the current moment and recognizes at least one candidate text sequence based on this speech, the electronic device 110 can add a node representing the at least one candidate text sequence to the prefix tree. Such a node added at the current moment can be referred to as the active node at the current moment. Continuing with reference to Figure 3, as shown in Figure 3, if the current moment is time T4, nodes 305-1 and 305-2 are active nodes at the current moment. Figure 3 uses a circle to represent the active nodes at the current moment. In some embodiments, the active node can also include at least one node added at the previous moment before the current moment. Specifically, considering the situation where the text sequences corresponding to two consecutive speech sounds are the same, the electronic device 110 can merge two consecutive nodes with the same corresponding text sequences based on the CTC criterion. As shown in Figure 3, if at time T4, the text unit determined by the electronic device 110 based on the speech includes the text unit "kick", since the node 304-1 corresponding to the text unit "kick" exists at time T3, the electronic device 110 can not add the node corresponding to the text unit "kick" obtained at time T4 after node 304-1. The electronic device 110 can directly add the matching score of the node of the text unit "kick" at time T4 to the node 304-1. For example, if the matching score of the node 304-1 is 55 at time T3 and the matching score of the node corresponding to the text unit "kick" at time T4 is 35, then at time T4, the electronic device 110 can merge the nodes so that the matching score of the node 304-1 added at time T3 is 90. The calculation method for adding the matching scores here can also be, for example, multiplication. For example, if the matching score of the node 304-1 is 0.7 at time T3 and the matching score of the node corresponding to the text unit "kick" at time T4 is 0.5, then at time T4, the matching score of the node 304-1 added at time T3 can be 0.35 after the nodes are merged. The present disclosure does not limit the specific adding method.
[0053] In some embodiments, for any node other than the root node in the prefix tree, the electronic device 110 may determine the node connected to the node before the node as the parent node of the node, and determine the node connected to the node after the node as the child node of the node. For example, node 304-1 and node 304-2 are both child nodes of node 303-1, and node 302-1 is the parent node of node 303-1.
[0054] For prefix tree 300, at time T1, electronic device 110 can identify text units "I" and "lying" from the voice acquired in real time. Electronic device 110 adds two child nodes corresponding to these two text units from root node 301, that is, adds node 302-1 and node 302-2. At time T2, electronic device 110 can identify text units "love" and "sigh" from the voice acquired in real time. Electronic device 110 adds nodes 303-1 and 303-2 corresponding to these two text units after each node added at time T1. That is, node 303-1 and node 303-2 are added after node 302-1 and node 302-2. Electronic device 110 can obtain 4 paths, and these 4 paths correspond to 4 text sequences "I love", "I sigh", "lying love", and "sigh".
[0055] The electronic device 110 can, for example, determine the score of the path based on the matching score of the node and the semantic score of the path. If the scores of the paths corresponding to the text sequences "I love" and "I sigh" are two paths among the four paths whose scores reach the score threshold and meet the predetermined number (for example, 2), the electronic device 110 can delete the live nodes corresponding to the remaining two paths, that is, delete the nodes added at time T2 in the path (for example, delete node 303-1 and node 303-2 connected to node 302-2). For the convenience of description, when describing the deletion of the path below in conjunction with Figure 3, in the absence of special examples, the default predetermined number for the path is 2, and each time there are at least two paths with scores higher than the score threshold. The electronic device 110 deletes multiple paths in the prefix tree to retain only the two paths with the highest corresponding scores and scores higher than the score threshold.
[0056] Furthermore, since nodes 303-1 and 303-2 connected to node 302-2 are deleted, node 302-2 is not an active node at the current moment and has no child nodes, the electronic device 110 also deletes node 302-2 from the prefix tree 300. Similarly, at time T3, if the text sequences "I love kicking" and "I love body" among the multiple paths obtained by the electronic device 110 have the highest scores, the electronic device 110 can delete the active nodes of the other paths other than these two paths at time T3. The electronic device 110 deletes node 303-2 connected after node 302-1 and nodes 304-1 and 304-2 connected after node 303-2. At time T4, if the text sequences "I love kicking" and "I love kicking shuttlecock" among the multiple paths obtained by the electronic device 110 have the highest scores, the electronic device 110 can delete the active nodes of the other paths other than these two paths at time T4. The electronic device 110 deletes the node 304 - 2 connected after the node 303 - 1 and the nodes 305 - 1 and 305 - 2 connected after the node 304 - 2 .
[0057] In block 250 , the electronic device 110 determines, based at least on the updated prefix tree, a first target text sequence that matches the first speech from at least one first candidate text sequence represented by at least one first node that has not been deleted.
[0058] In some embodiments, if the updated prefix tree is deleted to a single first node remaining, the electronic device 110 can determine the first candidate text sequence corresponding to the single first node as the first target text sequence. It should be noted that the remaining single first node does not mean that the number of paths containing the first node is 1. Continuing to refer to Figure 3, for the nodes 302-1 and 302-2 added to the prefix tree 300 by the electronic device 110 at time T1, if these two nodes are determined to be the first nodes and the node added at time T2 is determined to be the second node, then at time T2, the electronic device 110 deletes node 302-2 and the nodes 303-1 and 303-2 connected thereto based on the comparison of the path and the path score and the predetermined number. That is, the electronic device 110 only retains the two paths of node 301-node 302-1-node 303-1 and node 301-node 302-1-node 303-2 at time T2. It can be found that at this time, the prefix tree 300 only includes a single first node 302-1, but includes two first paths.
[0059] In some embodiments, the electronic device 110 may, for example, perform multiple rounds of deletion operations on multiple first nodes until the prefix tree includes only a single first node. In some embodiments, if after deleting at least one first node from the prefix tree by deleting at least one first path, the prefix tree includes two or more undeleted first nodes, the electronic device 110 may further continue the deletion operation based on the third voice.
[0060] Specifically, the electronic device 110 may, for example, add at least one third node representing at least one third candidate text sequence to the updated prefix tree based on at least one third candidate text sequence recognized from the third speech. The at least one third node is connected to each second node in the path corresponding to at least one first node that has not been deleted, and the third speech is collected immediately after the second speech. If the updated prefix tree includes more than two first nodes, the electronic device 110 may determine the scores corresponding to the multiple text sequences corresponding to the multiple second paths based on the semantics of the multiple text sequences corresponding to the multiple second paths corresponding to the two or more first nodes. The text sequence corresponding to each second path includes at least a combination of candidate text sequences represented by a first node, a second node connected to the first node, and a third node connected to the second node. The electronic device 110 may delete at least one second path with a score less than a score threshold from the updated prefix tree to delete the at least one first node, thereby obtaining a further updated prefix tree. The score threshold for the second path here may be the same as or different from the score threshold for the first path described above. The electronic device 110 may determine, based at least on the again updated prefix tree, a first target text sequence that matches the first speech from at least one first candidate text sequence represented by at least one first node that has not been deleted.
[0061] It can be understood that if at least two first nodes are still included after the first node is deleted based on the second path including the third node, the electronic device 110 can continue to perform the deletion operation on the first node based on the fourth voice, the fifth voice, etc. until the prefix tree only includes a single first node.
[0062] It can be understood that if the electronic device 110 only recognizes a first candidate text sequence based on the first speech, that is, the electronic device 110 only adds a first node in the prefix tree, the electronic device 110 can directly determine the first candidate text sequence corresponding to the first node as the first target text sequence.
[0063] In some embodiments, if no voice is acquired within the next predetermined time period, the electronic device 110 may, for example, select a first target text sequence from at least one first candidate text sequence based on the first matching score between each first candidate text sequence and the first voice in at least one first candidate text sequence. In this case, since the electronic device 110 no longer acquires voice within the next predetermined time period, the electronic device 110 may determine that voice acquisition is over, and the electronic device 110 may then select a first target text sequence from at least one first candidate text sequence based on the voice that has been acquired. The electronic device 110 may, for example, directly select a first target text sequence from at least one first candidate text sequence based on the first matching score between each first candidate text sequence and the first voice. The electronic device 110 may also, for example, determine the first candidate text sequence corresponding to the first node included in the path with the highest score as the first target text sequence based on the score of each path in the multiple paths included in the prefix tree, wherein each path includes at least a first node (for example, it may also include a historical node corresponding to a historical voice). The first target text sequence is the recognition result for the first voice.
[0064] FIG4 is a schematic diagram of a process 400 for performing a speech recognition task using a machine learning model according to some embodiments of the present disclosure. The machine learning model includes an encoder 430 and a decoder 460. The encoder 430 and the decoder 460 may both be encoders and decoders that can execute a CTC algorithm. For example, the encoder 430 and the decoder 460 may both include a CTC module (not shown) for executing the CTC algorithm.
[0065] Electronic device 110 can obtain a voice stream 410. This voice stream 410 can be, for example, a voice stream collected in real time by electronic device 110. When collecting voice stream 410, electronic device 110 can segment voice stream 410 into chunks according to a certain length and offset, for example, extracting audio features (also referred to as voice features) from 25ms segments. An audio feature is a one-dimensional vector, so voice stream 410 becomes a series of one-dimensional vectors arranged over time, which can be represented by a two-dimensional vector (T, D), where T represents time and D represents the feature dimension. To reduce computational complexity and extract deeper audio features, electronic device 110 can use a convolutional neural network to downsample the extracted audio features, converting feature vectors from multiple time points into a single feature vector. For example, if the downsampling factor is 4, the downsampled audio features become (t, D), where t = T / / 4. The downsampled features are divided into different chunks of fixed time lengths, such as chunks 401, 402, and 403 shown in FIG4 . The size of each block shown in Figure 4 is 8, that is, the features of 8 consecutive moments constitute a block. The electronic device 110 provides the blocks to the encoder 430. The attention mechanism and convolutional network in the encoder 430 can encode the audio features and learn the relationship between them and the modeling units (here, text units or text sequences). Combined with the CTC module, it outputs the probability of all modeling units at each moment (that is, the matching score). That is, it outputs the matching scores of the multiple text units corresponding to each moment. The electronic device 110 can be represented by a two-dimensional vector (t, vocab_size+1), where vocab_size refers to the size of the dictionary. The words in the dictionary are represented by single characters. For example, a dictionary can be ['I', 'love', 'kick', 'foot', 'ball', 'shuttlecock', 'son'], and the dictionary size is 7. Since the CTC criterion also introduces the special character "blank" to represent emptiness, there must be a bit in the dimension to represent the probability of the special character "blank".
[0066] Block-based streaming real-time speech recognition requires real-time intermediate results. For each block, electronic device 110 must provide recognition results from the beginning of the speech to the current block. Assume that the collected speech in Figure 4 is downsampled to 16 frames, i.e., the current time is t15. These frames constitute blocks 401 and 402. Block 402 is the block currently to be processed, while block 401 is the buffer used by the attention mechanism to calculate block 402. Assume that the original speech content corresponding to times t0-t15 is the full pronunciation of "I love kicking" and the partial pronunciation of "foot." Due to CTC spike latency, after decoding block 402, the recognition result may be "I love kicking" or even "I love," missing one or two text units compared to the actual speech content. These missing text units require subsequent speech to arrive and complete a block size (i.e., t23 and block 403 in the figure) before they can be determined. This waiting time directly affects the user experience.
[0067] In an embodiment of the present disclosure, zero-padded blocks can be combined. When block 402 is obtained at the current moment and block 403 is not obtained, 8 frames with features of 0 are directly added after block 402 to obtain block 420 (i.e., zero-padded blocks). Block 420 is encoded and decoded using the prediction capabilities of the masked language model-like encoder 430 and decoder 460 to predict the content corresponding to block 403. Encoder 430 can output encoding results 440 and encoding results 450 corresponding to blocks 402 and 420 in an associated manner. Decoder 460 can decode these two encoding results to output a text recognition result corresponding to block 402 and a prediction result corresponding to block 403. In this way, when only the t15 frame is obtained (i.e., block 402 is obtained and block 403 is not obtained), electronic device 110 may recognize "I love playing football" or even "I love playing soccer." It should be noted that the encoding and decoding of block 420 here does not affect the encoding and decoding of block 403 after the subsequent acquisition of block 403.
[0068] For the encoding and decoding operations for blocks 420 and 403, since the decoding based on zero-filled blocks is more complicated, each time, in addition to decoding the frames in the current block (i.e., block 402), the frames in the zero-filled blocks (i.e., block 420) must also be decoded. The electronic device 110 also needs to ensure that the decoding of block 420 does not affect the state of the decoded frames in block 402, otherwise the initial state of the decoding of block 403 will change, resulting in inconsistent final recognition results. The impact of zero-filled block decoding on the decoded state of the current block includes two aspects: first, the active node score corresponding to the last frame of the current block is mistakenly modified; second, the active node of the current block is mistakenly deleted. One possible situation of the first aspect mentioned above is that the candidate text unit obtained by decoding the first frame of the zero-filled block is the same as the candidate text unit decoded by the last frame of the current block. In this case, after merging according to the CTC criterion, no new node will be created, and the score originally saved in the current block will be overwritten by the new score. A possible scenario of the second aspect mentioned above is that the active node of the current block is no longer an active node during the subsequent zero-filled block decoding process and will be pruned, resulting in accidental deletion.
[0069] In order to ensure that the encoding and decoding of block 420 does not affect the encoding and decoding of block 403, the embodiments of the present disclosure also propose prefix tree node members and decoding methods. Figure 5 shows a schematic diagram of a key member 500 of a prefix tree according to some embodiments of the present disclosure. In response to the above-mentioned problem of the first aspect, the electronic device 110 can represent the scores of the current block decoding stage and the zero-filled block decoding stage with different variables (such as real_chunk and zero_chunk shown in Figure 5) to avoid the problem of mismodification. Since the CTC criterion involves the special character "blank", the electronic device 110 can calculate the two scores of the path ending with the special character "blank" and not ending with the special character "blank" during decoding, and add them up as the candidate score. Because the two scores are calculated in the same way, they can be uniformly represented by "score" for simplicity and versatility. Each block is decoded frame by frame in time. The score of a path in the current frame can be equal to the sum of the semantic score of the corresponding path in the previous frame and the matching score of a modeling unit in the frame, for example (score('CBA') = score('CB') + prob('A')). Therefore, "prev" and "cur" are used to save the path scores of the previous frame and the current frame respectively. In response to the second aspect of the above-mentioned problem, the electronic device 110 can use real_chunk_exists and zero_chunk_exists to respectively indicate whether the current node exists when decoding the current block and the zero-filled block. If the node does not exist, it does not need to be physically deleted directly from the prefix tree. It is only necessary to set these two flags. Only when the node flag is false and there is no child node, it needs to be physically deleted. When decoding the zero-filled block, it is stipulated that it can only be physically deleted from the prefix tree when both zero_chunk_exists and real_chunk_exists are false. Since the real_chunk_exists of the active node in the current block is true, even if the node needs to be deleted during the zero-fill block decoding phase, only zero_chunk_exists will be set to false, so it will not be physically deleted. In addition, the "char" in the node represents the characters contained in the current node (that is, text units and / or text sequences), and "parent" is used to point to the parent node of the current node. When obtaining real-time recognition results, the characters in all the nodes passed through are concatenated and reversed from the current node to the root node. The real-time recognition result is obtained. "Children" is used to point to all child nodes of the current node. Each child node is represented by a mapping between its character value and address.
[0070] FIG6 illustrates a schematic diagram of a prefix tree update process 600 according to certain embodiments of the present disclosure. After feature extraction and downsampling of the audio of the current block, T frames are obtained. The electronic device 110 adds a zero-padded block of length T frames and value 0 after the current block. After encoding and CTC of the two, the probability 601 of each modeling unit in the 2T-frame modeling unit is obtained, i.e., the matching score of each modeling unit.
[0071] In box 610, the electronic device 110 obtains the modeling unit probabilities for the current frame. In box 620, the electronic device 110 traverses the tree to obtain the active nodes (S_active) after decoding the previous frame. In box 630, the electronic device 110 selects the N1 modeling units with the highest probabilities in the current frame to form a set (top_ele). In box 640, for each element (ele) in the set, the electronic device 110 determines the active nodes and the tree nodes formed by merging them, and updates the corresponding scores. In box 650, the electronic device 110 traverses the tree and sets the N2 tree nodes with the highest scores to true, the others to false, and deletes redundant nodes. In box 660, the electronic device 110 determines whether the current frame is the last frame of the zero-padded block. In box 670, if the current frame is the last frame of the zero-padded block, real-time decoding is completed, the node with the highest score in the tree is found, and backtracking is performed to obtain the real-time recognition result. Otherwise, in box 680, the electronic device 110 uses the next frame as the current frame to continue processing the next frame. The above is an abstract summary of the decoding process. Because real-time identification and decoding of zero-padded blocks must consider the previously mentioned issues of incorrect modification of the current block score and incorrect node deletion, the above steps in Figure 6 also have different processing logic for the current block and zero-padded block decoding stages. The details are as follows (for ease of description, decoding the current block is referred to as stage one, and decoding the zero-padded block is referred to as stage two):
[0072] First, in phase one, nodes whose real_chunk_exists is true are selected as active nodes; in phase two, nodes whose zero_chunk_exists is true are selected as active nodes.
[0073] Secondly, for the nodes in the tree corresponding to the elements (decoded paths) in the active node in phase 1 and the new path obtained by merging the elements (either existing or newly created), their real_chunk_score_cur and zero_chunk_score_cur values are set to the sum of the real_chunk_scroe_prev value of the element in the active node and the element probability, and real_chunk_exists and zero_chunk_exists are both set to true. The values associated with the zero-padded chunk in the phase 1 node and the current chunk remain consistent so that the values associated with the zero-padded chunk can be used for decoding in phase 2. Phase 2 stores the intermediate decoding results in variables associated with the zero-padded chunk. If the node corresponding to the new path obtained by merging already exists in the tree, only the value of zero_chunk_score_cur is updated using zero_chunk_score_pre, and zero_chunk_exists is set to true. If the merged node does not exist in the tree, a new tree node is created, and the values associated with the zero-padded chunk in the new node are modified in the same way as if the node already exists. The scores related to the current chunk are all set to 0, and real_chunk_exists is set to false, indicating that the node was not created when the current chunk was decoded, making it easier to delete it later. It is easy to see that because the scores of phase 1 and phase 2 decoding are represented by zero_chunk_score and real_chunk_score respectively, the real_chunk_score will not be accidentally modified.
[0074] Third, after merging and updating the real_chunk_score_cur and zero_chunk_score_cur of the node, it is necessary to determine the better decoding path after decoding the current frame through any appropriate pruning method (such as beam pruning) as the initial candidate path for decoding the next frame. The pruning here refers to the above-mentioned deletion operation of nodes and paths in the prefix tree. In stage 1, the real_chunk_exists and zero_chunk_exists of the nodes retained after pruning are set to true, and the two variables of other nodes are set to false. If the real_chunk_exits of a node is false and there are no child nodes, the node is deleted from the tree. In the second stage, the zero_chunk_exitsts of the retained nodes after pruning are set to true, the zero_chunk_exitsts of other nodes are set to false, and real_chunk_exists remain unchanged. Only when the zero_chunk_exists and real_chunk_exists of the node are both false and there are no child nodes can it be deleted from the tree. In this way, the candidate nodes retained in the current block decoding stage will not be mistakenly deleted during the zero-filled block stage decoding (because real_chunk_exists is true). In addition, to improve decoding efficiency, based on the histogram pruning method, the present invention introduces threshold-based pruning, sets a score threshold s_threshold, and if the score_cur of the node is less than the difference between the score_cur and s_threshold of the largest active node in the tree, the node cannot be used as the active node after decoding.
[0075] Fourth, after decoding in stages 1 and 2, find the tree node with the highest zero_chunk_score_cur, backtrack to the root node, and concatenate the character fields of all nodes in the backtracking path and reverse the order. This is the real-time result of prompt-based speech recognition. The node with the highest zero_chunk_score_cur is chosen because it records the optimal candidate formed by decoding all frames in stages 1 and 2.
[0076] Therefore, the embodiments of the present disclosure solve the problem of erroneous modification and deletion of active nodes of the current block in tree structure-based decoding through special designs of the scores in tree nodes and whether there are members in the nodes, as well as different update strategies for member variable values in different decoding stages.
[0077] The above describes the details of the electronic device 110 encoding and decoding the voice stream and updating the prefix tree. The following describes an example scenario of such voice recognition. In some embodiments, such a voice recognition scheme can be applied to a conference scenario. For example, if the participants in the meeting expect that the electronic device 110 can collect the voices of the meeting participants in real time and present the corresponding text sequences, the electronic device 110 can use such a voice recognition scheme to perform the voice recognition task. In such a scenario, the electronic device 110 can present a text sequence corresponding to the user's voice in real time on the user interface. Such a text sequence can also be called subtitles, for example. Such a user interface can be, for example, a user interface for a conference application.
[0078] In some embodiments, before determining the first target text sequence (i.e., after acquiring the first speech and recognizing at least one first candidate text sequence based on the first speech), the electronic device 110 may select a first text sequence to be presented in the user interface from the at least one first candidate text sequence based on the first match score of each of the at least one first candidate text sequence with the first speech. For example, the electronic device 110 may determine the first candidate text sequence with the highest first match score among the at least one first candidate text sequence as the first text sequence to be presented in the user interface. In some embodiments, in addition to the match score, the electronic device 110 may also perform semantic analysis on the at least one first candidate text sequence to determine the semantic score corresponding to each of the at least one first candidate text sequence. The electronic device 110 may then accumulate the semantic score and the first match score corresponding to each of the at least one first candidate text sequence and determine the first candidate text sequence with the highest score (i.e., the first match score corresponding to the first candidate text sequence + the semantic score corresponding to the first candidate text sequence) as the first text sequence. It is understood that the subsequent determination of the second text sequence may also be combined with the semantic score corresponding to the candidate text sequence.
[0079] In some embodiments, when a second voice with a predetermined duration is not collected after the first voice, the electronic device 110 generates a predicted text sequence corresponding to the second voice based on at least the first text sequence. Specifically, when the first voice is the initially collected voice, the electronic device 110 can generate a predicted text sequence based on the first text sequence. The specific prediction method here can be the method of predicting the block corresponding to the zero-filled block based on the current block and the zero-filled block as introduced above. When the first voice is not the initially collected voice, the electronic device 110 can generate a predicted text sequence based on the first text sequence and the historical text sequence. The historical text sequence here can be, for example, the historical text sequence corresponding to the path with the highest score in the prefix tree. The electronic device 110 then presents the first text sequence and the predicted text sequence on the user interface.
[0080] In some embodiments, after collecting the second voice and recognizing at least one second candidate text sequence, the electronic device 110 may also select a second text sequence to be presented in the user interface from at least one second candidate text sequence based on the second matching score of each of the at least one second candidate text sequence and the second voice. Exemplarily, the electronic device 110 may determine a second candidate text sequence with the highest second matching score among the at least one second candidate text sequence as the second text sequence. If the predicted text sequence is different from the second text sequence, the electronic device 110 may switch the predicted text sequence to the second text sequence in the user interface.
[0081] Furthermore, if the determined first target text sequence is different from the first text sequence, the electronic device 110 may switch the first text sequence presented in the user interface to the first target text sequence after determining the first target text sequence. For example, if the electronic device 110 previously presented the first text sequence and the second text sequence, in response to the first target text sequence being different from the first text sequence, the electronic device 110 may switch to presenting the first target text sequence and the second text sequence.
[0082] In this way, the subtitles corresponding to the speech can be presented in real time, and the wrong text sequence presented in the subtitles can be switched to the correct text sequence automatically, which helps to improve the user experience.
[0083] In summary, the embodiments of the present disclosure can utilize a prefix tree structure to perform speech recognition by operating on each node in the prefix tree, thereby improving the efficiency of speech recognition while ensuring the accuracy of the speech recognition result.
[0084] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG7 shows a schematic structural block diagram of an apparatus 700 for speech recognition according to certain embodiments of the present disclosure. Apparatus 700 may be implemented as or included in electronic device 110. Each module / component in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.
[0085] As shown in Figure 7, the device 700 includes a first node adding module 710, which is configured to add at least one first node representing at least one first candidate text sequence to the prefix tree based on at least one first candidate text sequence recognized from the first speech; a second node adding module 720, which is configured to add at least one second node representing at least one second candidate text sequence to the prefix tree based on at least one second candidate text sequence recognized from the second speech, at least one second node is connected to each first node, and the second speech is collected immediately after the first speech; a score determining module 730, which is configured to add a first node to the prefix tree based on at least one first node obtained from the prefix tree if the prefix tree includes multiple first nodes. The semantics of the multiple text sequences corresponding to the multiple first paths are determined, and the scores corresponding to the multiple text sequences are determined, where the text sequence corresponding to each first path includes at least a combination of candidate text sequences represented by a first node and a second node connected thereto; a prefix tree updating module 740 is configured to delete at least one first path from the prefix tree to delete at least one first node if there is at least one first path with a score less than a score threshold, thereby obtaining an updated prefix tree; and a text sequence determination module 750 is configured to determine, based at least on the updated prefix tree, a first target text sequence that matches the first speech from at least one first candidate text sequence represented by at least one first node that has not been deleted. The apparatus 700 also includes a module configured to perform other operations described in the embodiments of the present disclosure.
[0086] FIG8 shows a block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 800 shown in FIG8 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 800 shown in FIG8 may be used to implement the electronic device 110 of FIG1 or the apparatus 700 of FIG7.
[0087] As shown in FIG8 , electronic device 800 is a general-purpose electronic device. Components of electronic device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 800.
[0088] The electronic device 800 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 800.
[0089] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0090] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 800 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0091] The input device 850 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 860 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 800 may also communicate with one or more external devices (not shown) via the communication unit 840 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 800, or with any device that allows the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0092] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for speech recognition, comprising: Based on at least one first candidate text sequence recognized from the first speech, adding at least one first node representing the at least one first candidate text sequence to the prefix tree; Based on at least one second candidate text sequence recognized from a second speech, adding at least one second node representing the at least one second candidate text sequence to the prefix tree, wherein the at least one second node is connected to each first node, and the second speech is collected immediately after the first speech; If the prefix tree includes a plurality of first nodes, determining scores corresponding to each of the plurality of text sequences based at least on the semantics of each of the plurality of text sequences corresponding to the plurality of first paths obtained from the prefix tree, wherein the text sequence corresponding to each first path includes at least a combination of candidate text sequences represented by a first node and a second node connected thereto; If there is at least one first path whose score is less than the score threshold, delete the at least one first path from the prefix tree to delete at least one first node, to obtain an updated prefix tree; as well as Based at least on the updated prefix tree, a first target text sequence matching the first speech is determined from at least one first candidate text sequence represented by at least one first node that has not been deleted.
2. The method according to claim 1, further comprising: If the prefix tree includes a single first node, a first candidate text sequence corresponding to the single first node is determined as the first target text sequence matching the first speech.
3. The method according to claim 1, wherein determining the scores corresponding to each of the plurality of text sequences comprises: Determine a plurality of first matching scores between a plurality of first candidate text sequences corresponding to the plurality of first nodes and the first speech, and a plurality of second matching scores between at least one second candidate text sequence corresponding to the at least one second node and the second speech; Determining a semantic score of each of the plurality of text sequences based on the semantics of each of the plurality of text sequences corresponding to the plurality of first paths; as well as For each first path among the multiple first paths, the score of the first path is determined based on a first matching score corresponding to a first node included in the first path, a second matching score corresponding to a second node, and a semantic score of the first path.
4. The method according to claim 1, wherein determining a first target text sequence matching the first speech from at least one first candidate text sequence represented by at least one first node that has not been deleted comprises: Based on at least one third candidate text sequence recognized from a third speech, adding at least one third node representing the at least one third candidate text sequence to the updated prefix tree, wherein the at least one third node is connected to each second node in a path corresponding to at least one first node that is not deleted, and the third speech is collected immediately after the second speech; If the updated prefix tree includes more than two first nodes, based on the semantics of the multiple text sequences corresponding to the multiple second paths corresponding to the more than two first nodes, determine the scores corresponding to the multiple text sequences corresponding to the multiple second paths, wherein the text sequence corresponding to each second path includes at least a combination of candidate text sequences represented by the first node, a second node connected to the first node, and a third node connected to the second node; If there is at least one second path whose score is less than the score threshold, deleting the at least one second path from the updated prefix tree to delete at least one first node, to obtain a prefix tree that is updated again; as well as At least based on the again updated prefix tree, a first target text sequence matching the first speech is determined from at least one first candidate text sequence represented by at least one first node that has not been deleted.
5. The method according to claim 1, wherein determining a first target text sequence matching the first speech from at least one first candidate text sequence represented by at least one first node that has not been deleted comprises: If the updated prefix tree is deleted until only a single first node remains, determining a first candidate text sequence corresponding to the single first node as the first target text sequence; as well as If no speech is acquired within the next predetermined time period, the first target text sequence is selected from the at least one first candidate text sequence based on a first matching score between each first candidate text sequence in the at least one first candidate text sequence and the first speech. The method according to claim 1 , wherein the first voice and the second voice are both voices that meet a predetermined duration.
7. The method according to claim 6, wherein the first voice and the second voice are both voices collected in real time, the first voice is the voice collected within the predetermined time period before the first moment, and the second voice is the voice collected within the predetermined time period after the first moment.
8. The method according to claim 1, wherein adding at least one first node representing the at least one first candidate text sequence to the prefix tree comprises: If the first speech is the initially collected speech, connecting the at least one first node after the root node of the prefix tree; as well as If the first speech is not the initially collected speech, the at least one node is connected to each node corresponding to the historical speech collected before the first speech in the prefix tree.
9. The method according to claim 1, further comprising: Before the first target text sequence is determined, a first text sequence to be presented in a user interface is selected from the at least one first candidate text sequence based on a first matching score between each of the at least one first candidate text sequence and the first speech.
10. The method according to claim 9, further comprising: If the second speech with a predetermined duration is not collected after the first speech, generating a predicted text sequence corresponding to the second speech based at least on the first text sequence; and The first text sequence and the predicted text sequence are presented in the user interface.
11. The method according to claim 10, further comprising: After collecting the second voice and recognizing the at least one second candidate text sequence, selecting a second text sequence to be presented in the user interface from the at least one second candidate text sequence based on a second matching score between the at least one second candidate text sequence and the second voice; and If the predicted text sequence is different from the second text sequence, the predicted text sequence is switched to the second text sequence in the user interface.
12. The method according to claim 9, further comprising: If the determined first target text sequence is different from the first text sequence, after determining the first target text sequence, the first text sequence presented in the user interface is switched to the first target text sequence. The method of claim 9 , wherein the user interface comprises a user interface of a conference application.
14. A device for speech recognition, comprising: A first node adding module, configured to add at least one first node representing at least one first candidate text sequence respectively to the prefix tree based on at least one first candidate text sequence recognized from the first speech; A second node adding module is configured to add at least one second node representing at least one second candidate text sequence respectively to the prefix tree based on at least one second candidate text sequence recognized from a second speech, wherein the at least one second node is connected to each first node, and the second speech is collected immediately after the first speech; a score determination module configured to determine scores corresponding to each of the plurality of text sequences based at least on the semantics of each of the plurality of text sequences corresponding to the plurality of first paths obtained from the prefix tree if the prefix tree includes a plurality of first nodes, wherein the text sequence corresponding to each first path includes at least a combination of candidate text sequences represented by a first node and a second node connected thereto; A prefix tree updating module, configured to delete at least one first path from the prefix tree to delete at least one first node if there is at least one first path whose score is less than a score threshold, to obtain an updated prefix tree; as well as The text sequence determination module is configured to determine, based at least on the updated prefix tree, a first target text sequence matching the first speech from at least one first candidate text sequence represented by at least one first node that has not been deleted.
15. An electronic device, comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 13 when executed by the at least one processing unit.
16. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Word graph rescoring method and system for deep learning language models
CN108415898A
Voice interaction method, electronic equipment and storage medium
CN114822532A
Speech recognition method and system based on domain classification and hot word prefix tree cluster search
CN115440197A
Speech recognition method, device, equipment and readable medium
CN117351963A
System for using silence in speech recognition
CN1307715A