Speech recognition post-processing method and device, storage medium, equipment and product

By matching the speech coded data output from the speech recognition model with the pre-constructed speech instruction structure tree, determining the target path and decoding, the problem of inaccurate speech recognition is solved and the accuracy and efficiency of speech interaction is improved.

CN120071935AActive Publication Date: 2025-05-30HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510495996.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-30
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In the prior art, the speech recognition model is prone to errors in individual words when recognizing the user's voice, resulting in inaccurate speech recognition and affecting the efficiency and experience of the user through the voice control device or software.

Method used

By obtaining the voice encoding data corresponding to the speech recognition results output by the speech recognition model, and determining the target path matching it from the pre-constructed voice command structure tree, by splicing and decoding the word segmentation of each node in the target path, the voice command text corresponding to the speech recognition results is obtained, thereby achieving accurate verification and correction of the speech recognition results.

Benefits of technology

It improves the accuracy and efficiency of speech recognition, improves the user's voice interaction experience, and ensures the accuracy of speech recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071935A_ABST
    Figure CN120071935A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech recognition, and particularly provides a speech recognition post-processing method and device, a storage medium, equipment and a product, and the method can comprise the steps: obtaining speech coding data corresponding to a speech recognition result outputted by a speech recognition model; determining a target path matched with the voice coding data from a pre-constructed voice instruction structure tree; wherein the voice instruction structure tree is constructed based on a voice interaction instruction text; the segmented word of each node on the voice instruction structure tree represents different instruction texts; and splicing the segmented words corresponding to each node in the target path and then decoding the segmented words to obtain a voice instruction text corresponding to the voice recognition result, and some embodiments of the application can realize correction of the voice recognition result and improve the accuracy of voice recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology. Specifically, it relates to a method, device, storage medium, equipment and product for post-processing speech recognition. Background Art

[0002] With the rapid development of large model technology, more and more devices or software products have accessed speech recognition models to enable users to perform corresponding operations by interacting with the devices or software through speech.

[0003] Currently, in the prior art, speech recognition models are used to recognize the speech content of users, and corresponding devices or software are controlled to perform corresponding operations through this speech content. When there are errors in individual words during the speech recognition of users by the speech recognition model, the speech recognition model will remind the user that it cannot be recognized and requires the user to speak again or not perform any operation; and due to the inaccurate speech recognized by the speech recognition model, users cannot accurately control relevant devices or software products through speech.

[0004] Therefore, how to provide a technical solution for a method of post-processing speech recognition with higher accuracy has become a technical problem that urgently needs to be solved. Summary of the Invention

[0005] Some embodiments of this application aim to provide a method, device, storage medium, equipment and product for post-processing speech recognition. Through the technical solutions of the embodiments of this application, accurate verification processing of speech recognition results can be achieved, the efficiency of speech recognition can be guaranteed, the accuracy and efficiency of speech interaction can be improved, and the user's speech interaction experience can be enhanced.

[0006] In a first aspect, some embodiments of this application provide a method for post-processing speech recognition, including: obtaining speech coding data corresponding to a speech recognition result output by a speech recognition model; determining a target path that matches the speech coding data from a pre-constructed speech instruction structure tree; wherein, the speech instruction structure tree is constructed based on speech interaction instruction texts; the word segmentation of each node on the speech instruction structure tree represents different instruction texts; splicing and decoding the word segments corresponding to each node in the target path to obtain a speech instruction text corresponding to the speech recognition result.

[0007] Some embodiments of this application perform matching processing on the speech coding data corresponding to the speech recognition result output by the speech recognition model and the speech instruction structure tree, determine the target path, splice and decode the word segments corresponding to the target path to obtain the speech instruction text, and can effectively verify the speech recognition result, guarantee the efficiency of speech recognition, improve the accuracy and efficiency of speech interaction, and enhance the user's speech interaction experience.

[0008] In some embodiments, determining a target path that matches the voice coding data from a pre-constructed voice instruction structure tree includes: obtaining data to be decoded, where the data to be decoded includes: the voice coding data and the node integration data corresponding to each batch of node data among multiple batches of node data pre-divided in the voice instruction structure tree; inputting the data to be decoded into a decoding network, and outputting the node scores of each node in each batch of node data; and determining the target path based on the node scores of each node.

[0009] In some embodiments of the present application, by inputting the data to be decoded into a decoding network, the node scores of each node are determined, and then the target path is obtained, so that the voice instruction text that best matches the voice recognition result can be obtained.

[0010] In some embodiments, the node integration data includes: batch node data and parent node data; the node integration data is obtained by the following method: concatenating the word segments of each node in each batch of node data to obtain the batch node data; and concatenating the word segments corresponding to the parent nodes of each node in each batch of node data to obtain the parent node data.

[0011] In some embodiments of the present application, by concatenating the word segments of each node in each batch of node data and the word segments of the corresponding parent nodes, the node integration data is obtained, providing data support for subsequent effective decoding.

[0012] In some embodiments, determining the target path based on the node scores of each node includes: obtaining the path scores of each path in multiple paths on the voice instruction structure tree through the node scores of each node; and taking the path corresponding to the maximum value among the path scores of each path as the target path.

[0013] In some embodiments of the present application, by calculating the path scores of each path and taking the path corresponding to the maximum value as the target path, the voice instruction text under the target path that best matches the voice recognition result can be accurately determined.

[0014] In some embodiments, before determining the target path that matches the voice coding data from a pre-constructed voice instruction structure tree, the method further includes: obtaining the voice interaction instruction texts acting on the product object, where there are multiple voice interaction instruction texts; converting each instruction text in the voice interaction instruction texts to obtain the word segments corresponding to each instruction text; and constructing the voice instruction structure tree according to the word segments corresponding to each instruction text.

[0015] In some embodiments of the present application, a voice instruction structure tree is constructed based on the voice interaction instruction texts, providing rich and effective data support for the subsequent verification and matching of voice recognition results.

[0016] In some embodiments, the method further includes: traversing the nodes in the voice command structure tree, and dividing the voice command structure tree into multiple batches of node data, where each batch of node data in the multiple batches of node data includes a set number of nodes, and there is no parent-child relationship between the nodes.

[0017] Some embodiments of the present application divide the voice command structure tree into multiple batches of node data, so as to facilitate subsequent batch matching of voice recognition results, and improve processing efficiency and accuracy.

[0018] In a second aspect, some embodiments of the present application provide an apparatus for post-processing voice recognition, including: an encoding module, configured to obtain voice encoding data corresponding to a voice recognition result output by a voice recognition model; a decoding module, configured to determine a target path matching the voice encoding data from a pre-constructed voice command structure tree; wherein, the voice command structure tree is constructed based on voice interaction command texts; the word segmentation of each node on the voice command structure tree represents different command texts; a text output module, configured to splice and decode the word segmentations corresponding to each node in the target path to obtain a voice command text corresponding to the voice recognition result.

[0019] In a third aspect, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any embodiment of the first aspect can be implemented.

[0020] In a fourth aspect, some embodiments of the present application provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the method described in any embodiment of the first aspect can be implemented.

[0021] In a fifth aspect, some embodiments of the present application provide a computer program product, the computer program product includes a computer program, wherein when the computer program is executed by a processor, the method described in any embodiment of the first aspect can be implemented. Description of the Drawings

[0022] In order to more clearly illustrate the technical solutions of some embodiments of the present application, the following will briefly introduce the drawings required to be used in some embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0023] Figure 1System diagram of post - processing for speech recognition provided for some embodiments of the present application; Figure 2 Method flowchart for constructing a speech instruction structure tree provided for some embodiments of the present application; Figure 3 Schematic diagram of a speech instruction structure tree provided for some embodiments of the present application; Figure 4 One of the method flowcharts for post - processing of speech recognition provided for some embodiments of the present application; Figure 5 Another method flowchart for post - processing of speech recognition provided for some embodiments of the present application; Figure 6 Block diagram of the composition of the device for post - processing of speech recognition provided for some embodiments of the present application; Figure 7 Schematic diagram of an electronic device provided for some embodiments of the present application. Detailed implementation manners

[0024] Next, the technical solutions in some embodiments of the present application will be described in conjunction with the accompanying drawings in some embodiments of the present application.

[0025] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0026] In the related art, with the emergence of AI interaction functions, in order to facilitate the interaction between users and electronic products, many electronic devices and application programs have introduced speech recognition models. When a user conducts voice interaction with a relevant electronic device or application program, due to the accent characteristics of each person, when the speech recognition model whisper performs speech recognition, there are sometimes cases of incorrect or missing character recognition. For example, for the voice instruction "Open a certain video, play the previous song, increase the volume by 20%", the speech recognition model whisper may recognize it as "A certain video, avoid a certain video, the previous song, increase by 20%" and other incorrect situations. In this case, the speech recognition model whisper usually reminds the user to speak again until an accurate voice instruction is recognized; or, the speech recognition model whisper will recommend relevant content for the user to select based on the recognition result.

[0027] As can be seen from the above - mentioned related art, in the prior art, the situation of incorrect or missing character recognition in the speech recognition model whisper cannot be corrected, which reduces the voice interaction efficiency between users and electronic products and affects the user experience.

[0028] In view of this, some embodiments of the present application provide a method for post-processing speech recognition, which can encode and decode the speech recognition result after being recognized by a speech recognition model, and select a speech instruction text that matches the speech recognition result from a pre-constructed speech instruction structure tree; wherein, the speech instruction structure tree is constructed by different instruction texts, so the speech instruction text corresponding to the obtained speech recognition result belongs to a standard speech instruction, realizing the timely verification or correction of the speech recognition result, thereby enabling the electronic product to serve the user accurately and providing a good speech interaction experience for the user.

[0029] The following combines the attached Figure 1 Exemplarily expounds the overall composition structure of the speech recognition post-processing system provided by some embodiments of the present application.

[0030] As Figure 1 shown, some embodiments of the present application provide a system diagram of speech recognition post-processing. The speech recognition post-processing system includes a terminal 100 and a speech recognition server 200. The user can perform a speech interaction with the terminal 100 itself or a certain application software on the terminal 100. After receiving the user's speech, the terminal 100 sends the speech to the speech recognition server 200. The speech recognition server 200 first recognizes the speech through the deployed speech recognition model to obtain a speech recognition result; then, the speech recognition server 200 verifies or corrects the speech recognition result based on the deployed speech instruction structure tree to obtain a final speech instruction text, and then sends it to the terminal 100, and the terminal 100 can perform related operations based on the speech instruction text.

[0031] In some other embodiments of the present application, if the terminal 100 can deploy a speech recognition model and a speech instruction structure tree to perform operations such as speech recognition, verification, and correction on the speech, the speech recognition server 200 may not be set at this time. Specifically, it can be set according to the actual situation, and the embodiments of the present application do not make specific limitations here. In addition, the terminal 100 can be a mobile terminal or a non-portable computer terminal, and the embodiments of the present application do not make specific limitations here.

[0032] In order to accurately verify or correct the speech recognition result output by the speech recognition model subsequently, it is first necessary to construct a speech instruction structure tree and deploy it in the speech recognition server 200. Based on this, the following combines the attached Figure 2 Exemplarily expounds the implementation process of constructing a speech instruction structure tree provided by some embodiments of the present application.

[0033] Please refer to the attached Figure 2 , Figure 2A method flowchart for constructing a voice command structure tree provided for some embodiments of the present application. The method for constructing the voice command structure tree may include: S210, obtaining the voice interaction command text acting on the product object, where the voice interaction command text includes multiple pieces.

[0034] For example, in some embodiments of the present application, a set of key voice commands are set according to the characteristics of the product object (such as an electronic device or an application program, etc.) for operating the product object. Taking a certain video software (as a specific example of the product object) as an example, the voice commands may include: open a certain video, close a certain video, increase the volume, turn off the volume, previous video, next video, etc. Taking a certain music playing software (as another specific example of the product object) as an example, the voice commands may include: open the music software, close the music software, increase the volume, decrease or turn off the volume, play the previous music, single loop, etc. The voice commands corresponding to different product objects are sorted into an instruction set list (as a specific example of the voice interaction command text). Among them, one product object corresponds to its specific instruction set list.

[0035] S220, converting each instruction text in the voice interaction command text to obtain the word segmentation corresponding to each instruction text.

[0036] For example, in some embodiments of the present application, for each instruction text, it is converted into an integer token (as a specific example of the word segmentation) by the tokenizer of the voice recognition model whisper. Among them, the Chinese of token is "token", which is a digital identifier used for authentication, security access control or data integrity protection in the field of information technology. The Chinese of tokenizer is "word segmenter".

[0037] S230, constructing the voice command structure tree according to the word segmentation corresponding to each instruction text.

[0038] For example, in some embodiments of the present application, a prefix tree is formed according to the tokens of each instruction text. For example, "open Tencent Video" and "open iQIYI" have the same prefix "open", then the tokens corresponding to "Tencent Video" and "iQIYI" have the same parent node, and the tree depth value of the node is saved; the voice command structure tree is constructed in this way. For example, as Figure 3 shown is the voice command structure tree corresponding to a certain video software, and nodes A~M represent different instruction texts.

[0039] In some embodiments of the present application, the method for constructing a voice command structure tree further includes: traversing the nodes in the voice command structure tree, and dividing the voice command structure tree into multiple batches of node data, where each batch of node data in the multiple batches of node data includes a set number of nodes, and there is no parent-child relationship between the nodes.

[0040] For example, in some embodiments of the present application, the nodes in the voice command structure tree are organized into batch data (i.e., batch node data) to facilitate subsequent verification or correction of the voice recognition results. For example, starting from the root node of the tree, take N nodes as a batch each time. There is no parent-child dependency relationship within the nodes of the same batch, and thus multiple batches of data are obtained. When traversing to node i, first take from the sibling nodes of i. If the number of sibling nodes of i is less than N, in order to ensure that the token prefix lengths in the same batch are the same, do not take other nodes; if the number of sibling nodes of i is sufficient for N, then take N nodes as a batch; start taking from the lower-level child nodes next; perform the above operations on each node in turn until all the nodes on the tree are taken, that is, each node belongs to a certain batch of data.

[0041] Specifically, in order to obtain multiple batches of node data, a first-in-first-out queue structure can be set up first, and write queue operations and take queue operations can be defined. Starting from the root node, write to the queue and take from the queue in a loop until all the nodes on the tree are added to the queue and all the nodes in the queue have entered the batch data. First, write the nodes to the queue, and then execute the following operations in turn: Take from the queue: Take nodes from the queue in turn until the number of nodes reaches the batch size (i.e., N) or the depth of the node is inconsistent with the depth of the nodes that have been taken, and obtain the nodes of the current batch, that is, obtain a batch of data containing N nodes or nodes with the depth meeting the requirements.

[0042] Write to the queue: Control the overall calculation scale by setting the maximum depth and the minimum score threshold. For each node of the current batch taken out, when the depth of the node is less than the preset maximum depth or the score of the node is greater than the preset minimum score threshold, add the child nodes of the node to the array to be written. At this time, after sorting the nodes in the array to be written in ascending order of the depth value, write them to the queue; that is, when the depth of the node is greater than the preset maximum depth and the score of the node is less than the preset minimum score threshold, no further comparison is made for this node, and thus a batch of data can be obtained.

[0043] Taking Figure 3 the voice command structure tree shown as an example, the token content in the nodes is: A: empty node; B: open, C: close, D: volume; E: QQ, F: album, G: album, H: increase, I: decrease, J: close; K: music, L: mailbox, M: 15%. When Figure 3During the traversal of the shown voice command structure tree, when there is no common parent node, an empty character node can be added as the root node, that is, the empty node A. Assume that the batch size N is 2. According to the above rules, starting from the child nodes of the root node, that is, the BCD nodes, the batch data taken in sequence is: BC, D, EF, GH, IJ, KL, M. It should be understood that in practical applications, the number of tree nodes may be very complex and large. Therefore, in addition to N, a preset maximum depth and a minimum score threshold need to be set to limit the number of nodes in the batch data. That is, under the two conditions of the preset maximum depth and the minimum score threshold, the batch size of the batch data may not be N, so as to facilitate the subsequent use of the batch data.

[0044] Specifically, in addition to the above rules, the way of sorting the voice command structure tree in batches can also be adjusted accordingly according to the actual situation. The embodiments of the present application are not limited thereto.

[0045] The following combines the attached Figure 4 Exemplarily illustrate the specific process of the voice recognition post-processing method provided by some embodiments of the present application.

[0046] Please refer to the attached Figure 4 , Figure 4 For a method flow of voice recognition post-processing provided by some embodiments of the present application, the voice recognition post-processing method may include: S410, obtain the voice encoding data corresponding to the voice recognition result output by the voice recognition model.

[0047] For example, in some embodiments of the present application, after obtaining the voice recognition result output by the voice recognition model for the user's voice, in order to avoid missing the beginning word, a blank data of a preset length (for example, 10000) is added in front. Convert the voice recognition result into mel spectrum data, and use an encoding network composed of one-dimensional convolution and residual attention layers (the number of residual attention layers can be set as needed) to encode the voice recognition result to obtain the voice encoding data.

[0048] S420, determine a target path matching the voice encoding data from a pre-constructed voice command structure tree; wherein, the voice command structure tree is constructed based on voice interaction command texts; the word segmentation of each node on the voice command structure tree represents different command texts.

[0049] For example, in some embodiments of the present application, after decoding the speech coding data through a decoding network composed of an embedding layer and several residual attention layers, the historical attention values on the speech instruction structure tree are combined to determine the target path in the speech instruction structure tree that is closest to the speech coding data. The speech instruction text corresponding to this target path is the speech instruction text that matches the speech recognition result. To reduce the number of decoding calculations, make full use of the already calculated results, and reduce repeated calculations, the historical attention values (used in the residual attention layer) of the current kvcache (full name Key-Value Cache) are saved for each node of the prefix tree (i.e., the speech instruction structure tree), and the kv cache of the node is initially empty.

[0050] It can be understood that in the embodiments of the present application, the speech coding data is compared and analyzed separately with the above-mentioned multiple batches of divided data to obtain the accurate speech instruction text corresponding to the speech recognition result, so that the terminal 100 can perform corresponding operations. Therefore, in some embodiments of the present application, S420 may include: S421, obtain the data to be decoded, where the decoded data includes: the speech coding data and the node integration data corresponding to each batch of node data in the multiple batches of node data pre-divided in the speech instruction structure tree. The node integration data includes: batch node data and parent node data; the node integration data is obtained by the following method: the word segments of each node in the multiple batches of node data are concatenated to obtain the batch node data; the word segments corresponding to the parent nodes of each node in the multiple batches of node data are concatenated to obtain the parent node data.

[0051] For example, in some embodiments of the present application, the implementation process of S421 is described by taking any one of the multiple batches of data as an example. In addition to using the speech coding data as the decoding input, for any batch of data, through the parent nodes of the nodes in this batch of data, all the tokens on the path of the nodes are found to obtain the node integration data for decoding input. That is, the tokens of all the nodes in any batch of data are concatenated to obtain the batch node data, which is used as the decoding input; at the same time, the parent node data obtained by concatenating the kv caches of the parent nodes of a batch of nodes is also used as the decoding input. For example, for node M (15%), the parent nodes H (increase) and D (volume) are found in sequence, and the parent node data obtained by concatenating the tokens corresponding to "increase the volume by 15%" is used as the decoding input.

[0052] S422, input the data to be decoded into the decoding network, and output the node scores of each node in the multiple batches of node data.

[0053] For example, in some embodiments of the present application, when the decoding network decodes the data to be decoded, after the voice decoding data and the token of each node in the data to be decoded are calculated by the embedding layer, the new kv values are obtained using the kv network of the residual attention layer, and are concatenated with the historical attention values in the kv cache. The attention values are calculated in the residual attention layer, and then the decoding result is obtained through the linear layer. After the decoding is completed, the updated kv values in any batch of data are assigned to each corresponding node, and at the same time, the decoding score (as a specific example of the node score) is assigned to each node, that is, the node score of each node can be obtained through the decoding network.

[0054] S423. Determine the target path based on the node scores of each node.

[0055] For example, in some embodiments of the present application, after obtaining the decoding scores of each node in the prefix tree, the comprehensive scores of each path in the prefix tree can be calculated to determine the target path.

[0056] In some embodiments of the present application, S423 may include: obtaining the path scores of each path in the multiple paths on the voice instruction structure tree through the node scores of each node; taking the path corresponding to the maximum value in the path scores of each path as the target path.

[0057] For example, in some embodiments of the present application, the decoding scores of each node on each path are weighted and averaged to obtain the comprehensive score of each path (as a specific example of the path score). The path with the largest comprehensive score among all paths is used as the target path. For example, in the voice instruction structure tree shown, there are still multiple paths, namely: ABEK, ABEL, ACG, etc. By scoring each path, the best target path is selected. This comprehensive score can indirectly represent the similarity between the speech recognition result and the speech interaction instruction text. In addition, in some embodiments, additional conditions can also be set, that is, when the comprehensive score is greater than the set threshold (for example, 0.1), it is adopted, otherwise, the speech recognition ASR (full name: Automatic Speech Recognition) result is obtained through the normal decoding method. Figure 3

[0058] S430. Concatenate the word segments corresponding to each node in the target path and decode them to obtain the speech instruction text corresponding to the speech recognition result.

[0059] For example, in some embodiments of the present application, after concatenating the tokens of each node in the target path, they are decoded by a tokenizer to obtain the original text (as a specific example of the speech instruction text) and then output.

[0060] The following will exemplarily elaborate on the specific process of the speech recognition post - processing method provided by some embodiments of the present application in conjunction with the appended Figure 5 drawings.

[0061] Please refer to the appended Figure 5 drawings, Figure 5 which shows the flowchart of a speech recognition post - processing method provided by some embodiments of the present application.

[0062] The following will exemplarily elaborate on the above - mentioned implementation process.

[0063] S510: Based on the tokens corresponding to the speech interaction instruction text acting on the product object, construct a speech instruction structure tree, and divide the nodes in the speech instruction structure tree into batches.

[0064] S520: Obtain the speech coding data corresponding to the speech recognition result output by the speech recognition model, and obtain the node integration data.

[0065] S530: Input the speech coding data and the node integration data into the decoding network to determine the decoding scores of each node in the speech instruction structure tree.

[0066] S540: Based on the decoding scores of each node, calculate the path scores for each path in the speech instruction structure tree.

[0067] S550: Take the path corresponding to the maximum value among all the path scores as the target path.

[0068] S560: Concatenate and decode the tokens corresponding to each node in the target path to obtain the speech instruction text corresponding to the speech recognition result.

[0069] It should be noted that the specific implementation processes of S510 - S560 can refer to the method embodiments provided above. To avoid repetition, the detailed descriptions are appropriately omitted here.

[0070] As can be seen from some embodiments of the present application above, the present application constructs a prefix tree structure, reduces the number of decoding times, accelerates the comparison and verification of decoding scores for a large number of data, realizes the correction of speech recognition results for a fixed speech interaction instruction set, and improves the accuracy of speech recognition.

[0071] Please refer to Figure 6 the drawings, Figure 6 which shows the block diagram of the composition of the speech recognition post - processing device provided by some embodiments of the present application. It should be understood that this speech recognition post - processing device corresponds to the above - mentioned method embodiments and can execute each step involved in the above - mentioned method embodiments. The specific functions of this speech recognition post - processing device can refer to the descriptions above. To avoid repetition, the detailed descriptions are appropriately omitted here.

[0072] Figure 6 The speech recognition post-processing device includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the speech recognition post-processing device, and the speech recognition post-processing device includes: an encoding module 610, which is used to obtain speech encoding data corresponding to the speech recognition result output by the speech recognition model; a decoding module 620, which is used to determine a target path matching the speech encoding data from a pre-constructed speech instruction structure tree; wherein the speech instruction structure tree is constructed based on the speech interaction instruction text; the word segmentation of each node on the speech instruction structure tree represents different instruction texts; a text output module 630, which is used to splice and decode the word segmentations corresponding to each node in the target path to obtain the speech instruction text corresponding to the speech recognition result.

[0073] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.

[0074] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the operations of the method corresponding to any of the above methods provided in the above embodiments.

[0075] Some embodiments of the present application further provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operations corresponding to any of the above methods provided in the above embodiments.

[0076] like Figure 7 As shown, some embodiments of the present application provide an electronic device 700, which includes: a memory 710, a processor 720, and a computer program stored in the memory 710 and executable on the processor 720, wherein the processor 720 can implement a method as described in any of the above embodiments when reading the program from the memory 710 through a bus 730 and executing the program.

[0077] Processor 720 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 720 can be a microprocessor.

[0078] The memory 710 can be used to store instructions executed by the processor 720 or data related to the instruction execution process. These instructions and / or data can include code for implementing some or all of the functions of one or more modules described in the embodiments of the present application. The processor 720 of the embodiments of the present disclosure can be used to execute the instructions in the memory 710 to implement the methods shown above. The memory 710 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.

[0079] The above are only the embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0080] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

[0081] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

Claims

1. A method for post-processing of speech recognition, characterized in that: include: Obtaining speech coding data corresponding to the speech recognition result output by the speech recognition model; Determining a target path matching the voice coding data from a pre-constructed voice command structure tree; wherein the voice command structure tree is constructed based on the voice interaction command text; and the word segmentation of each node on the voice command structure tree represents different command texts; The word segments corresponding to each node in the target path are concatenated and decoded to obtain the voice instruction text corresponding to the voice recognition result.

2. The method according to claim 1, characterized in that Determining a target path matching the voice coding data from a pre-constructed voice instruction structure tree includes: Acquire data to be decoded, wherein the decoded data includes: the speech coding data and node integration data corresponding to each batch of node data in a plurality of batches of node data pre-divided in the speech command structure tree; Inputting the data to be decoded into a decoding network, and outputting the node score of each node in each batch of node data; The target path is determined based on the node score of each node.

3. The method according to claim 2, characterized in that The node integration data includes: batch node data and parent node data; the node integration data is obtained by the following method: Concatenating the word segments of each node in each batch of node data to obtain the batch of node data; The word segments corresponding to the parent nodes of each node in each batch of node data are concatenated to obtain the parent node data.

4. The method according to claim 2 or 3, characterized in that The determining the target path based on the node score of each node includes: Obtaining a path score of each of the multiple paths on the voice command structure tree through the node score of each node; The path corresponding to the maximum value of the path scores of each path is taken as the target path.

5. The method according to any one of claims 1 to 3, characterized in that Before determining the target path matching the voice coding data from the pre-constructed voice instruction structure tree, the method further includes: Acquire the voice interaction instruction text acting on the product object, wherein the voice interaction instruction text includes multiple items; Convert each instruction text in the voice interaction instruction text to obtain a word segment corresponding to each instruction text; The voice command structure tree is constructed according to the word segments corresponding to each command text.

6. The method according to claim 5, characterized in that The method further comprises: The nodes in the voice command structure tree are traversed, and the voice command structure tree is divided into multiple batches of node data, wherein each batch of node data in the multiple batches of node data includes a set number of nodes, and there is no parent-child relationship between the nodes.

7. A device for post-processing speech recognition, characterized in that: include: The encoding module is used to obtain the speech encoding data corresponding to the speech recognition result output by the speech recognition model; A decoding module is used to determine a target path matching the voice coding data from a pre-constructed voice command structure tree; wherein the voice command structure tree is constructed based on the voice interaction command text; and the word segmentation of each node on the voice command structure tree represents different command texts; The text output module is used to concatenate and decode the word segments corresponding to each node in the target path to obtain the voice instruction text corresponding to the voice recognition result.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program executes the method according to any one of claims 1 to 6 when executed by a processor.

9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the computer program executes the method according to any one of claims 1 to 6 when being run by the processor.

10. A computer program product, characterized in that The computer program product comprises a computer program, wherein the computer program executes the method according to any one of claims 1 to 6 when executed by a processor.

Citation Information

Patent Citations

  • Control instruction determination method and device, electronic equipment and storage medium

    CN112017662A

  • Speech recognition method and device, storage medium and electronic equipment

    CN112133285A

  • Speech recognition method and device, equipment and medium

    CN112863489A

  • Intention generation method, server, voice control system and readable storage medium

    CN113239178A

  • Data processing method and equipment

    CN114328525A

Cited By

  • Voice instruction recognition method and device

    CN121096341A