Error correction method and device for voice intention recognition and storage medium
By converting speech recognition information into text information and performing intention recognition, forming a combined map, the problem of low accuracy of speech intention recognition is solved, and higher accuracy of speech intention recognition and user experience is achieved.
Patent Information
- Application Number
- CN202510035485.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-27
AI Technical Summary
The accuracy of speech intention recognition is low, resulting in poor user experience.
By converting the speech recognition information into text information, preprocessing and intent recognition are performed, a combined map of intent recognition information is formed, and error correction is carried out through a visual interactive interface to improve the accuracy of speech intent recognition.
Through the combined display of text and maps, the error correction process of speech recognition information is intuitively enhanced, the accuracy of speech intention recognition is improved, and the accuracy of users' understanding of the system is enhanced.
Smart Images

Figure CN120048247A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent interaction technologies, and in particular, to an error correction method, device, and storage medium for speech intent recognition. Background Art
[0002] With the development of speech recognition and natural language processing technologies, speech input has become one of the preferred human-computer interaction means in various intelligent devices. In related technologies, due to the influence of environmental noise and specific application scenarios of users during the recognition process, there are often some deviations between the accuracy of speech and intent recognition and the true intent of users, thus unable to meet user needs and resulting in a poor user experience. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide an error correction method for speech intent recognition, which effectively solves the technical problem of low accuracy of speech intent recognition.
[0004] The above technical problem is solved by the following technical solutions:
[0005] An error correction method for speech intent recognition, the method comprising:
[0006] Input speech recognition information and convert the speech recognition information into text information;
[0007] Preprocess the text information;
[0008] Perform intent recognition on the preprocessed text information to obtain intent recognition information;
[0009] Form a combined graph of the intent recognition information based on the intent recognition information;
[0010] Transmit the text information and the combined graph to a visual interaction interface;
[0011] Perform error correction on the speech recognition information based on the visual interaction interface to obtain error-corrected speech recognition information.
[0012] Compared with the background art, the beneficial effects of the error correction method for speech intent recognition of the present invention are as follows: First, the system receives the input speech recognition information. For example, after a user emits speech through a speech input device and the speech recognition process is performed, the obtained speech recognition information. Then, according to the language model in the system, the speech recognition information can be converted into text information, presenting the speech content in the form of text, which is convenient for the system to further analyze the content expressed by the user according to the text processing logic. After obtaining the text information, preprocess the text information to facilitate subsequent matching and comparison with the target database and more refined analysis of the text semantics.
[0013] Further, perform intent recognition on the preprocessed text information according to relevant algorithms and models in natural language processing. By analyzing elements such as vocabulary and semantic relationships in the text information, determine the true intention that the user wants to express, and obtain clear intent recognition information. Then, according to the obtained intent recognition information, construct a combined graph corresponding to the intent recognition information, so that the semantic structure of the sentence and the association between semantic elements can be intuitively presented through the combined graph.
[0014] Transmit the processed text information and the constructed combined graph to the visual interaction interface, which can display the speech recognition information in the form of text information, and at the same time present the intent recognition information in the form of a combined graph, enabling the user to intuitively see the text content corresponding to the voice input and the semantic structure relationship. Finally, using the display function provided by the visual interaction interface, the user can view the text information and the combined graph; in the case where the system cannot automatically correct the speech recognition information, manual intervention can be performed through the visual operation interface. For example, click or select operations are performed on the visual operation interface to obtain the corrected speech recognition information, making the system's understanding of the user's voice input more accurate, and subsequent tasks can also be executed based on more accurate information. By using the method of combining text and graph to display the recognition content of the voice command in the background, the relatively abstract process of algorithm recognition is visualized, making the error correction process of the speech recognition information more intuitive, facilitating user understanding, and effectively improving the accuracy of speech intent recognition.
[0015] In one embodiment, the preprocessing of the text information includes:
[0016] Perform word segmentation on the text information to obtain multiple words;
[0017] Perform similarity matching between the multiple words and preset words in the target database.
[0018] In one embodiment, the performing similarity matching between the multiple words and preset words in the target database includes:
[0019] In response to successful similarity matching, record the node type of the word as successfully matched, and construct a candidate word list from the preset words with a similarity greater than the first target value;
[0020] In response to failed similarity matching, replace the word with the preset word with a similarity greater than the first target value and the highest similarity, record the node type of the word as automatically replaced, and construct a candidate word list from the other preset words;
[0021] In response to the failure of the similarity matching and the non-existence of preset vocabulary in the target database with a similarity greater than the first target value, record the node type of the vocabulary as unrecognized.
[0022] In one embodiment, after performing the similarity matching between the multiple vocabularies and the preset vocabulary in the target database, the method further includes:
[0023] Based on the multiple candidate vocabulary lists and the node types, combine the multiple candidate vocabulary lists to form multiple combined statements;
[0024] Perform similarity comparison between the multiple combined statements and the text information, and sort them according to the similarity;
[0025] Remove the combined statements with a similarity less than the second target value.
[0026] In one embodiment, the forming of the combined graph of the intent recognition information based on the intent recognition information includes:
[0027] Perform intent understanding on the multiple combined statements;
[0028] Based on the combined statements after intent understanding, and in combination with the point set, edge set, and the similarity, combine to form the combined graph of the intent recognition information; the point set represents the semantic nodes of the intent, slot, and entity corresponding to the speech recognition information, and the edge set represents the association relationship between different semantic nodes.
[0029] In one embodiment, the method further includes:
[0030] Determine whether there is a legal instruction in the visual interaction interface, where the legal instruction represents information that includes intent, slot, entity and meets the requirements of preset rules;
[0031] In response to the non-existence of the legal instruction in the visual interaction interface, input a target node into the visual interaction interface;
[0032] Associate the target node with the combined graph to correct the combined graph.
[0033] In one embodiment, after obtaining the corrected speech recognition information, the method further includes:
[0034] Generate a correction instruction based on the corrected speech recognition information;
[0035] Perform keyword mapping on the speech recognition information in the correction instruction and the speech recognition information in the original instruction to form the mapping relationship between the correction instruction and the original instruction, where the original instruction is the instruction generated according to the input speech recognition information;
[0036] Store the mapping relationship in a target database.
[0037] On the other hand, the present invention also provides an error correction device for speech intention recognition, including:
[0038] An input module, configured to input speech recognition information and convert the speech recognition information into text information;
[0039] A preprocessing module, configured to preprocess the text information;
[0040] An intention recognition module, configured to perform intention recognition on the preprocessed text information to obtain intention recognition information;
[0041] A combination module, configured to form a combined graph of intention recognition information based on the intention recognition information;
[0042] A transmission module, configured to transmit the text information and the combined graph to a visual interaction interface;
[0043] An error correction module, configured to correct the speech recognition information based on the visual interaction interface to obtain corrected speech recognition information.
[0044] In one embodiment, it further includes:
[0045] A backtracking module, configured to generate an error correction instruction based on the corrected speech recognition information; perform keyword mapping on the speech recognition information in the error correction instruction and the speech recognition information in the original instruction to form a mapping relationship between the error correction instruction and the original instruction, where the original instruction is an instruction generated according to the input speech recognition information; store the mapping relationship in a target database.
[0046] On the other hand, the present invention also provides an electronic device, including: a memory, a processor, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete mutual communication through the communication bus;
[0047] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the error correction method for speech intention recognition as described above.
[0048] On the other hand, the present invention also provides a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a processor to execute the error correction method for speech intention recognition as described above when executed. Description of the Drawings
[0049] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a schematic flowchart of an error correction method for speech intent recognition according to an embodiment of the present invention;
[0051] Figure 2 It is a schematic diagram of a Web system architecture according to an embodiment of the present invention;
[0052] Figure 3 It is a schematic diagram of a Web interface layout method according to an embodiment of the present invention;
[0053] Figure 4 It is a schematic flowchart of an error correction method for speech intent recognition according to another embodiment of the present invention;
[0054] Figure 5 It is a schematic flowchart of an error correction method for speech intent recognition according to another embodiment of the present invention;
[0055] Figure 6 It is a schematic diagram of the structure of an error correction device for speech intent recognition according to an embodiment of the present invention;
[0056] Figure 7 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Specific Embodiments
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0058] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0059] Such asFigure 2 As shown in the figure, an embodiment of the present invention provides a Web system based on the browser / server (B / S) architecture, including a presentation layer, an interface layer, a business layer, and an operation support layer. Among them, the B / S architecture includes a front-end Web interface and a back-end business processing service; the front-end Web interface provides a visual operation interface to display the speech and intent recognition results and related UI interaction controls, and assists the Web system to complete speech and intent error correction; the back-end business processing service includes an initial metadata modeling module, an upstream message receiving module, and an error correction information storage module. Among them, the UI interaction controls include buttons, lists, modal boxes, etc.
[0060] In the presentation layer, customized UI controls are integrated in a Web programming manner and presented in a browser rendering manner, which can be adaptively displayed on mobile terminals or clients, etc.; among them, the Web interface layout method is as Figure 3 shown.
[0061] In the interface layer, Nginx (a high-performance Web server and reverse proxy server) is used as a reverse proxy to provide the service routes exposed by the business layer to the presentation layer; in the Web system, the interaction protocol includes the Hypertext Transfer Protocol (HTTP) sent by the Web side and WebSocket (a protocol for full-duplex communication on a single TCP connection) pushed by the server side.
[0062] In the business layer, the metadata modeling module models the metadata in a specific domain and stores it in a graph database, and constructs a corresponding thesaurus for the key node information for subsequent input retrieval. The upstream message receiving module is responsible for listening to the upstream speech and intent recognition results, formatting them into interface rendering data, and pushing them to the Web side through the WebSocket protocol. The error correction information sent by the front end through the HTTP protocol is formatted and stored in the graph database for subsequent training of the intent recognition model or automatic replacement of error correction words. The operation support layer includes the graph database (Neo4j), message bus (message middleware), thesaurus database (ElasticSearch), and relational database (Mysql) required by the business layer. Among them, the Web system can run on Linux / Windows servers or other operating systems, and can be specifically selected according to actual usage requirements.
[0063] According to an embodiment of the present invention, as Figure 1 shown, there is also provided a method for correcting speech intent recognition, including the following steps:
[0064] Step S100, input speech recognition information, and convert the speech recognition information into text information;
[0065] Step S200: Preprocess the text information;
[0066] Step S300: Perform intent recognition on the preprocessed text information to obtain intent recognition information;
[0067] Step S400: Form a combined graph of the intent recognition information based on the intent recognition information;
[0068] Step S500: Transmit the text information and the combined graph to the visual interaction interface;
[0069] Step S600: Correct the speech recognition information based on the visual interaction interface to obtain the corrected speech recognition information.
[0070] In this embodiment, first, the system receives the input speech recognition information. For example, after the user emits speech through a speech input device and undergoes speech recognition processing, the obtained speech recognition information. Then, according to the language model in the system, the speech recognition information can be converted into text information, presenting the speech content in the form of text, which is convenient for the system to further analyze the content expressed by the user according to the text processing logic. After obtaining the text information, preprocess the text information to facilitate subsequent matching and comparison with the target database and more refined analysis of the text semantics.
[0071] Further, perform intent recognition on the preprocessed text information according to relevant algorithms and models in natural language processing. By analyzing elements such as vocabulary and semantic relationships in the text information, to determine the true intention that the user wants to express, and obtain clear intent recognition information. Then, according to the obtained intent recognition information, construct a combined graph corresponding to the intent recognition information, so that the semantic structure of the sentence and the association between semantic elements can be intuitively presented through the combined graph.
[0072] Transmit the processed text information and the constructed combined graph to the visual interaction interface, which can display the speech recognition information in the form of text information, and at the same time present the intent recognition information in the form of a combined graph, enabling users to intuitively see the text content corresponding to the speech input and the semantic structure relationship. Finally, using the display function provided by the visual interaction interface, users can view the text information and the combined graph; in the case where the system cannot automatically correct the speech recognition information, manual intervention can be performed through the visual operation interface. For example, click or select operations can be performed on the visual operation interface to obtain the corrected speech recognition information, making the system's understanding of the user's speech input more accurate, and subsequent tasks can also be executed based on more accurate information. By using the method of combining text and graph to display the recognition content of the speech command in the background, the relatively abstract process of algorithm recognition is visualized, making the error correction process of speech recognition information more intuitive and easy for users to understand, and can also effectively improve the accuracy of speech intent recognition.
[0073] In one embodiment, before step S100, it further includes: initialization of the graph database and initialization of the thesaurus database.
[0074] Specifically, the initialization of the graph database includes summarizing all instruction sets in the upstream intent understanding service, and constructing intents, slots, and entities in natural language processing as nodes in the graph database; each node attribute includes: name, label, and pinyin abbreviation, a total of 3 attributes; the edge types include: the edge between intent and slot, and the edge between slot and entity, 2 types. The initialization of the thesaurus database includes that the system constructs a specific thesaurus prefix tree with the pinyin abbreviations of the vocabulary. The thesaurus prefix tree includes a root node and several child nodes. The root node does not contain characters, and the remaining child nodes contain a character and one or more pointers to child nodes. The path from the root node to a certain child node is constructed into a string, which becomes the pinyin abbreviation or the prefix of the abbreviation of the keyword vocabulary. Thus, keywords within the scope of intent understanding, such as intents, slots, entity names, etc., can be retrieved efficiently.
[0075] As Figure 4 shown, in one embodiment, step S200 includes the following steps:
[0076] Step S210: Perform word segmentation on the text information to obtain multiple words;
[0077] Step S220: Perform similarity matching between the multiple words and the preset words in the target database.
[0078] In this embodiment, the target database can be the target database. The complete text information is split into multiple relatively independent words according to the grammar rules and word formation characteristics of the language, laying a foundation for subsequent operations such as comparison with the thesaurus database. Through word segmentation, the semantic composition of the sentence can be analyzed more precisely.
[0079] Each word obtained through word segmentation is successively matched with the words in the thesaurus database for similarity. The thesaurus database is a powerful search engine that can efficiently store and retrieve a large amount of word data. The system uses the thesaurus database to find the words in the thesaurus that are similar to the words in the text information, and calculates the similarity value between the two through a specific algorithm to determine the level of matching between the two, providing a basis for subsequent processing of different situations.
[0080] As Figure 5 shown, in one of the embodiments, step S220 includes the following steps:
[0081] Step S221, in response to successful similarity matching, record the node type of the word as successfully matched, and construct a candidate word list from the preset words with a similarity greater than the first target value;
[0082] Step S222, in response to failed similarity matching, replace the word with the preset word with a similarity greater than the first target value and the highest similarity, record the node type of the word as automatically replaced, and construct a candidate word list from other preset words;
[0083] Step S223, in response to failed similarity matching and the absence of a preset word with a similarity greater than the first target value in the target database, record the node type of the word as unrecognized.
[0084] In this embodiment, when a certain word is successfully matched with the word in the thesaurus database, the system will mark the word as successfully matched and can be denoted as S, thereby identifying that there is a corresponding appropriate word for the word in the thesaurus. At the same time, a candidate word list is constructed from the preset words with a similarity greater than the first target value, so as to facilitate auxiliary functions in subsequent scenarios such as further analyzing the sentence semantics or there being multiple possible understandings. Among them, the first target value can be 0.7, 0.8, 0.9, etc., and can be set according to the usage requirements.
[0085] If a certain word fails in the similarity matching with the thesaurus database, at this time, multiple preset words with similarity greater than the first target value are obtained according to the thesaurus database, and the preset word with the highest similarity among all the preset words is used to replace the word in the text information. At the same time, the system marks this word as automatically replaced and records it as A; and further constructs the remaining preset words into a candidate word list, which is also to provide more possible word choices for subsequent analysis to make the processing of sentences more flexible and accurate.
[0086] If the similarity matching of a certain word fails and no preset word with similarity greater than the first target value can be found in the thesaurus database, then this word is marked as unrecognized and recorded as E; correspondingly, the candidate word list of this word is empty.
[0087] In one embodiment, after step S220, it further includes:
[0088] Based on multiple candidate word lists and node types, combine multiple candidate word lists to form multiple combined sentences;
[0089] Compare the similarity of multiple combined sentences with the text information and sort them according to the similarity;
[0090] Remove the combined sentences with similarity less than the second target value.
[0091] In this embodiment, the system transmits the result list after similarity matching to the intent understanding preprocessing module. Among them, the result list includes key elements such as nodes, node types, and multiple candidate word lists. Then, according to the intent understanding preprocessing module, multiple candidate word lists are combined to form multiple combined sentences. For example, word B and word C are combined into a new sentence form. By combining multiple candidate word lists to construct different sentence possibilities as comprehensively as possible, it helps to explore various potential intents of the sentence. For the formed multiple combined sentences, compare the similarity of multiple combined sentences with the text information and sort them according to the similarity. Through sorting, it can be more intuitively seen which combined sentences are closer to the original input text information in semantics. Further remove the combined sentences with similarity less than the second target value to reduce the processing burden of the subsequent intent understanding module and improve the processing accuracy at the same time. Among them, the second target value can be 0.01, 0.05, 0.1, 0.2, etc.; specifically, it can be set according to actual usage requirements. In addition, instructions with the same intent can be merged to facilitate the integration of repeated or similar intent expressions, so that the subsequent intent understanding module receives an optimized and sorted sentence combination and instruction content, which is more conducive to accurately grasping the core intent of the user input sentence and making corresponding processing.
[0092] In one embodiment, a combined graph of intent recognition information is formed based on the intent recognition information, including: understanding the intent of multiple combined statements; combining the combined statements after intent understanding, together with the point set, edge set, and similarity, to form a combined graph of intent recognition information; the point set represents semantic nodes of intents, slots, and entities corresponding to the speech recognition information, and the edge set represents the association relationships between different semantic nodes.
[0093] In this embodiment, the intent understanding module deeply analyzes each combined statement to dig out the true purpose that the user wants to express through the combined statement. When the intent understanding module completes the intent understanding of each combined statement, it will sequentially send the combined statements after intent understanding to the graph calculation module. The graph calculation module will query the graph database to obtain the point set, edge set, and similarity information of the combined statement, and further perform combined calculations on the point set, edge set, and similarity information to integrate the scattered point, edge, and similarity data, thereby forming a complete combined graph. The combined graph can intuitively present the semantic structure of the statement and the association situation between each semantic node, and finally return the combined graph to the front-end visualization module for further display to the user. Among them, the point set represents semantic nodes of intents, slots, and entities corresponding to the speech recognition information, and the edge set represents the association relationships between different semantic nodes.
[0094] In an implementable manner, relevant visual encoding is performed on different types of point sets and edge sets based on the front-end visualization module in combination with the similarity. Specifically, for nodes with a higher similarity, they are rendered in green, and at the same time, the edges connected to them are shown in a thick solid line. For nodes with a lower similarity, they are rendered in red, and the corresponding edges are shown in a dashed line. To help users more intuitively understand the content of the graph.
[0095] In one embodiment, it further includes:
[0096] Determine whether there is a legal instruction in the visual interaction interface, and the legal instruction represents information that includes intent, slot, entity, and meets the requirements of the preset rules;
[0097] In response to the absence of a legal instruction in the visual interaction interface, input a target node in the visual interaction interface;
[0098] Associate the target node with the combined graph to correct the combined graph.
[0099] In this embodiment, after a series of processes, if there is still no legal instruction in the visual interaction interface, that is, there is no information in the visual interaction interface that contains complete semantic nodes such as intent, slot, entity, etc. and meets the requirements of preset rules, then the user needs to participate in further improving the combined graph information. At this time, the user can input the target node through the front-end customized keyboard and then click "associate" to associate the target node with the combined graph, so as to correct the combined graph information. For example, if a certain key entity node is missing in the combined graph, the user can manually input and associate it to the corresponding position to make the combined graph more complete and accurate, and better reflect the intention that the user wants to convey.
[0100] In an implementable manner, the system will automatically send the legal instruction with the highest similarity to the background engine status statistical calculation module in sequence. The background engine status statistical calculation module will summarize the current status information of the system in real time. Taking the battle game system as an example, the status information includes: the number and status of the equipment and personnel on both sides of the enemy and us, etc. If the legal instruction conflicts with the current engine status, the background engine status statistical calculation module will identify the legal instruction as "error" and automatically skip it; at the same time, the front end clears the content related to the legal instruction and automatically processes the next input statement; to ensure that the processing flow of the entire system can continue and proceed orderly, and will not be confused or perform unreasonable operations due to incorrect instructions. If the legal instruction does not conflict with the current engine status, the background engine status statistical calculation module will identify the legal instruction as "correct" and send it to the downstream module to execute the legal instruction, so as to achieve the purpose that the user wants to achieve through the input statement and complete the entire complete process from user input to system execution. It can be understood that the user can send the legal instruction to the downstream module by means of keyboard input and manual click.
[0101] In one of the embodiments, after step S600, it further includes:
[0102] Generate an error correction instruction based on the error-corrected speech recognition information;
[0103] Perform keyword mapping on the speech recognition information in the error correction instruction and the speech recognition information in the original instruction to form a mapping relationship between the error correction instruction and the original instruction, where the original instruction is the instruction generated according to the input speech recognition information;
[0104] Store the mapping relationship in the target database.
[0105] In this embodiment, an error correction instruction is generated based on the corrected speech recognition information, and then an original instruction is generated based on the input speech recognition information. When the backtracking module receives the error correction instruction sent by the front end, it maps the keywords included in the speech recognition information in the error correction instruction to the keywords included in the speech recognition information in the original instruction, thereby forming a mapping relationship between the error correction instruction and the original instruction, which includes "original text" -> "corrected text" and "original intention" -> "corrected intention". Among them, "original text" -> "corrected text" records the content of the original statement initially input by the user and the corresponding correct and reasonable corrected statement text after error correction, presenting the transformation process of the text from error to correct at the text level. "Original intention" -> "corrected intention" records the intention expressed by the original statement and what the accurate intention is after correction, thereby helping to grasp the changes before and after error correction at a deep semantic level.
[0106] Furthermore, the mapping relationship between the error correction instruction and the original instruction is stored in the target database, so that these information can be stored permanently and stably, and will not be lost due to factors such as the temporary running state change of the system. This facilitates providing a reliable data basis for subsequent multiple application scenarios and conveniently calling the organized error correction mapping relationship data at any time.
[0107] Furthermore, in the offline state, the mapping relationship between the original instruction and the error correction instruction stored in the target database can be used to train the intention understanding model. The intention understanding model can continuously adjust its own parameters and internal logic to improve the accuracy of judging the intentions of different statements, so that it can more accurately grasp the intentions when actually processing user input subsequently, reduce the occurrence of misunderstanding situations, and optimize the performance of the entire intention understanding.
[0108] Even further, when the system encounters a statement expression similar to the previous error again during the real-time operation process, it can quickly find the corresponding mapping relationship in the target database, obtain the correct words that should be replaced, correct the statement in time, replace the wrong words with the correct ones, thereby improving the accuracy of real-time processing of user input, avoiding the recurrence of the same error, and enhancing the user experience.
[0109] In a specific embodiment, a scenario where a player controls a game through voice commands is used to illustrate step by step the flow steps of the error correction method for speech intention recognition provided in this embodiment and the key module functions in the Web system.
[0110] First, the graph database is initialized. In the scenario of controlling the game, it mainly includes instructions such as game start / stop, viewing equipment, and operating equipment. Taking the zooming in of the game map as an example, its corresponding intention, slot, and entity relationship are:
[0111] Intention: Enlarge the map (map_larger);
[0112] Slots: value, num_unit;
[0113] Entities: value entity (entity values supported by the system, 1, 2, 5, 10…); unit entity (unit entity values supported by the system, multiples…);
[0114] Example of the construction statement for the graph database:
[0115] Intention node: MERGE(n:intention{name:"map_larger",initials:"FDDT",label:"Enlarge the map"});
[0116] Slot node: MERGE(n:slot{name:"value",initials:"SZ",label:"Value"});
[0117] Entity node: MERGE(n:entity_value{name:"1",initials:"1",label:"1"});
[0118] Edge between intention and slot: MATCH(a:intention{name:'map_larger'}),(b:slot{name:'value'})CREATE(a)-[:relation]->(b);
[0119] Edge between slot and entity: MATCH(a:entity{name:'value'}),(b:entity_value{name:'1'})CREATE(a)-[:entity_value]->(b);
[0120] Secondly, initialize the thesaurus database. The system thesaurus is a prefix tree generated from the initials attributes of all graph nodes. Still taking "Enlarge the map" as an example, the relevant prefix tree fragment is shown below. Each node contains a tree structure with two parts: candidates (list of candidate words) and nodes (set of lower-level nodes). Starting from the root node and searching layer by layer downward, the range of the candidate word list will gradually decrease, and the actual retrieval level is less than 4. This thesaurus is used when interacting with the customized keyboard at the front end.
[0121]
[0122] Further, the recognition data pushed by the server includes the speech recognition result and the intent recognition result; taking "amplify by Figure 2 times" as an example, the intent recognition - post - processing module will receive the following format information:
[0123]
[0124] The intent recognition - post - processing module formats the information obtained by querying the initialization graph and ElasticSearch candidate words as follows:
[0125]
[0126]
[0127] This information is pushed to the front - end. The speech recognition list on the front - end displays the recognition text information (ASR field) in reverse chronological order, and the combined graph renders the intent graph through the returned point set (nodes field) and edge set (edges field) in an integrated cytoscape.
[0128] Further, the above steps are illustrated according to the following examples; Example 1 "amplify the map" is recognized correctly.
[0129] At this time, the word segmentation result is "amplify" and "map". The list of similar candidate words for "amplify" includes "enlarge", "increase", etc., and there is no candidate word list for "map"; the intent pre - processing module combines them into "amplify the map" and returns the implicit word slot "times", with the entity default, and splices it into the instruction "amplify the map to the default amplification factor"; the graph calculation module adds other candidate word lists of the current implicit entity, "2 times", "5 times", etc., and the associated candidate word list, including "reduce the map", "move the map", etc.; the engine status statistical calculation module obtains the current map status value. If it has not been amplified to the default amplification factor, the instruction is executed; if it has already reached the default amplification factor, the instruction is skipped; during this process, the user can also select other instructions through the front - end page for manual intervention.
[0130] Example 2 "home game" is recognized incorrectly, and the correct instruction is "load the game".
[0131] The word segmentation result is "home" and "game". The list of similar candidate words for "home" includes "load" (homophonic similarity) and has the highest similarity, and automatic replacement is completed; the candidate word list for "game" includes "scenario"; the intent pre - processing module combines them into "load the game"; the graph calculation module adds the candidate word list (associated words), other entity word lists such as "special effects animation", "timeline", etc. The engine status statistical calculation module skips the instruction if the current game has been loaded; during this process, the user can also select other instructions through the front - end page for manual intervention.
[0132] In Example 3, the recognition of "View M16 No. 1" was incorrect. The correct instruction is "View F16 Aircraft No. 1".
[0133] a. The word segmentation results are "View", "No. 1", "M16". The list of similar candidate words for No. 1 includes "No. 10", "No. 101", etc., and "M16" includes "F16", "M161", etc.
[0134] b. The intent preprocessing module combines them into multiple instructions, such as "View M16 No. 1", "View M16 No. 10", etc., and sorts them according to the combination similarity, eliminating the instruction "Query M161 No. 101" whose similarity is lower than the threshold.
[0135] c. The intent understanding service recognizes the intent as "View game equipment", the slot as "equipment name", and the entity value as "M16".
[0136] d. The graph calculation module hands over the queried point set and edge set to the front-end visualization module for graph rendering, and the user can see a combined graph composed of multiple similar instructions.
[0137] e. The engine status statistical calculation module verifies the instructions one by one in combination with the current game state, and eliminates illegal instructions (for example, there is no M16 No. 1 equipment in the game).
[0138] f. Automatically execute the corrected instruction.
[0139] Furthermore, when the user sends the corrected intent to the Web server through the front-end interface operation, at this time, the backtracking module stores the original instruction and the corrected instruction in the Mysql table, and the fields include:
[0140] `id` bigint(20) NOT NULL AUTO_INCREMENT COMMENT 'Index',
[0141] `text` varchar(1024) NOT NULL COMMENT 'Text recognized by asr',
[0142] `intent_model` varchar(1024) DEFAULT NULL COMMENT 'Intent recognized by the model',
[0143] `intent_select` varchar(1024) DEFAULT NULL COMMENT 'Corrected intent',
[0144] `slot_model` varchar(1024) DEFAULT NULL COMMENT 'Model-recognized slot entity',
[0145] `slot_select` varchar(1024) DEFAULT NULL COMMENT 'Corrected slot entity';
[0146] Meanwhile, map keywords for each part of the incorrect instruction and the correct instruction, and dynamically update the index constructor of ElasticSearch.
[0147] As shown in Example 2 of the error correction process, add the mapping between "home" and "load". In subsequent occurrences of the same error, the candidate word list obtained after retrieving ElasticSearch will include the "load" node.
[0148] Meanwhile, the formatted data stored in the above MySQL can be used for offline training of the intent understanding model and calculation of accuracy and recall rate to complete the optimization and evaluation of the upstream service.
[0149] On the other hand, as Figure 6 shown, an embodiment of the present invention further provides an error correction device for speech intent recognition, including:
[0150] An input module 1010, configured to input speech recognition information and convert the speech recognition information into text information;
[0151] A preprocessing module 1020, configured to preprocess the text information;
[0152] An intent recognition module 1030, configured to perform intent recognition on the preprocessed text information to obtain intent recognition information;
[0153] A combination module 1040, configured to form a combined graph of the intent recognition information based on the intent recognition information;
[0154] A transmission module 1050, configured to transmit the text information and the combined graph to a visual interaction interface;
[0155] An error correction module 1060, configured to correct the speech recognition information based on the visual interaction interface to obtain the corrected speech recognition information.
[0156] In one embodiment, the preprocessing module 1020 includes:
[0157] A word segmentation unit, configured to perform word segmentation on the text information to obtain a plurality of words;
[0158] A matching unit, configured to perform similarity matching between the plurality of words and preset words in a target database.
[0159] In one embodiment, the matching unit includes:
[0160] A first matching subunit, configured to, in response to successful similarity matching, record the node type of the vocabulary as successfully matched, and construct a candidate vocabulary list from preset vocabularies with a similarity greater than a first target value;
[0161] A second matching subunit, configured to, in response to failed similarity matching, replace the vocabulary with a preset vocabulary having a similarity greater than the first target value and the highest similarity, record the node type of the vocabulary as automatically replaced, and construct a candidate vocabulary list from other preset vocabularies;
[0162] A third matching subunit, configured to, in response to failed similarity matching and the absence of a preset vocabulary with a similarity greater than the first target value in the target database, record the node type of the vocabulary as unrecognized.
[0163] In one embodiment, the combination module 1040 includes:
[0164] A first combination unit, configured to combine multiple candidate vocabulary lists based on the multiple candidate vocabulary lists and the node types to form multiple combined statements;
[0165] A comparison unit, configured to compare the similarities of the multiple combined statements with the text information and sort them according to the similarities;
[0166] A removal unit, configured to remove combined statements with a similarity less than a second target value.
[0167] In one embodiment, the combination module 1040 further includes:
[0168] An intention understanding unit, configured to perform intention understanding on the multiple combined statements;
[0169] A second combination unit, configured to, based on the combined statements after intention understanding, and in combination with a point set, an edge set, and the similarities, combine and form a combined graph of intention recognition information; the point set represents semantic nodes corresponding to the speech recognition information, the intention, the word slot, and the entity, and the edge set represents the association relationship between different semantic nodes.
[0170] In one embodiment, the input module 1010 includes:
[0171] A judgment unit, configured to judge whether there is a legal instruction in the visual interaction interface, where the legal instruction represents information including an intention, a word slot, an entity, and meeting the requirements of a preset rule;
[0172] An input unit, configured to, in response to the absence of a legal instruction in the visual interaction interface, input a target node in the visual interaction interface;
[0173] A correction unit for correlating a target node with a combined graph to correct the combined graph.
[0174] In one embodiment, it further includes:
[0175] A backtracking module for generating an error correction instruction based on the error-corrected speech recognition information; performing keyword mapping on the speech recognition information in the error correction instruction and the speech recognition information in the original instruction to form a mapping relationship between the error correction instruction and the original instruction, where the original instruction is an instruction generated according to the input speech recognition information; storing the mapping relationship in a target database.
[0176] Figure 7 The structural schematic diagram of an embodiment of an electronic device provided by an embodiment of the present invention is shown. The specific implementation of the electronic device in the specific embodiment of the present invention is not limited.
[0177] As Figure 7 shown, the electronic device may include: a processor 502, a communication interface 504, a memory 506, and a communication bus 508.
[0178] Wherein: the processor 502, the communication interface 504, and the memory 506 communicate with each other through the communication bus 508. The communication interface 504 is used to communicate with network elements of other devices such as a client or other servers. The processor 502 is used to execute a program 510, and specifically can execute the relevant steps in the above-mentioned embodiment of the error correction method for speech intent recognition.
[0179] Specifically, the program 510 may include program code, and the program code includes computer-executable instructions.
[0180] The processor 502 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the electronic device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0181] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0182] Specifically, the program 510 can be called by the processor 502 to cause the electronic device to execute the relevant steps in the above-described embodiment of the error correction method for speech intent recognition.
[0183] Those of ordinary skill in the art can understand that Figure 7 the structure shown is only illustrative and does not limit the structure of the above device. For example, the electronic device may further include more or fewer components than those shown in Figure 7 or have a different configuration from that shown in Figure 7 the illustration.
[0184] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading through a network the original computer code stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the method described herein can be processed by such software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiment is implemented.
[0185] In the specific content of the above specific embodiment, the technical features can be combined arbitrarily without contradiction. For the sake of brevity of description, not all possible combinations of the above technical features are described. However, as long as the combinations of these technical features do not exist in contradiction, they should all be considered as the scope recorded in this specification.
[0186] The specific content of the above specific embodiment only expresses several embodiments of the present invention, and its description is relatively specific and detailed, but it cannot be understood as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A method for correcting speech intention recognition, characterized in that: The method comprises: Input voice recognition information, and convert the voice recognition information into text information; Preprocessing the text information; Performing intent recognition on the preprocessed text information to obtain intent recognition information; forming a combination map of intent recognition information based on the intent recognition information; Transmitting the text information and the combined graph to a visual interactive interface; The speech recognition information is corrected based on the visual interactive interface to obtain corrected speech recognition information.
2. The method for correcting speech intention recognition according to claim 1, characterized in that: The preprocessing of the text information includes: Performing word segmentation processing on the text information to obtain multiple words; The plurality of words are matched with preset words in the target database for similarity.
3. The method for correcting speech intention recognition according to claim 2, characterized in that: The similarity matching of the plurality of words with preset words in the target database includes: In response to a successful similarity match, recording the node type of the vocabulary as a successful match, and constructing the preset vocabulary with a similarity greater than the first target value as a candidate vocabulary list; In response to a similarity match failure, replacing the vocabulary with a preset vocabulary having a similarity greater than the first target value and the highest similarity, recording the node type of the vocabulary as automatic replacement, and constructing the other preset vocabulary into a candidate vocabulary list; In response to similarity matching failure and the target database not having a preset vocabulary with a similarity greater than the first target value, recording the node type of the vocabulary as unrecognized.
4. The method for correcting speech intention recognition according to claim 3, characterized in that: After the similarity matching of the plurality of the words with the preset words in the target database is performed, the method further includes: Based on the plurality of candidate vocabulary lists and the node types, combining the plurality of candidate vocabulary lists to form a plurality of combined sentences; Comparing the similarities between the plurality of combined sentences and the text information, and sorting them according to the similarities; The combined sentences whose similarity is less than the second target value are removed.
5. The method for correcting speech intention recognition according to claim 4, characterized in that: The forming a combination graph of intent recognition information based on the intent recognition information includes: Performing intention understanding on the plurality of combined sentences; Based on the combined sentences after the intention is understood, and in combination with the point set, edge set and the similarity, a combined graph of the intention recognition information is formed; the point set is represented as the semantic nodes of the intention, word slot and entity corresponding to the speech recognition information, and the edge set is represented as the association relationship between different semantic nodes.
6. The method for correcting speech intention recognition according to claim 1, characterized in that: The method further comprises: Determine whether there is a legal instruction in the visual interactive interface, where the legal instruction is characterized as information containing intent, word slot, entity and meeting the requirements of preset rules; In response to the absence of the legal instruction in the visual interaction interface, inputting a target node in the visual interaction interface; The target node is associated with the combined graph to modify the combined graph.
7. The method for correcting speech intention recognition according to claim 1, characterized in that: After obtaining the corrected speech recognition information, the method further includes: generating an error correction instruction based on the corrected speech recognition information; Perform keyword mapping of the speech recognition information in the error correction instruction with the speech recognition information in the original instruction to form a mapping relationship between the error correction instruction and the original instruction, wherein the original instruction is an instruction generated according to the input speech recognition information; The mapping relationship is stored in the target database.
8. A speech intention recognition error correction device, characterized in that: include: An input module, used for inputting speech recognition information and converting the speech recognition information into text information; A preprocessing module, used for preprocessing the text information; An intention recognition module, used to perform intention recognition on the preprocessed text information to obtain intention recognition information; A combination module, used for forming a combination map of intent recognition information based on the intent recognition information; A transmission module, used for transmitting the text information and the combined graph to a visual interactive interface; The error correction module is used to correct the speech recognition information based on the visual interactive interface to obtain the corrected speech recognition information.
9. The error correction device for speech intention recognition according to claim 8, characterized in that: Also includes: A backtracking module is used to generate a correction instruction based on the corrected speech recognition information; perform keyword mapping on the speech recognition information in the correction instruction and the speech recognition information in the original instruction to form a mapping relationship between the correction instruction and the original instruction, wherein the original instruction is an instruction generated according to the input speech recognition information; and store the mapping relationship in a target database.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the error correction method for speech intent recognition according to any one of claims 1 to 7 when executed.