An entity error correction method and intelligent device for multilingual semantic understanding
Through the combination of sound code algorithm and knowledge graph, the problem of lack of training data in the multilingual speech recognition system is solved, cross-language entity error correction is achieved, semantic understanding and entity recognition accuracy is improved, and user experience is improved.
Patent Information
- Application Number
- CN202210394592.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-14
AI Technical Summary
In the absence of large-scale training data, it is difficult for multilingual speech recognition systems to effectively correct the insensitive character order of short text entities, and the training data set does not match the application scenarios of voice assistants, resulting in low accuracy of semantic understanding and entity recognition and poor user experience.
The phonological code algorithm is used to encode error-correcting entities, find matching candidate entities in the phonological code database, and filter them using the knowledge graph to obtain the result entities, providing a unified multilingual semantic understanding framework.
In the absence of large-scale training data, across the influence of different languages, realizing entity error correction for texts in different languages, improving the accuracy of semantic understanding and entity recognition, and improving the performance and user experience of multilingual speech recognition products.
Smart Images

Figure CN114817465B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and particularly to an entity error correction method and an intelligent device for multilingual semantic understanding. Background Art
[0002] With the development of intelligent voice interaction technology, the voice interaction function has gradually become a standard configuration of intelligent terminal products. Users can use the voice interaction function to control intelligent terminal products by voice, and perform a series of operations such as watching videos, listening to music, checking the weather, and controlling the TV.
[0003] The process of controlling an intelligent terminal product by voice is usually that a voice recognition model recognizes the voice input by the user as text. Then, a semantic understanding model analyzes the text in terms of morphology, syntax, and semantics to understand the user's intention. Finally, the control end controls the intelligent terminal product to perform corresponding operations according to the understanding result.
[0004] In practical applications, both the voice recognition model and the semantic understanding model have errors, and the errors accumulated in the two stages directly affect the voice recognition quality of the entire system. Therefore, existing voice recognition systems are all configured with an error correction model to correct entities in the user request statement to improve the quality of the entire voice recognition system. Currently, multilingual error correction mainly focuses on the research of text grammar correction, and usually uses a large amount of text data for deep learning model training.
[0005] However, since the entities in multilingual voice recognition are usually short texts, these short text entities are insensitive to the character order and have no semantic information such as grammar, so it is difficult to correct errors based on text grammar. In addition, the training data sets used in current multilingual voice recognition do not match the application scenarios of voice assistants, and it is also difficult to collect a large amount of text data in real scenarios for model training. Currently, products for multilingual voice recognition cannot well meet the usage needs of users. Therefore, there is an urgent need for an entity error correction method that can cross the influence of different languages and is multilingual and universal in real time in the case of lack of large-scale training data. Summary of the Invention
[0006] This application provides an entity error correction method and an intelligent device for multilingual semantic understanding, which are used to solve the problem that since the entities in multilingual voice recognition are usually short texts, these short text entities are insensitive to the character order and have no semantic information such as grammar, so it is difficult to correct errors based on grammar. In addition, the training data sets used in current multilingual voice recognition do not match the application scenarios of voice assistants, and it is also difficult to collect a large amount of text data in real scenarios for model training.
[0007] In a first aspect, an embodiment of the present application provides an entity error correction method for multilingual semantic understanding. The method includes: obtaining an entity to be error-corrected, where the entity to be error-corrected is an entity obtained after semantic analysis and processing of a request statement input by a user;
[0008] Encoding the entity to be error-corrected using a phonetic shape code algorithm;
[0009] Searching for a candidate entity that matches the encoded entity to be error-corrected in a phonetic shape code database;
[0010] Filtering the candidate entity according to a knowledge graph to obtain a result entity, where the knowledge graph describes the association relationship between the candidate entities, and the result entity is the candidate entity that has an association relationship with other candidate entities.
[0011] In a second aspect, an embodiment of the present application provides an intelligent device for entity error correction in multilingual semantic understanding. The intelligent device includes:
[0012] An entity to be error-corrected acquisition unit for performing: obtaining an entity to be error-corrected, where the entity to be error-corrected is an entity obtained after semantic analysis and processing of a request statement input by a user;
[0013] An encoding unit for performing: encoding the entity to be error-corrected using a phonetic shape code algorithm;
[0014] A candidate entity search unit for performing: searching for a candidate entity that matches the encoded entity to be error-corrected in a phonetic shape code database;
[0015] A filtering unit for performing: filtering the candidate entity according to a knowledge graph to obtain a result entity, where the knowledge graph describes the association relationship between the candidate entities, and the result entity is the candidate entity that has an association relationship with other candidate entities.
[0016] The technical solution provided by the present application has the following beneficial effects: After obtaining the entity to be error-corrected, the entity to be error-corrected is encoded using a phonetic shape code algorithm. Search for a candidate entity that matches the encoded entity to be error-corrected in a phonetic shape code database. Filter the candidate entity according to a knowledge graph to obtain a result entity. Among them, the knowledge graph describes the association relationship between the candidate entities, and the result entity is the candidate entity that has an association relationship with other candidate entities. The entity error correction method and intelligent device for multilingual semantic understanding provided by the present application provide a unified framework for multilingual semantic understanding, can cross the influence of different languages in the case of lack of large-scale training data, realize entity error correction for texts in different languages, thereby improving the accuracy of semantic understanding and entity recognition, improving the performance of multilingual speech recognition products, and further improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0018] Figure 1 Exemplarily shows a schematic diagram of the principle of voice interaction according to some embodiments;
[0019] Figure 2 Exemplarily shows a schematic diagram of the framework of an entity error correction module according to some embodiments;
[0020] Figure 3 Exemplarily shows a schematic diagram of the process of an entity error correction method for multilingual semantic understanding according to some embodiments;
[0021] Figure 4 Exemplarily shows a schematic diagram of the encoding process of Chinese phonetic shape code according to some embodiments;
[0022] Figure 5 Exemplarily shows a schematic diagram of the specific example framework of an entity error correction method according to some embodiments;
[0023] Figure 6 Exemplarily shows a schematic diagram of the process of another entity error correction method according to some embodiments;
[0024] Figure 7 Exemplarily shows a schematic diagram of the process of another entity error correction method according to some embodiments;
[0025] Figure 8 Exemplarily shows a schematic diagram of a multilingual phonetic shape code encoding algorithm according to some embodiments;
[0026] Figure 9 Exemplarily shows a schematic diagram of the process of another entity error correction method according to some embodiments;
[0027] Figure 10 Exemplarily shows a schematic diagram of the framework of an intelligent device for entity error correction for multilingual semantic understanding according to some embodiments;
[0028] Figure 11 Exemplarily shows a flowchart of a multilingual voice assistant application according to some embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the objectives and implementation manners of this application clearer, the following will clearly and completely describe the exemplary implementation manners of this application in conjunction with the accompanying drawings in the exemplary embodiments of this application. Apparently, the described exemplary embodiments are only a part rather than all of the embodiments of this application.
[0030] It should be noted that the brief description of terms in this application is only for facilitating the understanding of the subsequent described implementation manners, rather than intending to limit the implementation manners of this application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0031] The terms "include" and "have" and any variations thereof are intended to cover but not be exclusive of inclusion. For example, a product or device including a series of components does not necessarily have to be limited to all the clearly listed components, but may include other components not clearly listed or inherent to these products or devices.
[0032] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to this element.
[0033] To clearly illustrate the embodiments of this application, the following will be combined with Figure 1 Describe a speech recognition network architecture provided by an embodiment of this application.
[0034] See Figure 1 , Figure 1 Is a schematic diagram of a speech recognition network architecture provided by an embodiment of this application. Figure 1 In, the intelligent device is used to receive the input information and output the processing result of this information. The speech recognition service device is an electronic device deployed with a speech recognition service, the semantic service device is an electronic device deployed with a semantic service, and the business service device is an electronic device deployed with a business service. The electronic device here may include a server, a computer, etc. The speech recognition service, semantic service (also called semantic engine), and business service here are web services that can be deployed on the electronic device. Among them, the speech recognition service is used to recognize audio as text, the semantic service is used to perform semantic parsing on the text, and the business service is used to provide specific services such as weather query service, music query service, etc. In one embodiment, Figure 1 In the shown architecture, there may be multiple entity service devices deployed with different business services, or one or more entity service devices may integrate one or more functional services.
[0035] In some embodiments, the following is based on Figure 1The process of the illustrated architecture for processing the information of the input intelligent device is described by way of example. Taking the information of the input intelligent device as a query statement input by voice as an example, the above process may include the following three processes:
[0036] Speech recognition: After receiving a query statement input by voice, the intelligent device can upload the audio of the query statement to the speech recognition service device, so that the speech recognition service device can recognize the audio as text through the speech recognition service and return it to the intelligent device. In one embodiment, before uploading the audio of the query statement to the speech recognition service device, the intelligent device can perform noise reduction processing on the audio of the query statement. Here, the noise reduction processing may include steps such as removing echo and ambient noise.
[0037] Semantic understanding: The intelligent device uploads the text of the query statement recognized by the speech recognition service to the semantic service device, so that the semantic service device can perform semantic parsing on the text through the semantic service to obtain the business field, intention, etc. of the text.
[0038] Semantic response: According to the semantic parsing result of the text of the query statement, the semantic service device issues a query instruction to the corresponding business service device to obtain the query result given by the business service. The intelligent device can obtain the query result from the semantic service device and output it. As an embodiment, the semantic service device can also send the semantic parsing result of the query statement to the intelligent device, so that the intelligent device can output the feedback statement in the semantic parsing result.
[0039] It should be noted that Figure 1 The illustrated architecture is only an example and does not limit the protection scope of the present application. In the embodiments of the present application, other architectures can also be used to implement similar functions. For example, all or part of the three processes can be completed by the intelligent device, which will not be elaborated here.
[0040] In some embodiments, Figure 1 The illustrated intelligent device may be a display device, such as a smart TV. The function of the speech recognition service device can be realized by the cooperation of a sound collector and a controller set on the display device. The functions of the semantic service device and the business service device can be realized by the controller of the display device or by the server of the display device.
[0041] In some embodiments, the intelligent device supports the voice interaction function, and the voice interaction function can be used as a standard configuration of the intelligent terminal product. Users can use the voice interaction function to realize voice control of the intelligent terminal product and perform a series of operations such as watching videos, listening to music, checking the weather, and controlling the TV.
[0042] The process of a voice-controlled intelligent terminal product is usually that the voice recognition module recognizes the voice input by the user as text. Then, the semantic analysis module analyzes the text in terms of morphology, syntax, and semantics to understand the user's intention. Finally, the control end controls the intelligent terminal product to perform corresponding operations according to the understanding result.
[0043] In practical applications, there are errors in both the voice recognition model and the semantic understanding model. The errors accumulated in the two stages directly affect the voice recognition quality of the entire system. Therefore, in some embodiments, the voice recognition system is configured with an error correction model to correct the entities in the user request statement to improve the quality of the entire voice recognition system.
[0044] For example, in Chinese, the user request "I want to watch the movie XX of Wu XX", where Wu XX is an entity. The error correction module of the voice system may correct "Wu XX" to better reach the user's actual intention and improve the user experience. It should be noted that the entity here refers to a named entity, which is an entity with a specific meaning, mainly including personal names, place names, organization names, proper nouns, etc.
[0045] In some embodiments, the error correction model is mainly based on two methods: the rule-based method and the machine learning-based method. The rule-based method refers to correcting homophonic errors, fuzzy sound errors, etc. according to language pronunciation rules and combining user usage habits. The machine learning-based method refers to constructing multiple decision models, training the models with a large amount of "error - corrected" text, and finally obtaining a unified error correction result from the decision models.
[0046] For the error correction method of multiple languages, it is impossible to simply reuse the error correction method in Chinese. The main problems are as follows: Chinese is a pictographic language, while most languages in multiple languages are alphabetic languages, and the logical meaning of alphabetic language characters is fundamentally different from that of Chinese; the pronunciation habits of different languages are different, and it is impossible to effectively construct a rule-based correction candidate set; a unified framework is needed to provide a reuse function for the recognition of different languages.
[0047] In some embodiments, the error correction of multiple languages mainly focuses on the error correction research of text grammar, and usually uses a large amount of text data for deep learning model training. However, since the entities in multiple language voice recognition are usually short texts, these short text entities are not sensitive to the character order and have no semantic information such as grammar, so it is difficult to correct errors based on grammar. In addition, the training data sets used in multiple language voice recognition do not match the application scenarios of voice assistants, and it is also difficult to collect a large amount of text data in real scenarios for model training. The products for multiple language voice recognition cannot well meet the user's usage needs. Therefore, there is an urgent need for a real-time entity error correction method that can cross the influence of different languages and is common for multiple languages in the case of lacking a large amount of training data.
[0048] To solve the above problems, the present application provides an entity error correction method for multilingual semantic understanding. This method provides a unified framework for multilingual semantic understanding, and can, in the case of a lack of large-scale training data, overcome the influence of different languages, realize entity error correction for texts in different languages, thereby improving the accuracy of semantic understanding and entity recognition, enhancing the performance of multilingual speech recognition products, and further enhancing the user experience.
[0049] Before elaborating on the method flow of the present application, the technical terms involved in the present application are explained:
[0050] The multilingual semantic understanding model uses the pre-trained model LaBSE (multilingual embedding vector model) of large-scale multilingual corpus data to encode and analyze the request text, and perform intent judgment and entity recognition on it. For example, if the user's request is "search for XX by Tom", after passing through the multilingual semantic parsing model, the output semantic parsing result is: "intent": "video.search", "actor": "Tom", "title": "XX", where the intent is the judged intent to search, and both actor and title are entities.
[0051] The Metaphone algorithm is a phonetic code algorithm that encodes according to the English pronunciation method of text. After being improved and perfected in recent years to Metaphone3. This algorithm is mainly a phonetic code encoding algorithm for English pronunciation rules, and its main purpose is to encode texts with similar pronunciations into the same key values. Using the pronunciation rules of English words, distinguishing between vowel letters and consonant letters, the consonant letters are encoded for the input text according to pre-set rules. For example, the consonant letters in the text "volume up" are "vlmp", so the entire text is encoded as "FLMP".
[0052] Although the Metaphone algorithm is developed for English, its basic idea is based on language pronunciation rules, and this algorithm can be extended to other alphabetic scripts similar to English. For example, the Metaphone algorithm for Brazilian Portuguese (pt-br) has become a solution for local city databases in Brazil, and scholars have also open-sourced Metaphone algorithms based on French, Spanish, Russian, etc. on GitHub (a hosting platform for open-source and private software projects).
[0053] This application mainly corrects the entities recognized by the semantic understanding model based on the entity error correction module in the multilingual speech assistant. Such as Figure 2Schematic diagram of the framework of the entity error correction module shown. The input of the entity error correction module is the parsing result of the user request statement by the multilingual semantic understanding model, and the parsing result includes the intent and the entity. The entity error correction module of this application mainly corrects the input entity. The entity error correction module mainly includes a phonetic and shape code recall sub-module and a knowledge graph retrieval sub-module. Based on Figure 2 the entity error correction module shown, as Figure 3 shown in the flowchart of the entity error correction method for multilingual semantic understanding. The method includes the following steps:
[0054] Step S101: Obtain the entity to be corrected, where the entity to be corrected is the entity obtained after semantic analysis and processing of the user input request statement.
[0055] The voice text is obtained by parsing the user input voice signal. Specifically, the user inputs a voice signal within the distance range where the terminal device receives the signal. The terminal device can collect the user input voice signal through a microphone, and then identify the voice text from the voice signal. In the embodiments of this application, the voice text can be recognized by a voice recognition server. The semantic server performs semantic analysis and processing on the voice text. It should be noted that when the semantic server performs semantic analysis and processing on the voice text, it can use the multilingual semantic understanding model mentioned above to parse the voice text to obtain the intent and the entity. The specific process of performing semantic analysis and processing on the voice text can adopt existing technologies, and this application does not elaborate on it in detail.
[0056] Step S102: Encode the entity to be corrected using the phonetic and shape code algorithm.
[0057] This step mainly uses the phonetic and shape code algorithm to encode the parsed entity, and the parsed entity is the entity to be corrected. The input of the intelligent voice assistant is the user voice, so the places where text errors occur only exist in the errors of voice recognition as text. Therefore, the pronunciation of the text is an important basis for error correction. This application encodes the entity to be corrected based on the existing Metaphone algorithm. As mentioned above, Metaphone is a phonetic and shape code algorithm for English. It encodes the text according to the English pronunciation rules and encodes the texts with similar pronunciations into the same phonetic and shape code. The Metaphone algorithm is designed for English and can also be extended to languages with similar phonetic writing (such as English, Spanish, Russian, etc.). Therefore, this application can use the Metaphone algorithm for different languages to encode the entity to be corrected.
[0058] In addition to using the Metaphone algorithm for phonetic writing, it is also possible to develop a phonetic and shape code algorithm for languages with pictographic writing based on pronunciation rules. For example, in Chinese, it is encoded based on the pinyin rules, and in Japanese, it is encoded through the fifty sounds. Figure 4This is an example of Chinese phonetic shape code encoding. For the text to be encoded, first convert the text into the corresponding pinyin through the PyPinyin tool, and encode the initials and finals in the pinyin respectively. During the encoding process, considering the similarity of pinyin pronunciations, for example, the initial n and the initial l, and the finals "an" and "ang" can be encoded with the same code.
[0059] Step S103, search for candidate entities in the phonetic shape code database that match the encoded entity to be corrected.
[0060] A large number of text entities are pre-stored in the phonetic shape code database of this application. The entities in the phonetic shape code database are also generated through the same encoding process, so the entities in the phonetic shape code database are also in encoded form. This step matches the encoded form of the entity to be corrected with the encoded form data in the phonetic shape code database to search for candidate entities that match the entity to be corrected. It should be noted that the matching of the entity to be corrected and the candidate entity can be that the encoding of the entity to be corrected is the same as the encoding of the candidate entity, or the encoding of the entity to be corrected contains the encoding of the candidate entity. This application does not limit the matching form.
[0061] Step S104, screen the candidate entities according to the knowledge graph to obtain result entities, where the knowledge graph describes the association relationships between various entities, and the result entities are the candidate entities that have an association relationship with other candidate entities.
[0062] This step is mainly to obtain the final error correction result. Specifically, the entity to be corrected and the candidate entities are put into the knowledge graph for query. The entities in the phonetic shape code database are basically the same as the entities in the knowledge graph. The knowledge graph refers to a knowledge base used to enhance the function of the search engine, aiming to describe various entities or concepts and their relationships existing in the real world. It constitutes a huge semantic network graph, where the nodes represent entities or concepts, and the edges are composed of attributes or relationships. All the data in the knowledge graph can be expressed and stored using RDF (Resource Description Framework).
[0063] Illustrate with examples. Represent the complex relationships between entities in the form of triples such as [entity, attribute, attribute value], [entity1, relationship, entity2], etc. For example, [XX (movie name), starring, Wu Moumou], [Yu Moumou, place of birth, Beijing], etc. It should be noted that in the embodiments of this application, the knowledge graph can also be represented and stored in other forms, and this application does not make specific limitations on this.
[0064] After screening the candidate entities, the result entities are obtained, and finally the valid intents can be output according to the result entities. Specifically, the results output by the knowledge graph retrieval are encapsulated into a unified format to facilitate the downstream terminal to execute commands. Entities of some intents also need to be converted into a unified format. For example, some fixed setting items of the TV need to be converted into the TV language; time formats in different languages need to be converted into digital formats that the terminal can parse, such as "2 hours" in English is converted into "{h: 2, m: 0, s: 0}".
[0065] For the entity error correction method for steps S101 to S014, it is described through specific examples as follows: Figure 5 shown below:
[0066] First, obtain the entities to be error-corrected [Title: Barnaby] and [Actor: David]. Use the phonetic and shape code algorithm to encode the above two entities to be error-corrected, and obtain the encodings [Title: FTPRNPRJ] and [Actor: TFTKRLN] respectively. Search for candidate entities in the phonetic and shape code database. Specifically, recall entities with similar encodings in the phonetic and shape code database: [Title: Barnaby, Bridge, Barrage] and [Actor: David]. At this time, three Title entities and one Actor entity are recalled. Then, screen all the recalled entities according to the knowledge graph.
[0067] Specifically, the knowledge graph is a knowledge base that describes various entities or concepts existing in the real world and their relationships. The movie starring actor David is Barrage, and its representation in the knowledge graph is [David, main character (starring), Barrage]. Therefore, at this time, after screening by the knowledge graph, the result entities [Title: Barrage] and [Actor: David] are obtained. Correct the entity to be error-corrected [Title: Barnaby] to the result entity [Title: Barrage]. Finally, according to the obtained result entities, output the valid intent: Search for the movie Barrage starring actor David.
[0068] The entity error correction method for multi-language semantic understanding provided by this application provides a unified framework for multi-language semantic understanding. It can, in the case of lacking large-scale training data, overcome the influence of different languages, realize entity error correction for texts in different languages, thereby improving the accuracy of semantic understanding and entity recognition, enhancing the performance of multi-language speech recognition products, and further enhancing the user experience.
[0069] In some embodiments, corresponding phonetic shape code databases can be set in advance for different languages. Before searching for candidate entities that match the entity to be corrected in the phonetic shape code database, first determine the language type of the entity to be corrected, and then search in the corresponding phonetic shape code database. For example, for the entity to be corrected "spider", first determine that the language type of this entity to be corrected is English. Therefore, candidate entities that match this entity to be corrected can be searched in the English phonetic shape code database. By distinguishing different languages, the search efficiency can be improved when recalling candidate entities according to the phonetic shape code database.
[0070] In some embodiments, corresponding phonetic shape code databases can be set in advance for different services. Before searching for candidate entities that match the entity to be corrected in the phonetic shape code database, first determine the service type of the entity to be corrected, and then search in the corresponding phonetic shape code database. In specific service requirements, usually only some types of entities need to be corrected. For example, in the TV voice assistant service, only video names, person names, and channel names need to be corrected.
[0071] As Figure 6 shown in the example, when constructing a phonetic shape code database related to the actual service, different types of collected entities are respectively encoded using the phonetic shape codes of a single language, and the encoded results are used as key values, and the text of the original entity is used as the value and stored in the phonetic shape code database S. Among them, for example, Schannel represents the phonetic shape code database related to channels, Stitle represents the phonetic shape code database related to video names, and Sactor represents the phonetic shape code database related to person names.
[0072] For the entity to be corrected "text", recall candidate entities E with similar pronunciations in the following way:
[0073] E phoneric = S(F phonetic (text))
[0074] For example, when the entity to be corrected "text" is "Barnaby", its entity type is "title". After the key value encoded by the phonetic shape code algorithm F is "FTPRNPRJ", recall in the database Stitle. Stitle(FTPRNPRJ) can recall three entities {"Barnaby", "Barrage", "Bridge"} stored in the knowledge base, and use these three recalled entities as the candidate set for subsequent error correction. In this way, by dividing different phonetic shape code databases according to different service types, the search efficiency can be improved when recalling candidate entities according to the phonetic shape code database.
[0075] In the actual application of a multilingual recognition system, the candidate entities recalled by the above method may be excessive in number or zero in number. It is necessary to post-process the recalled candidate entities for error correction to improve the quality of the system. For a large multilingual recognition system, the number of candidate entities collected is in the tens of millions. When recalling candidate entities similar in pronunciation to the entity to be error-corrected from the phonetic and shape code database, a huge number of candidate entities may be recalled. Excessive candidate entities will bring a burden on calculation and query for subsequent error correction, wasting computing resources and time. Therefore, it is necessary to clean the number of recalled candidate entities.
[0076] Specifically, if the number of candidate entities recalled from the phonetic and shape code database is greater than the number threshold N, a sorting method based on the edit distance is used to clean the candidate entities. The edit distance (Minimum Edit Distance, MED) refers to the minimum number of times required to transform one text into another text, and is used to measure the difference degree between two texts. When the number of recalled candidate entities exceeds the pre-set threshold N, calculate the edit distance between each entity in the candidate entities and the entity to be error-corrected, sort according to the edit distance, and select the TopN candidate entities with smaller edit distances for subsequent error correction.
[0077] In some embodiments, if the number of candidate entities recalled from the phonetic and shape code database is zero, a substring is intercepted from the entity to be error-corrected. Then, the substring is encoded using the phonetic and shape code algorithm, and finally, candidate entities matching the encoded substring are searched for in the phonetic and shape code database.
[0078] Exemplarily, a multilingual semantic understanding model may recognize the title entity in "search for film Red Dog" as "film Red Dog". Since the entity to be error-corrected contains the code of "film" during encoding, the number of candidate entities recalled for this entity to be error-corrected in the phonetic and shape code database is zero. At this time, the substring "Red Dog" of the entity to be error-corrected "film Red Dog" can be intercepted. After encoding the substring "Red Dog", it is retrieved and recalled from the database S. This re-retrieval in the database through substring encoding is implemented recursively to find the longest substring that can match the result, and its candidate result is used as the final phonetic and shape code recall result.
[0079] The above embodiments are all cases where only one language is included in a user request statement. In actual applications, there may also be cases where a user request statement contains multiple languages. There are mainly two reasons for this situation: there is a certain regional correlation among users of different languages. For example, there is a regional correlation between Chinese and Japanese users. Among Malay users, because there are many overseas Chinese, they will also carry some Chinese in their requests; the wide use of English and the richness of media resources lead to many users entering business requests with English when the language of the voice assistant is not English.
[0080] For example, French users may have a request like "rechercher le (French) film spider (English)". When the existing Automatic Speech Recognition (ASR) model faces such a situation of mixed languages, the accuracy rate will decrease, and it will probabilistically mis-recognize the pronunciation of non-native languages as texts in the native language. For example, the request voice of a Malay user "buka sofa" will be mis-recognized as the text "buka soufa" by the ASR model. The recall of single-language phonetic and graphic codes is not applicable to this type of scenario, and the accuracy rate of the multi-language voice assistant system in the face of such problems is also relatively low.
[0081] To solve the above problems, in this embodiment, the entity error correction method is improved based on the above embodiments. Specifically, when no valid entity (an entity that can generate a valid intention) is recalled from the phonetic and graphic code database after encoding the entity to be corrected, the phonetic and graphic code algorithm of the candidate other language is used for secondary recall. There are two ways to determine the candidate language here: determining the possible language type of the text based on the characters of the text; speculating other languages that the users of this language may use based on prior knowledge.
[0082] Exemplarily, as Figure 7 shown in the schematic diagram of the multi-language phonetic and graphic code recall process. When the French user's request statement is "rechercher le film man", the recognized named entity text is "man". The candidate set E after recalling the entity text "man" through the French phonetic and graphic code is [mania]. The recalled entity text "mania" cannot generate a valid intention. Then the phonetic and graphic code algorithm F of other languages is used for secondary recall.
[0083] First, based on the analysis of the characters of the entity text, it is determined that "man" may belong to the English language. At the same time, based on the language habits of French users, it is speculated based on prior knowledge that the neighboring languages of French are languages such as English, Spanish, Italian, and Portuguese (Spanish, Italian, Portuguese, English, etc. belong to the same language family as French and there is also a regional correlation). Combining the information from both aspects, it is concluded that the language that can be used for secondary recall is English. Using the English phonetic and graphic code algorithm Fen( Figure 7The candidate entity E obtained after the second recall is [man]. At this time, a valid intent (for example, searching for the movie man) can be generated based on the recalled entity text man.
[0084] The improved method of the above embodiment is fundamentally a way of combining the phonetic and graphic coding algorithms of different languages. However, when the ASR model faces mixed requests in multiple languages, the errors in its output text are diverse and uncertain. Therefore, the improved method of the above embodiment cannot be applied to all situations of ASR recognition errors.
[0085] In order to solve the above problems, a further improved method is to consider the pronunciation problem of multiple languages when designing a single language phonetic coding algorithm, that is, to make the texts in different languages with similar pronunciations similarly encoded after being encoded in their respective languages. For example, the encoding "soufa" in English and the encoding "sofa" in Chinese are similarly encoded. When recalling candidate entities, similarity calculation is used to recall candidate entities with similar pronunciations in other languages. Large quantities of text data with similar pronunciations in different languages can be collected to train the semantic understanding model, and the semantic understanding model can be a LABSE model.
[0086] The network structure of the encoding model can adopt the Transformer model, and the contrastive learning method is used so that after the text pairs with similar pronunciations in different languages are encoded by the model, the feature vectors of the two are close in the encoding space, while the feature vectors of the text pairs with large pronunciation differences in different languages are far apart in the encoding space. For example, the feature vector of the first entity in English and the feature vector of the second entity in Chinese exist in the encoding space. If the pronunciations of the first entity and the second entity are similar, the feature vectors of the two are close in the encoding space. If the pronunciations of the first entity and the second entity are very different, the feature vectors of the two are far apart in the encoding space.
[0087] An example deep learning network is Figure 8 As shown, the entity text "soufa" in English and the entity text "sofa" in Chinese have similar pronunciations. When encoded by their respective encoding algorithms, they can share encoding parameters so that the obtained encoded feature vectors are closer in the encoding space. When searching for candidate entities, if the screened candidate entity "soufa" cannot output a valid intent, the candidate entity "soufa" can be directly replaced with the candidate entity "sofa" which is closer. This can not only further improve the accuracy of semantic understanding and entity recognition, but also shorten the system calculation time and improve the efficiency of semantic understanding.
[0088] In some embodiments, when screening candidate entities according to a knowledge graph, if there is a pre-stored entity in the knowledge graph that matches the candidate entity, the pre-stored entity is determined as the result entity. If there is no pre-stored entity in the knowledge graph that matches the candidate entity, but there is a pre-stored entity that matches a sub-character in the candidate entity, the pre-stored entity is determined as the result entity. It should be noted that the match between the pre-stored entity and the candidate entity can be exactly the same or have an inclusion relationship, and the present application does not limit this.
[0089] Exemplarily, if the candidate entity is "film man", and there is no pre-stored entity in the knowledge graph that matches the candidate entity, but there is a pre-stored entity that matches the sub-character "man", then the pre-stored entity "man" is determined as the final result entity.
[0090] In some embodiments, the association relationships between candidate entities can also be queried to eliminate inappropriate candidate entities. Exemplarily, the English request "search for man by Tom" will output two types of entities after passing through a semantic understanding model: the movie name is "man"; the actor name is "Tom". This type of entity may recall multiple candidate movie names and actor names after phonetic and shape code encoding. For example, "man" may recall multiple movies related to "man" but not related to the actor "Tom". Through knowledge graph query, candidate entities that are associated with both the entity "man" and the entity "Tom" are filtered out, that is, only entities that are related to both "man" and "Tom" are retained.
[0091] In some embodiments, if there are still multiple result entities after the screening process in the above embodiments, multiple valid intents will be generated, which is not conducive to responding to user requests. Therefore, TF-IDF (Term Frequency-Inverse Document Frequency) can be used to score and rank multiple result entities, and the one with the highest score is selected as the final result entity. Specifically, other knowledge attributes of the entity can be referred to, such as the rating of a movie, the year and popularity of a movie. Entities with a newer year and higher popularity have a higher score. This can not only effectively respond to user requests, but also make the final response result more in line with the user's expectations to a greater extent.
[0092] In some embodiments, an instruction set to be matched can be constructed for common instructions of a voice assistant. When the user's request meets a preset condition (such as the length being less than a preset length), the original request text of the user is directly encoded with a phonetic and shape code, and then matched with the instruction set to be matched. If there is a matching result, the pre-stored semantic result is directly given to the downstream task. At this time, the constructed instruction set to be matched is the phonetic and shape code database. It should be noted that the matching here can be that the two are exactly the same, or there is an inclusion relationship between the two, and this application does not limit this.
[0093] Exemplarily, as Figure 9 shown in the specific embodiment, when the user's request in English is "pose", according to the experience of the multilingual voice assistant, the user's request should be "pause", but the speech-to-text recognition is incorrect and it is recognized as "pose". At this time, the multilingual semantic understanding model cannot effectively determine the intention of "pose" and cannot issue a command to the terminal. After encoding "pose" with the phonetic and shape code algorithm, the encoding "PS" is obtained. Then, it is retrieved in the instruction set to be matched, and "pause" with the same encoding "PS" is obtained. Then, it can be further screened in the knowledge graph. In this embodiment, since the candidate entity is found in the instruction set to be matched, it can also not be screened in the knowledge graph, and directly use "pause" as the result entity. Further retrieve in the instruction set to obtain the instruction "control.pause" associated with "pause", and send this instruction to the terminal for execution. This can not only further improve the accuracy of semantic understanding and entity recognition, but also shorten the system calculation time and improve the efficiency of semantic understanding.
[0094] An embodiment of the present application provides an intelligent device for entity error correction in multilingual semantic understanding, which is used to execute Figure 2 the corresponding embodiment, as Figure 10 shown, the intelligent device provided by the present application at least includes:
[0095] An entity to be corrected acquisition unit U1001, which is used to execute: acquire an entity to be corrected, where the entity to be corrected is an entity obtained after semantic analysis and processing of a request statement input by a user;
[0096] An encoding unit U1002, which is used to execute: encode the entity to be corrected by using a phonetic and shape code algorithm;
[0097] A candidate entity search unit U1003, which is used to execute: search for a candidate entity that matches the encoded entity to be corrected in a phonetic and shape code database;
[0098] A screening unit U1004 for performing: screening the candidate entities according to a knowledge graph to obtain result entities, where the knowledge graph describes the association relationships between the candidate entities, and the result entities are the candidate entities having association relationships with other candidate entities.
[0099] Based on the entity error correction method of the above embodiments, the present application further provides a multilingual voice assistant for multilingual semantic understanding, as Figure 11 shown in the application flow chart of the multilingual voice assistant, where the content in the dashed box is the entity error correction method provided by the present application. After the speech-to-text model receives the user-requested speech, it converts the user-requested speech into text, and then inputs it into the multilingual semantic understanding model to output an intention, which includes entities.
[0100] If the output intention needs to perform entity error correction, encode the entity using the phonetic and shape code algorithm, and recall candidate entities in the phonetic and shape code data path. Then, through knowledge graph query, obtain the result entities. Finally, encapsulate the knowledge graph retrieval output into a unified format for downstream terminals to execute commands conveniently.
[0101] If the output intention does not need to perform entity error correction, perform phonetic and shape code matching between the entity text and the text in the instruction set to be matched. If the entity text and the text in the instruction set to be matched have phonetic and shape code matching, then encapsulate the text in the instruction set to be matched into a unified format and output it. If the entity text and the text in the instruction set to be matched do not have phonetic and shape code matching, then encapsulate the entity text into a unified format and output it.
[0102] Examples of implementations of the present invention have been described above. For the purpose of describing the claimed subject matter, of course, it is not possible to describe every conceivable combination of components or methods, but it should be realized that many other combinations and permutations of this innovation are possible. Accordingly, the claimed subject matter is intended to embrace all such changes, modifications, and variations that fall within the spirit and scope of the appended claims. In addition, the above description of the illustrated implementations of the present application, including what is described in the "Abstract", is not intended to detail or limit the disclosed implementations to the precise forms disclosed. Although specific implementations and examples are described in the present application for illustrative purposes, as those skilled in the relevant art can recognize, various modifications considered to be within the scope of such implementations and examples are possible.
[0103] In addition, the terms "example" or "exemplary" are used in this application to mean "serving as an example, instance, or illustration". Any aspect or design described as "exemplary" in this application is not necessarily to be understood as being preferred or advantageous relative to other aspects or designs. Instead, the use of the terms "example" or "exemplary" is intended to present concepts in a concrete manner.
Claims
1. An entity error correction method for multilingual semantic understanding, characterized in that, Including: Obtain multiple types of entities to be corrected, where the entities to be corrected are entities obtained after semantic analysis of a request statement input by a user; Encode the multiple types of entities to be corrected using a phonetic and shape code algorithm; Obtain the language type of the multiple types of entities to be corrected; call a phonetic and shape code database according to the language type; Search for candidate entities that match the encoded multiple types of entities to be corrected in the phonetic and shape code database; where at least one candidate entity is matched for each type of entity to be corrected; when the number of candidate entities recalled from the phonetic and shape code database based on the encoded entity to be corrected is zero, intercept a substring from the entity to be corrected; encode the substring using a phonetic and shape code algorithm, and search for candidate entities that match the encoded substring in the phonetic and shape code database; Filter the candidate entities according to a knowledge graph to obtain result entities, where the knowledge graph describes the association relationships between the candidate entities, and the result entities are candidate entities that have association relationships among at least one candidate entity that matches other types of entities to be corrected.
2. The entity error correction method for multilingual semantic understanding according to claim 1, characterized in that, Search for candidate entities that match the encoded entity to be corrected in the phonetic and shape code database. The specific steps are as follows: Obtain the business type of the entity to be corrected; Call the phonetic and shape code database according to the business type of the entity to be corrected; Search for candidate entities that match the encoded entity to be corrected in the phonetic and shape code database.
3. The entity error correction method for multilingual semantic understanding according to claim 1, wherein Search for candidate entities that match the encoded entity to be corrected in the phonetic and shape code database. The specific steps are as follows: When the number of candidate entities recalled from the phonetic and shape code database based on the encoded entity to be corrected is greater than the quantity threshold N, calculate the edit distance between all the candidate entities and the entity to be corrected; Sort all the candidate entities according to the edit distance, and determine the first N candidate entities sorted in ascending order of the edit distance as the final candidate entities.
4. The entity error correction method for multilingual semantic understanding according to claim 1, characterized in that, After an entity is encoded, it has a feature vector. The distance between the feature vectors of a first entity and a second entity in the encoding space matches the degree of pronunciation similarity between the first entity and the second entity, where the language types of the first entity and the second entity are different.
5. The entity error correction method for multilingual semantic understanding according to claim 1, wherein, Filter the candidate entities according to a knowledge graph to obtain result entities. The specific steps are as follows: When there is a pre-stored entity that matches the candidate entity in the knowledge graph, determine the pre-stored entity as the result entity; When there is no pre-stored entity that matches the candidate entity in the knowledge graph, and there is a pre-stored entity that matches a sub-character in the candidate entity in the knowledge graph, determine the pre-stored entity as the result entity.
6. The entity error correction method for multilingual semantic understanding according to claim 1, characterized in that, The phonetic and shape code database is a preset common instruction set, and the candidate entities are entities used to generate common instructions.
7. An intelligent device for entity error correction in multilingual semantic understanding, characterized in that, Including: An entity acquisition unit to be corrected, configured to execute: Obtain multiple types of entities to be corrected, where the entities to be corrected are entities obtained after semantic analysis of a request statement input by a user; Coding unit, for performing: encoding the multiple types of entities to be error-corrected by using a phonetic and shape code algorithm; Candidate entity search unit, for performing: obtaining the language type of the multiple types of entities to be error-corrected; calling a phonetic and shape code database according to the language type; searching in the phonetic and shape code database for candidate entities that match the encoded multiple types of entities to be error-corrected; wherein, at least one candidate entity is matched to each type of entity to be error-corrected; when the number of candidate entities recalled from the phonetic and shape code database according to the encoded entity to be error-corrected is zero, intercepting a substring from the entity to be error-corrected; encoding the substring by using a phonetic and shape code algorithm, and searching in the phonetic and shape code database for candidate entities that match the encoded substring; Filtering unit, for performing: filtering the candidate entities according to a knowledge graph to obtain result entities, where the knowledge graph describes the association relationships between the candidate entities, and the result entities are candidate entities having association relationships among at least one candidate entity that matches other types of entities to be error-corrected.
Citation Information
Patent Citations
Semantic error correction method, electronic equipment and storage medium
CN111291571A
Online video course content summary generation method
CN113343026A