Intention recognition method and device, computer processing device, and storage medium
By extracting and matching entity paths from the intent knowledge graph, the problems of inaccuracy and high cost in existing intent recognition methods are solved, achieving efficient and accurate user intent recognition.
Patent Information
- Application Number
- CN202210919268.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-01
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2042-08-01
AI Technical Summary
Existing intent recognition methods cannot accurately identify user intent and require a large amount of manually labeled training data, resulting in high costs, low efficiency, and poor user experience.
By extracting the entity set of the text to be identified from a pre-built intent knowledge graph, determining the entity path associated with the entity set, and matching it to determine the intent, manual annotation is reduced and recognition efficiency is improved.
It achieves accurate recognition of user intent, reduces the need for manual annotation, lowers training costs, and improves the efficiency of intent recognition and user experience.
Smart Images

Figure CN116126994B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an intent recognition method, apparatus, computer processing device, and storage medium. Background Technology
[0002] Intent recognition is an important part of human-computer interaction systems. It transforms the content of user dialogue into a way that computers can understand, analyzes user needs, and provides relevant content that meets those needs back to the user. Through intent recognition, users can find the intent information they want. Therefore, intent recognition is widely used in various industries and application scenarios.
[0003] However, most existing intent recognition methods cannot accurately identify user intent, resulting in a poor user experience. Furthermore, they require a large amount of manually annotated training data to train the model, which is very labor-intensive, has high training costs, poor inference performance, and affects user experience. Summary of the Invention
[0004] The main technical problem addressed by this application is to provide an intent recognition method, apparatus, device, and storage medium that can avoid the problem of low efficiency in intent knowledge graph construction caused by manual annotation.
[0005] To solve the above-mentioned technical problems, one technical solution adopted in this application is: to provide an intent recognition method, the method comprising: extracting a first entity set from the text to be recognized;
[0006] Based on the first entity set, determine the second entity set associated with the first entity set in the pre-constructed intent knowledge graph;
[0007] Based on the second entity set, at least one entity path corresponding to the second entity set is determined in the intent knowledge graph; wherein, the entity path includes at least two entities with an association relationship;
[0008] The first entity set is matched with at least one entity path to obtain the target entity path, and the intent of the text to be identified is determined based on the target entity path.
[0009] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide an intent recognition device, the intent recognition device comprising:
[0010] The extraction module is used to extract the first entity set from the text to be recognized.
[0011] The retrieval module is used to determine a second entity set associated with the first entity set in a pre-constructed intent knowledge graph, based on the first entity set; wherein the intent knowledge graph is built based on a corpus and is used to represent the relationships between entities.
[0012] An entity path determination module is used to determine at least one entity path corresponding to the second entity set in the intent knowledge graph, based on the second entity set; wherein the entity path includes at least two entities with an association relationship;
[0013] A matching module is used to match a first entity set with at least one entity path to obtain a target entity path, and to determine the intent of the text to be identified based on the target entity path.
[0014] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium that stores executable instructions, which are executed by a processor as described above in the intent recognition method.
[0015] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide a computer processing device, the computer processing device including a processor; and a memory arranged to store computer executable instructions, the executable instructions being configured to be executed by the processor, the executable instructions including steps for performing the above-described intent recognition method.
[0016] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium that stores program data, which, when executed by a processor, is used to implement the intent recognition method as described above.
[0017] The beneficial effects of this application are as follows: by identifying a second entity set associated with the first entity set in a pre-constructed intent knowledge graph, the matching range between the first entity set and the intent knowledge graph is narrowed. Then, based on the second entity set, the entity path corresponding to the second entity set is determined, and the first entity set is matched with the entity path. By matching the first entity set with the entity path, for example, similarity matching can be used to filter out target entity paths that are more closely related to the text to be identified. Then, based on the target entity path, the intent of the text to be identified can be determined, and the intent that the user wants to express can be accurately obtained. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0019] Figure 1 This is a flowchart illustrating the first embodiment of the intent recognition method provided in this application;
[0020] Figure 2 This is a schematic diagram of the structure of a subgraph in the intent knowledge graph provided in this application;
[0021] Figure 3 This is a schematic diagram of the structure of a subgraph in the intent knowledge graph provided in this application;
[0022] Figure 4 This is a schematic diagram of the structure of a subgraph in the intent knowledge graph provided in this application;
[0023] Figure 5 This is a schematic diagram of the structure of a subgraph in the intent knowledge graph provided in this application;
[0024] Figure 6 This is a schematic diagram of the structure of an embodiment of the intent recognition device provided in this application;
[0025] Figure 7 This is a schematic diagram of the structure of an embodiment of the computer processing device provided in this application;
[0026] Figure 8 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0028] The following explains the technical terms and abbreviations involved in this invention.
[0029] Knowledge Graph: A semantic network designed to describe conceptual entities in the objective world and the relationships between them; sometimes also called a knowledge base.
[0030] Traditional knowledge graphs have the concepts of nodes and edges.
[0031] Among them, a node represents an information entity or the attribute value of an entity.
[0032] Edge: can represent the relationship between two connected entities or a certain attribute of an entity.
[0033] Entity: An entity is the basic unit of a knowledge graph and an important linguistic unit that carries information in text.
[0034] Generally, a knowledge graph can be composed of triples: <entity1, relation, entity2> or <entity, attribute, attribute value>.
[0035] For example: <Celebrity, plays-in, NBA>, <Celebrity, height, 2.29m>.
[0036] Mention: A linguistic fragment in natural text that expresses an entity.
[0037] Please see Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the intent recognition method provided in this application. The method includes:
[0038] Step 110: Extract the first entity set from the text to be recognized.
[0039] In step 110, the text to be recognized can be a user-inputted statement or voice content, and the first entity set can be the entity set obtained after entity recognition, i.e., the entity Mention set.
[0040] For example, assuming a user inputs a question like "conditions for transferring Shenzhen household registration", the entity Mention set obtained after entity recognition can be {"Shenzhen", "household registration", "transfer", "conditions"}, that is, {"Shenzhen", "household registration", "transfer", "conditions"} can be used as the first entity set.
[0041] In some embodiments, the text to be recognized can be first input into a trained entity recognition model to obtain the entities of the text to be recognized, and then a first entity set can be constructed based on the obtained entities. The entities in the text to be recognized can be multiple semantic information contained in reasoning based on the input user question.
[0042] In some embodiments, entity recognition models can be obtained through a training phase using traditional machine learning or deep learning methods.
[0043] For example, assuming the input recognition text is the question "Conditions for transferring Shenzhen household registration", the input to the trained entity recognition model will result in the entities in the text to be recognized being Shenzhen, household registration, transfer, and conditions. The first entity set constructed based on these entities can be {"Shenzhen", "household registration", "transfer", "conditions"} or {"Shenzhen", "household registration", "transfer"} or {"Shenzhen", "household registration"}.
[0044] Step 120: Based on the first entity set, determine the second entity set associated with the first entity set in the pre-constructed intent knowledge graph.
[0045] For step 120: You can first collect a corpus, and then pre-construct an intent knowledge graph based on the collected corpus.
[0046] For example, the corpus collected regarding household registration transfer could be as follows:
[0047] Conditions for transferring Shenzhen household registration;
[0048] Procedures for transferring a Shenzhen household registration within the city;
[0049] Conditions for relocating from Shenzhen to Beijing;
[0050] Conditions for relocating from Beijing to Shenzhen;
[0051] Procedures for transferring within Shenzhen;
[0052] Materials transferred within Shenzhen.
[0053] The intent knowledge graph constructed based on the listed household registration transfer corpus can include Figures 2-3 The sub-map shown.
[0054] In addition, the collected corpus can also be:
[0055] The conditions for applying for a credit loan at a certain bank;
[0056] Risks of credit loans;
[0057] What procedures are required to apply for a credit loan at a certain bank?
[0058] What are the current policies regarding credit loans?
[0059] What are the repayment methods for personal loans?
[0060] Intent knowledge graphs constructed from corpora of credit loan data can, for example, Figure 4 As shown.
[0061] It should be noted that the collected corpus can include content related to various industries, as well as topics such as entertainment, people, and food. Therefore, the intent knowledge graph pre-constructed based on the collected corpus can encompass relevant knowledge from various industries, such as the aforementioned... Figures 2 to 4 It can be one of the subgraphs in the intent knowledge graph.
[0062] Assuming the first entity set is {“Shenzhen City”, “Household Registration”, “Transfer”, “Condition”} mentioned in step 110, then the second entity set can be a pre-constructed intent knowledge graph centered on household registration, and may include... Figure 2 and Figure 3 .
[0063] In some embodiments, when collecting a corpus, clustering or other methods can be used in the corpus samples to determine the intents contained in the corpus, and an intent knowledge graph can be constructed through similar intents.
[0064] In some embodiments, after constructing an intent knowledge graph based on the collected corpus, when user-inputted text is received, semantic information existing in the corpus can be annotated using template regularization matching. For example, if the input text is "Shenzhen household registration transfer conditions", the semantic information existing in the corpus can be annotated based on the terms "household registration" and "transfer conditions" in the regularization matching samples.
[0065] Since the pre-constructed intent knowledge graph may encompass relevant knowledge from various industries and include a large number of sub-graphs, in order to facilitate the selection of a second entity set related to the first entity set in the text to be identified, in some embodiments, the head entity in the first entity set will be determined first. Specifically, the head entity in the first entity set can be determined based on the dependency relationship between each entity in the first entity set in the text to be identified.
[0066] For example, the first entity set is {"Shenzhen City", "Household Registration", "Transfer", "Condition"} mentioned in step 110. Here, Shenzhen City is the location, Household Registration is the subject, Transfer is the action performed on the subject Household Registration, Condition can be said to be the use attribute of the action Transfer, and Shenzhen City can be said to be the location attribute of the action Transfer. Therefore, based on the dependency relationship between the four entities "Shenzhen City", "Household Registration", "Transfer", and "Condition", the head entity of the first entity set can be determined to be Household Registration.
[0067] Once the head entities in the first entity set are determined, in some embodiments, a second entity set can be selected based on the determined head entities in the first entity set. Specifically, based on the determined head entities and the first entity set, a second entity set associated with the first entity set can be filtered in the intent knowledge graph.
[0068] For example, if the head entity in the first entity set is "household registration", then in the intent knowledge graph, subgraphs related to "household registration", "household registration", or "household registration book" can be filtered out. These subgraphs related to "household registration", "household registration", or "household registration book" can be used as the second entity set associated with the first entity set.
[0069] The second entity set, with household registration as the head entity, can be as follows: Figure 2 As shown, or as Figure 3 As shown, it can also be based on Figure 2 and Figure 3 The content was further integrated to obtain similar results. Figure 5 A schematic diagram.
[0070] Since the preset intent knowledge graph encompasses many aspects, in order to narrow down the scope of filtering, in some embodiments, a candidate entity set associated with the first entity set will be filtered first. The specific approach is as follows:
[0071] In the intent knowledge graph, identify the candidate entity set associated with the first entity set;
[0072] Taking the first entity set as {“Shenzhen City”, “Household Registration”, “Transfer”, “Conditions”} mentioned in step 110 as an example, the entities “Shenzhen City”, “Household Registration”, “Transfer”, and “Conditions” are compared with… Figure 5 After matching all entities in the list of Shenzhen, Beijing, Zhongshan, Dongguan, Henan, Huizhou, relocation, relocation, processing, advantages and disadvantages, procedures, materials, household registration, conditions, collectives, and individuals, candidate entities can be obtained: Shenzhen, household registration, transfer, and conditions. The set containing these candidate entities can be used as the candidate entity set.
[0073] 2) Based on the header entities in the first entity set, filter the candidate entity set to obtain the second entity set, wherein the header entities of the entity paths corresponding to the second entity set are the same as the header entities in the first entity set.
[0074] For example, based on the head entity of the first entity set being "household registration", a set with "household registration" as the head entity is selected from the candidate entity sets containing "Shenzhen", "household registration", "transfer", and "condition" as the second entity set. The candidate entity sets contain "Shenzhen", "household registration", "transfer", and "condition", but may be a set with "Shenzhen" as the head entity, or may contain "Shenzhen", "household registration", "transfer", and "condition" but with the head entity not being "Shenzhen" and "household registration". Therefore, it is necessary to select a set with "household registration" as the head entity from the candidate entity sets to ensure that the head entity of the entity path corresponding to the second entity set is the same as the head entity of the first entity set, which facilitates the subsequent path generation and sorting stages.
[0075] In some embodiments, candidate entity sets can be recalled from all node sets in a preset intent knowledge graph by means of inverted index or vector similarity sorting.
[0076] In some embodiments, candidate paths can be recalled in a pre-built intent knowledge graph based on the recalled candidate entity set. The recalled paths can take the form of <head entity>, <relationship>, <entity>, etc.
[0077] Step 130: Based on the second entity set, determine at least one entity path corresponding to the second entity set in the intent knowledge graph; wherein the entity path includes at least two entities with an association relationship.
[0078] Specifically, taking the second entity set listed in step 120, which is centered on household registration, as an example, the entity path corresponding to household registration is determined in the intent knowledge graph.
[0079] For example, the physical path could be: household registration → Shenzhen → relocation → Zhongshan;
[0080] Or, household registration → Shenzhen → relocation → Huizhou;
[0081] Or, household registration → Shenzhen → advantages and disadvantages;
[0082] Alternatively, you can apply for a household registration in Shenzhen.
[0083] Or, household registration → Shenzhen → transfer → collective → conditions;
[0084] Or, household registration → Shenzhen → transfer → individual → conditions;
[0085] Or, household registration → Shenzhen → transfer → individual → married → conditions;
[0086] Or, household registration → Shenzhen → transfer → individual → unmarried → conditions;
[0087] Alternatively, you can transfer your household registration to Shenzhen, or meet the following conditions.
[0088] It should be noted that an entity path includes at least two related entities.
[0089] Step 140: Match the first entity set with at least one entity path to obtain the target entity path, and determine the intent of the text to be identified based on the target entity path.
[0090] For example, suppose the first entity set, namely {"Shenzhen", "household registration", "transfer", "conditions"}, is matched with the entity paths listed in step 120. For instance, similarity matching can be used. If the similarity of the main path "household registration → Shenzhen → transfer" is relatively high, then the main path "household registration → Shenzhen → transfer" is taken as the target entity path. Based on the target entity path, the intent of the text to be identified is determined to be "household registration processing".
[0091] To further expand the content of the intent knowledge graph, in some embodiments, once the intent of the text to be identified is determined, the determined intent of the text to be identified is used as a hyperedge to update the intent knowledge graph.
[0092] Specifically, in an intent knowledge graph, a "hyperedge" is defined to represent a set of entities or terms, and intents can represent equivalence relationships between them. For example, the sample "How to apply for a Beijing household registration" is transformed into three entities in the intent knowledge graph: "household registration," "Beijing," and "apply," which are used as node information in the graph. Connecting these nodes forms a hyperedge "household registration -> Beijing -> apply," which is then integrated into the graph. The intent corresponding to the hyperedge is "household registration application."
[0093] For example, after matching the first entity set {"Shenzhen City", "Household Registration", "Transfer", "Condition"} with the entity paths listed in step 120, the paths with high similarity, such as Household Registration → Shenzhen → Transfer → Collective → Condition; or Household Registration → Shenzhen → Transfer → Individual → Condition; or Household Registration → Shenzhen → Transfer → Individual → Married → Condition; or Household Registration → Shenzhen → Transfer → Individual → Unmarried → Condition, can extract Household Registration → Shenzhen → Transfer as a hyperedge from these paths. The intent corresponding to the hyperedge Household Registration → Shenzhen → Transfer can be "Household Registration Processing". Using Household Registration → Shenzhen → Transfer as a hyperedge can connect more other entities and terms, thereby updating the intent knowledge graph. This avoids the problems of high manual cost, training cost, and poor reasoning performance caused by the traditional approach of manually annotating and expanding the content of the intent knowledge graph.
[0094] Furthermore, since the entities in the first entity set are text, and text is an abstract entity of cognition, before matching the first entity set with at least one entity path using the model, each entity needs to be converted into a numerical vector or matrix as the standard input to the machine learning model. The model can be an algorithmic model or a neural network model, and the conversion of entities into vectors can specifically include the following steps:
[0095] Step 210: Convert the entities in the first entity set into corresponding vectors, and convert the entities in each entity path into corresponding vectors.
[0096] For example, suppose the model has a pre-trained dictionary of [household registration, transfer, Shenzhen, of, condition]. This dictionary can be regarded as a bag of words with a capacity of 5. The vector corresponding to each entity can be represented as: household registration: [1, 0, 0, 0, 0]; transfer: [0, 1, 0, 0, 0]; Shenzhen: [0, 0, 1, 0, 0]; of: [0, 0, 0, 1, 0]; condition: [0, 0, 0, 0, 1].
[0097] Here, 1 can represent the number of times each entity appears, and the position of 1 can represent the sorting position of that entity in the dictionary of [household registration, transfer, Shenzhen, of, condition].
[0098] Similarly, the entities in each entity path can be converted into corresponding vectors, which will not be listed here.
[0099] Step 220: Perform similarity matching between the vector corresponding to the first entity set and the vector corresponding to each entity path to obtain the similarity score for each entity path.
[0100] One approach is to use an inverted index or TF-IDF (Term Frequency-Inverse Document Frequency) vector matching to vectorize all entities, and then perform similarity matching on the corresponding entities.
[0101] For example, lightweight calculation methods such as edit distance and keyword weighted distance can be used to perform similarity matching and sorting on the vectors corresponding to the first entity set and the vectors corresponding to each entity path, so as to obtain the similarity score of each entity path.
[0102] For information on similarity calculations, please refer to the following two sentences:
[0103] For example, sentence A: I like watching TV, but I don't like watching movies.
[0104] Sentence B: I don't like watching TV or movies.
[0105] The entity corpus for sentences A and B is: I like to watch TV and movies, no, also.
[0106] In sentence A, "I" appears once, "like" appears twice, "watch" appears twice, "TV" appears once, "movie" appears once, "no" appears once, and "also" appears zero times.
[0107] In sentence B, "I" appears once, "like" appears twice, "watch" appears twice, "TV" appears once, "movie" appears once, "no" appears twice, and "also" appears once.
[0108] Based on the frequency of each entity, the word frequency vector of sentence A can be [1, 2, 2, 1, 1, 1, 0], and the word frequency vector of sentence B can be [1, 2, 2, 1, 1, 2, 1].
[0109] Suppose the angle between these two vectors is α. We can judge the similarity between the vectors by the size of the angle α. Generally speaking, the smaller the angle, the more similar they are.
[0110] The size of the included angle α can be calculated by taking the cosine of the angle. For example, the cosine of the included angle α between the two vectors is cosα=0.938. Since the cosine value of 0.938 is close to 1, the included angle α is close to 0 degrees, which means that the two vectors are very similar. Therefore, the similarity between the two sentences can be determined.
[0111] The similarity between two vectors is obtained through cosine similarity, and thus the similarity between two sentences can be derived. It should be noted that the above method is only one way to obtain similarity scores in this application; this application does not specify which method is used to obtain the similarity scores.
[0112] Step 230: Sort the paths based on their similarity scores and determine the target entity paths based on the sorting results.
[0113] For example, after sorting the similarity scores of each entity path, the path with the highest similarity score, i.e., the top1 path, is selected as the target entity path, and the intent corresponding to the target entity path is used as the intent output of the text to be identified.
[0114] For example, taking the first entity set as {"Shenzhen City", "Household Registration", "Transfer", "Conditions"} mentioned in step 110 as an example, the similarity score with the path Household Registration → Shenzhen → Transfer is 90%; the similarity with the paths Household Registration → Shenzhen → Migration Out and Household Registration → Shenzhen → Migration In is 30%; and the similarity with the path Household Registration → Shenzhen → Advantages and Disadvantages is 20%. Therefore, the path Household Registration → Shenzhen → Transfer is taken as the target entity path, and the user's intended meaning can be determined through this target entity path.
[0115] Furthermore, the path "Household Registration → Shenzhen → Transfer" can be used as the main path, which can include multiple sub-paths, such as "Household Registration → Shenzhen → Transfer → Conditions", "Household Registration → Shenzhen → Transfer → Collective → Conditions", etc. By searching the sub-paths under the main path, the user's intent can be further confirmed.
[0116] By comparing similarity scores, the path with the highest similarity score is selected as the target entity path. This ensures that the information the machine provides to the user is as similar as possible to the text the user inputs, thereby accurately outputting the user's intended meaning and avoiding situations where the user's intent cannot be recognized, which could lead to a poor user experience.
[0117] The intent recognition method of this application can be applied to various application scenarios, such as human-computer interaction scenarios. For example, conversational AI is where machines can engage in conversations similar to humans. For instance, the machine can first collect various topic corpora, and then construct intent knowledge graphs corresponding to each topic based on the corpora. When the machine receives a user's voice input, it will convert the user's voice into text, understand the meaning of the text, and search for the best response from the preset intent knowledge graph based on the semantics of the text, and then provide feedback.
[0118] By way of example, a more complete embodiment is provided herein in conjunction with the above embodiments.
[0119] Taking the user-input question "Conditions for transferring Shenzhen household registration" as an example, the intent recognition method of this application may specifically include the following steps:
[0120] 1) The machine performs entity recognition and extraction on the user-input question "Conditions for transferring Shenzhen household registration", and converts it into the first entity set ["Shenzhen", "Household Registration", "Transfer", "Conditions"];
[0121] 2) Based on the semantics of the text and the extracted first entity set, the machine selects Shenzhen City, household registration, transfer, and conditions as candidate entities from the preset intent knowledge graph, and determines the set containing these candidate entities as the candidate entity set associated with the first entity set.
[0122] 3) Based on the dependencies between the entities Shenzhen City, Household Registration, Transfer, and Conditions, the machine determines Household Registration as the head entity in the first entity set ["Shenzhen City", "Household Registration", "Transfer", "Conditions"].
[0123] 4) The machine selects a set with household registration as the head entity from the candidate entity set of Shenzhen City, household registration, transfer, and conditions based on the head entity of the first entity set, so as to ensure that the head entity of the entity path corresponding to the second entity set is the same as the head entity of the first entity set.
[0124] 5) Based on the second entity set, which is the set of entities with household registration as the head, recall multiple entity paths corresponding to household registration in the intent knowledge graph, and then determine at least one entity path corresponding to the second entity set.
[0125] 6) Convert the entities “Shenzhen City”, “Household Registration”, “Transfer”, and “Condition” in the first entity set into corresponding vectors, and convert the entities in the multiple entity paths corresponding to Household Registration into corresponding vectors.
[0126] 7) Perform similarity matching between the vectors corresponding to the first entity set and the vectors in the multiple entity paths corresponding to the household registration, so as to obtain the similarity score of each entity path corresponding to the household registration.
[0127] 8) Sort the similarity scores of each entity path corresponding to the household registration, and select the path with the highest similarity score, Household Registration → Shenzhen → Transfer, as the target entity path based on the sorting results.
[0128] 9) Based on the path "Household Registration → Shenzhen → Transfer", determine that the user's intention is "Household Registration Processing" and provide the corresponding information to the user accordingly.
[0129] 10) The machine uses the household registration → Shenzhen → transfer as the intent hyperedge and updates the intent knowledge graph.
[0130] Please see Figure 6 , Figure 6 This is a schematic diagram of an embodiment of the intent recognition device provided in this application. The intent recognition device 60 includes an extraction module 61, a retrieval module 62, an entity path determination module 63, and a matching module 64. The functions of each module are as follows:
[0131] Extraction module 61 is used to extract the first entity set from the text to be recognized;
[0132] The retrieval module 62 is used to determine, based on the first entity set, a second entity set associated with the first entity set in a pre-constructed intent knowledge graph; wherein, the intent knowledge graph is built based on a corpus set and is used to represent the association relationships between entities;
[0133] Entity path determination module 63 is used to determine at least one entity path corresponding to the second entity set in the intent knowledge graph based on the second entity set; wherein the entity path includes at least two entities with an association relationship;
[0134] Matching module 64 is used to match the first entity set with at least one entity path to obtain the target entity path, and to determine the intent of the text to be identified based on the target entity path.
[0135] Optionally, the retrieval module 62 is further configured to determine the head entities in the first entity set; based on the determined head entities and the first entity set, to filter the second entity set associated with the first entity set in the intent knowledge graph.
[0136] Optionally, the retrieval module 62 is further configured to determine the head entity in the first entity set based on the dependency relationships between the entities in the first entity set in the text to be identified.
[0137] Optionally, the retrieval module 62 is further configured to determine, in the intent knowledge graph, a candidate entity set associated with the first entity set; and filter the candidate entity set based on the head entities in the first entity set to obtain the second entity set, wherein the head entities of the entity paths corresponding to the second entity set are the same as the head entities in the first entity set.
[0138] Optionally, the matching module 64 is further configured to perform similarity matching between the first entity set and the at least one entity path, and select the entity path with the highest similarity as the target entity path; and determine the intent of the text to be identified based on the target entity path.
[0139] Optionally, the matching module 64 is further configured to convert entities in the first entity set into corresponding vectors, and to convert entities in each entity path into corresponding vectors;
[0140] The vector corresponding to the first entity set is matched with the vector corresponding to each entity path to obtain a similarity score for each entity path.
[0141] The paths of each entity are sorted based on their similarity scores, and the target entity path is determined based on the sorting results.
[0142] Optionally, the matching module 64 is also used to update the intent knowledge graph based on the determined intent of the text to be identified as a hyperedge.
[0143] Please see Figure 7 , Figure 7 This is a schematic diagram of an embodiment of the computer processing device 90 provided in this application. The computer processing device 90 includes a memory 91 and a processor 92. The memory 91 is used to store program data, and the processor 92 is used to execute the program data to implement the following method:
[0144] Extract the first entity set from the text to be recognized;
[0145] Based on the first entity set, determine the second entity set associated with the first entity set in the pre-constructed intent knowledge graph;
[0146] Based on the second entity set, determine at least one entity path corresponding to the second entity set in the intent knowledge graph;
[0147] The first entity set is matched with at least one entity path to obtain the target entity path, and the intent of the text to be identified is determined based on the target entity path.
[0148] Optionally, when executing program data, processor 92 is also used to implement the following methods:
[0149] Determine the head entity in the first entity set;
[0150] Based on the determined head entity and the first entity set, the second entity set associated with the first entity set is filtered in the intent knowledge graph.
[0151] Optionally, when executing program data, processor 92 is also used to implement the following methods:
[0152] Based on the dependency relationships between the entities in the first entity set in the text to be identified, the head entity in the first entity set is determined.
[0153] Optionally, when executing program data, processor 92 is also used to implement the following methods:
[0154] In the intent knowledge graph, a candidate entity set associated with the first entity set is determined; based on the head entities in the first entity set, the candidate entity set is filtered to obtain the second entity set, wherein the head entities of the entity paths corresponding to the second entity set are the same as the head entities in the first entity set.
[0155] Optionally, when executing program data, processor 92 is also used to implement the following methods:
[0156] The first entity set is matched with the at least one entity path for similarity, and the entity path with the highest similarity is selected as the target entity path; the intent of the text to be identified is determined based on the target entity path.
[0157] Optionally, when executing program data, processor 92 is also used to implement the following methods:
[0158] Convert the entities in the first entity set into corresponding vectors, and convert the entities in each entity path into corresponding vectors;
[0159] The vector corresponding to the first entity set is matched with the vector corresponding to each entity path to obtain a similarity score for each entity path.
[0160] The paths of each entity are sorted based on their similarity scores, and the target entity path is determined based on the sorting results.
[0161] Optionally, when executing program data, processor 92 is also used to implement the following methods:
[0162] The intent knowledge graph is updated based on the determined intent of the text to be identified as a hyperedge.
[0163] Optionally, in one embodiment, the intent recognition device 90 may be a chip, a field-programmable gate array (FPGA), a microcontroller, etc. The chip may be a processing chip, such as a CPU, GPU, MCU, etc., or a storage chip, such as DRAM, SRAM, etc.
[0164] Please see Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the computer-readable storage medium 100 provided in this application. The computer-readable storage medium 100 stores program data 101, which, when executed by a processor, is used to implement the following method:
[0165] Extract the first entity set from the text to be recognized;
[0166] Based on the first entity set, determine the second entity set associated with the first entity set in the pre-constructed intent knowledge graph;
[0167] Based on the second entity set, determine at least one entity path corresponding to the second entity set in the intent knowledge graph;
[0168] The first entity set is matched with at least one entity path to obtain the target entity path, and the intent of the text to be identified is determined based on the target entity path.
[0169] Optionally, when the program data 101 is executed by the processor, it is also used to implement the following methods:
[0170] Identify the head entities in the first entity set; based on the identified head entities and the first entity set, filter the second entity set associated with the first entity set in the intent knowledge graph.
[0171] Optionally, when the program data 101 is executed by the processor, it is also used to implement the following methods:
[0172] Based on the dependency relationships between the entities in the first entity set in the text to be identified, the head entity in the first entity set is determined.
[0173] Optionally, when the program data 101 is executed by the processor, it is also used to implement the following methods:
[0174] In the intent knowledge graph, a candidate entity set associated with the first entity set is determined; based on the head entities in the first entity set, the candidate entity set is filtered to obtain the second entity set, wherein the head entities of the entity paths corresponding to the second entity set are the same as the head entities in the first entity set.
[0175] Optionally, when the program data 101 is executed by the processor, it is also used to implement the following methods:
[0176] The first entity set is matched with the at least one entity path for similarity, and the entity path with the highest similarity is selected as the target entity path; the intent of the text to be identified is determined based on the target entity path.
[0177] Optionally, when the program data 101 is executed by the processor, it is also used to implement the following methods:
[0178] Convert the entities in the first entity set into corresponding vectors, and convert the entities in each entity path into corresponding vectors;
[0179] The vector corresponding to the first entity set is matched with the vector corresponding to each entity path to obtain a similarity score for each entity path.
[0180] The paths of each entity are sorted based on their similarity scores, and the target entity path is determined based on the sorting results.
[0181] Optionally, when the program data 101 is executed by the processor, it is also used to implement the following methods:
[0182] The intent knowledge graph is updated based on the determined intent of the text to be identified as a hyperedge.
[0183] When the embodiments of this application are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0184] The above are merely embodiments of this application and do not limit the scope of this patent application. Any equivalent structural or procedural changes made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.
Claims
1. An intention recognition method, characterized by, The method comprises: extracting a first entity set in the text to be identified; determining a second entity set associated with the first entity set in a pre-constructed intent knowledge graph according to the first entity set; determining at least one entity path corresponding to the second entity set in the intent knowledge graph according to the second entity set; matching the first entity set with the at least one entity path to obtain a target entity path, and determining the intent of the text to be identified according to the target entity path; updating the intent knowledge graph according to the determined intent of the text to be identified as a hyperedge.
2. The method of claim 1, wherein, The method comprises: determining a head entity in the first entity set; screening the second entity set associated with the first entity set in the intent knowledge graph based on the determined head entity and the first entity set.
3. The method of claim 2, wherein, The method comprises: determining the head entity in the first entity set according to the dependency relationship of each entity in the first entity set in the text to be identified.
4. The method of claim 2, wherein, The method comprises: determining a candidate entity set associated with the first entity set in the intent knowledge graph; screening the candidate entity set based on the head entity in the first entity set to obtain the second entity set, wherein the head entity of the entity path corresponding to the second entity set is the same as the head entity in the first entity set.
5. The method of claim 1, wherein, The method comprises: performing similarity matching on the first entity set and the at least one entity path, and selecting the entity path with the largest similarity as the target entity path; determining the intent of the text to be identified according to the target entity path.
6. The method of claim 5, wherein, The method comprises: converting the entities in the first entity set into corresponding vectors, and converting the entities in each entity path into corresponding vectors; performing similarity matching on the vectors corresponding to the first entity set and the vectors corresponding to each entity path to obtain the similarity scores of each entity path; sorting the similarity scores of each entity path, and determining the target entity path according to the sorting result.
7. An intention recognition apparatus characterized by comprising: The method comprises: an extraction module for extracting a first entity set in the text to be identified; a retrieval module for determining a second entity set associated with the first entity set in a pre-constructed intent knowledge graph according to the first entity set; wherein the intent knowledge graph is established based on a corpus set and is used to represent the association relationship between entities. The entity path determination module is configured to determine at least one entity path corresponding to the second entity set in the intent knowledge graph according to the second entity set, wherein the entity path comprises at least two entities having a correlation relationship. The matching module is configured to match the first entity set with the at least one entity path to obtain a target entity path, and determine the intent of the text to be recognized according to the target entity path. The matching module is further configured to update the intent knowledge graph by taking the determined intent of the text to be recognized as a hyperedge.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores executable instructions, and the executable instructions are executed by the processor to implement the method in any one of claims 1-6.
9. A computer processing device, characterized by, The device comprises: a processor; and a memory arranged to store computer executable instructions configured to be executed by the processor, the computer executable instructions comprising steps for implementing the method in any one of claims 1-6.
Citation Information
Patent Citations
Method and device for determining answers to questions, computer equipment and storage medium
CN112287095A
Chemical knowledge graph construction method and device, and intelligent question and answer method and device for chemical knowledge
CN112948566A