An information error correction method and device based on a knowledge graph
By using the knowledge graph to checksum correction of the person name-job name triple in the text entered by the user, the problem of difficult to reflect user intentions when correcting errors in the prior art is solved, and the accuracy and user experience of the error correction system are improved.
Patent Information
- Application Number
- CN202010725467.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-07-24
AI Technical Summary
In the prior art, when correcting the error of matching names-job names in the text entered by the user, it is difficult to accurately reflect the user's intentions, resulting in a poor user experience.
By obtaining a triple to be detected that includes the first type of mention words and the second type of mention words, verifying them with a preset knowledge graph, and correcting errors based on the corresponding entities of the mention words in the knowledge graph, priority is given to returning error correction results with high text similarity based on the first type and the second type of mention words.
It improves the accuracy and rationality of the error correction system, can analyze the potential intentions of users from multiple angles, and provides error correction results that are more in line with user needs.
Smart Images

Figure CN113971217B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an information error correction method and device based on a knowledge graph, a computing device, and a computer-readable storage medium. Background Art
[0002] In the existing error correction system, when it is detected that there is an error in the matching of a person name and a position name in the text input by the user, the error of the user is often corrected only by fixed rules, that is, taking the position name as the standard or taking the person name as the standard. At present, although this method can ensure that the result after modification must be correct in the matching of the person name and the position name, it is very difficult to ensure that the modified result is what the user needs. For example, the text input by the user is "Director Zhang San of A", and the user's intention is "Deputy Director Zhang San of A", and the text input by the user is "Director Li Si of A", and the user's intention is "Director Li Xiaosi of A". The above two situations tend to be errors caused by the confusion of the positive and deputy positions of personnel and the omission or omission of characters in the person name respectively. However, if the error correction is only based on the position name or only based on the person name, the user's intention cannot be accurately reflected, resulting in a poor user experience. Summary of the Invention
[0003] In view of this, the embodiments of the present application provide an information error correction method and device based on a knowledge graph, a computing device, and a computer-readable storage medium to solve the technical defects existing in the prior art.
[0004] According to the first aspect of the embodiments of this specification, an information error correction method based on a knowledge graph is provided, including:
[0005] Obtaining a to-be-detected triple containing a first type of mentioned word and a second type of mentioned word;
[0006] Verifying the to-be-detected triple according to a preset knowledge graph, and when the to-be-detected triple fails the verification, correcting the second type of mentioned word according to the second type of entity corresponding to the first type of mentioned word in the knowledge graph;
[0007] When the second type of entity and the second type of mentioned word do not meet the error correction condition, correcting the first type of mentioned word according to the first type of entity corresponding to the second type of mentioned word in the knowledge graph.
[0008] According to the second aspect of the embodiments of this specification, an information error correction device based on a knowledge graph is provided, including:
[0009] A triple acquisition module configured to obtain a to-be-detected triple containing a first type of mentioned word and a second type of mentioned word;
[0010] The first error correction module is configured to verify the triple to be detected according to a preset knowledge graph, and correct the second type of mentioned words according to the second type of entities corresponding to the first type of mentioned words in the knowledge graph when the triple to be detected fails the verification;
[0011] The second error correction module is configured to correct the first type of mentioned words according to the first type of entities corresponding to the second type of mentioned words in the knowledge graph when the second type of entities and the second type of mentioned words do not meet the error correction conditions.
[0012] According to the third aspect of the embodiments of the present specification, a computing device is provided, including a memory, a processor, and computer instructions stored on the memory and executable on the processor. When the processor executes the instructions, the steps of the information error correction method based on the knowledge graph are implemented.
[0013] According to the fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer instructions that implement the steps of the information error correction method based on the knowledge graph when executed by a processor.
[0014] This application extracts a triple to be detected including a first type of mentioned words and a second type of mentioned words from the text input by the user, and uses the graph structure relationship of the knowledge graph to correct the writing of the extracted triple to be detected. In the process of correcting the matching error of the triple to be detected, error correction results with higher text similarity are returned based on the first type of mentioned words and the second type of mentioned words respectively according to the priority, so as to support the analysis of the user's potential intention from multiple angles, effectively increasing the accuracy and rationality of the error correction system. Description of the Drawings
[0015] Figure 1 is a structural block diagram of the computing device provided by the embodiments of the present application;
[0016] Figure 2 is a flowchart of the information error correction method based on the knowledge graph provided by the embodiments of the present application;
[0017] Figure 3 is a flowchart of constructing the triple to be detected provided by the embodiments of the present application;
[0018] Figure 4 is a flowchart of error judgment of the triple to be detected provided by the embodiments of the present application;
[0019] Figure 5 is another flowchart of the information error correction method based on the knowledge graph provided by the embodiments of the present application;
[0020] Figure 6It is another flowchart of the information error correction method based on the knowledge graph provided by the embodiments of the present application;
[0021] Figure 7 It is a flowchart of a specific error correction application provided by the embodiments of the present application;
[0022] Figure 8 It is a flowchart of another specific error correction application provided by the embodiments of the present application;
[0023] Figure 9 It is a flowchart of another specific error correction application provided by the embodiments of the present application;
[0024] Figure 10 It is a flowchart of another specific error correction application provided by the embodiments of the present application;
[0025] Figure 11 It is a schematic structural diagram of the information error correction device based on the knowledge graph provided by the embodiments of the present application. Detailed implementation manners
[0026] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.
[0027] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0028] It should be understood that although the terms first type, second type, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first type may also be referred to as the second type, and similarly, the second type may also be referred to as the first type. First, the noun terms related to one or more embodiments of the present invention are explained.
[0029] Knowledge Graph: That is, Knowledge Graph, a semantic network aiming to describe the conceptual entities in the objective world and the relationships between them, and sometimes also called Knowledge Base.
[0030] Triple: A kind of identification method for knowledge graphs. Common forms include (Entity 1, Relationship, Entity 2) or (Entity, Attribute, Attribute Value). For example, (Yao Ming, Plays for, NBA), (Yao Ming, Height, 2.29m).
[0031] Entity: That is, Entity. Entities are the basic units of knowledge graphs and important language units that carry information in text.
[0032] Mention: That is, Mention. It refers to the language fragment that expresses an entity in natural text.
[0033] Object Properties: Relationships are used to describe the connections between entities. For example, Zhang San's father is Zhang Er, where "father" is the relationship.
[0034] Data Properties: Attributes are used to describe the inherent characteristics of entities. For example, Zhang San's age is twenty-four, where "age" is the attribute.
[0035] Entity Linking: That is, EntityLinking. It refers to mapping the mentions in text to entities in a given knowledge base, which plays a very interesting fundamental role in many fields, such as question answering, semantic search, and information extraction.
[0036] Named Entity Recognition: It refers to automatically identifying named entities from the original data corpus. Since entities are the most basic elements in knowledge graphs, the integrity, accuracy, recall rate, etc. of their extraction will directly affect the quality of knowledge graph construction. We can divide the methods of entity extraction into four types: extraction based on encyclopedia sites or vertical sites, methods based on rules and dictionaries, methods based on statistical machine learning, and extraction methods for the open domain.
[0037] Syntactic Analysis: Syntactic analysis is also a fundamental task in natural language processing. It analyzes the syntactic structure of sentences (subject-predicate-object structure) and the dependency relationships between words (parallel, subordinate, etc.). Through syntactic analysis, it can lay a foundation for application scenarios of natural language processing such as semantic analysis, sentiment tendency, and opinion extraction.
[0038] Edit Distance: That is, Edit Distance. It refers to the minimum number of edit operations required to convert one string into another between two strings. Edit operations include replacing one character with another, inserting a character, and deleting a character. Generally speaking, the smaller the edit distance, the greater the similarity between the two strings.
[0039] In this application, an information error correction method and apparatus, a computing device, and a computer-readable storage medium based on a knowledge graph are provided, and will be described in detail one by one in the following embodiments.
[0040] Figure 1 FIG. 4 shows a structural block diagram of a computing device 100 according to an embodiment of the present specification. The components of the computing device 100 include, but are not limited to, a memory 110 and a processor 120. The processor 120 is connected to the memory 110 through a bus 130, and a database 150 is used to store data.
[0041] The computing device 100 further includes an access device 140, and the access device 140 enables the computing device 100 to communicate via one or more networks 160. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 140 may include one or more of any type of wired or wireless network interfaces (for example, a network interface card (NIC)), such as an IEEE802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0042] In an embodiment of the present specification, the above components of the computing device 100 and Figure 1 other components not shown therein may also be connected to each other, for example, through a bus. It should be understood that Figure 1 the shown structural block diagram of the computing device is only for example purposes and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0043] The computing device 100 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (for example, a smart phone), a wearable computing device (for example, a smart watch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or a PC. The computing device 100 may also be a mobile or stationary server.
[0044] Among them, the processor 120 may execute Figure 2 the steps in the method shown. Figure 2 FIG. 24 is a schematic flowchart of an information error correction method based on a knowledge graph according to an embodiment of the present application, including steps 202 to 206.
[0045] Step 202: Obtain the triple to be detected that contains the first type of mention words and the second type of mention words.
[0046] In the embodiments of the application, as Figure 3 shown, the said step 202 includes step 302 to step 304.
[0047] Step 302: Perform entity named entity recognition on the text statement input by the user to obtain the first type of mention words and the second type of mention words.
[0048] In the above embodiments, by using multi-dimensional entity named entity recognition methods such as extraction based on encyclopedia sites or vertical sites, methods based on rules and dictionaries, methods based on statistical machine learning, and extraction methods for the open domain, etc., perform entity named entity recognition on the text statement input by the user, so as to obtain the first type of mention words and the second type of mention words as specific targets.
[0049] In a specific application, the first type of mention words may be the names of people input by the user, such as "Zhang San", "Li Si" or "Wang Wu", etc., and the second type of mention words may be the position names in the enterprises or administrative organs input by the user, such as "chairman", "legal representative", "director" or "deputy director", etc.
[0050] Step 304: Determine the target path relationship between the first type of mention words and the second type of mention words through syntactic analysis, and construct the triple to be detected according to the path relationship.
[0051] In the above embodiments, determine the corresponding relationship that can connect the first type of mention words and the second type of mention words, that is, the target path relationship, through syntactic analysis, so as to construct the triple to be detected with the first type of mention words and the second type of mention words as nodes and the target path relationship as the path. The triple to be detected includes (the first type of mention words, the target path relationship, the second type of mention words).
[0052] Taking the above specific application as an example, when the first type of mention words, the second type of mention words and the target path relationship are "Li Si", "director" and "position" respectively, the triple to be detected may be (Li Si, position, director).
[0053] This application obtains the first type of mention words and the second type of mention words in the text statement input by the user through entity named entity recognition and syntactic analysis and constructs the triple to be detected, so as to accurately obtain the knowledge content that needs to be corrected in the text statement.
[0054] Step 204: Verify the triple to be detected according to a preset knowledge graph. In the case where the triple to be detected fails the verification, correct the second type of mention word according to the second type of entity corresponding to the first type of mention word in the knowledge graph.
[0055] In an embodiment of the present application, the present application first obtains a large amount of structured data in the target field from the network or existing knowledge databases, and then extracts entities, relationships between entities, and attributes of entities and relationships from the structured data, so as to construct a knowledge graph for the target field.
[0056] In a specific application, the target field may be the government affairs field, that is, a large amount of structured government affairs data is obtained from government affairs information disclosure websites or existing government affairs knowledge databases, such as person names, administrative institution names, position names, etc., so as to construct a government affairs knowledge graph including person name-position name-administrative institution name.
[0057] In the embodiment of the present application, as Figure 4 shown, verifying the triple to be detected according to a preset knowledge graph includes steps 402 to 406.
[0058] Step 402: Perform entity linking on the first type of mention word and the second type of mention word, and respectively obtain the first mapped entity of the first type of mention word mapped in the knowledge graph and the second mapped entity of the second type of mention word mapped in the knowledge graph.
[0059] In the above embodiment, since the mention words input by different users may have text errors such as typos or similar-shaped characters, resulting in ambiguity, entity disambiguation and entity alignment are performed through the method of entity linking, and the first mapped entity of the first type of mention word mapped in the knowledge graph and the second mapped entity of the second type of mention word mapped in the knowledge graph are respectively obtained.
[0060] Step 404: Determine a first similarity between the first type of mention word and the first mapped entity and a second similarity between the second type of mention word and the second mapped entity.
[0061] Specifically, the step 404 includes the following steps:
[0062] S4041: Calculate a first edit distance between the first type of mention word and the first mapped entity, and determine the first similarity according to a ratio of the first edit distance to the number of characters of the first mapped entity.
[0063] In the above embodiments, the minimum number of editing characters required to replace the first type of mentioned word with the first mapped entity is calculated as the first edit distance, and then the first similarity is determined according to the ratio of the first edit distance to the number of characters of the first mapped entity. For example, the first type of mentioned word input by the user is the person's name "Li Si", and the first mapped entity corresponding to the person's name "Li Si" in the knowledge graph is the person's name "Li Xiaosi". At the same time, the first edit distance is "1" and the number of characters of the first mapped entity is "3", then the first similarity is 0.33.
[0064] S4042: Calculate the second edit distance between the second type of mentioned word and the second mapped entity, and determine the second similarity according to the ratio of the second edit distance to the number of characters of the second mapped entity.
[0065] Step S4042 is the same as the content described in step S4041, and the present application will not specifically elaborate here.
[0066] Step 406: Determine whether both the first similarity and the second similarity are less than a preset verification threshold; if so, it is determined that the verification is passed and the process ends; if not, it is determined that the verification is not passed and it is confirmed that there is an error in the triple to be detected.
[0067] In the above embodiments, the calculated first similarity and second similarity are compared with a preset verification threshold. When both the first similarity and the second similarity are less than the preset verification threshold, it indicates that both the first type of mentioned word and the second type of mentioned word can correspond to the entities preset in the knowledge graph, so it is presumed that both the first type of mentioned word and the second type of mentioned word input by the user are correct and no correction is required. When the first similarity and / or the second similarity is greater than the preset verification threshold, it indicates that the first type of mentioned word and / or the second type of mentioned word cannot correspond to the entities preset in the knowledge graph, so it is presumed that there is an error in the first type of mentioned word and / or the second type of mentioned word input by the user and correction is required.
[0068] Optionally, the verification threshold can be set to 0.25 to 0.5.
[0069] The present application calculates the edit distance between the mentioned word and the mapped entity, determines the similarity according to the ratio of the edit distance to the number of characters of the mapped entity, and uses the edit distance and entity linking technology to judge whether there is an error in the triple to be detected, so as to accurately and quickly judge the error of the triple to be detected within a reasonable range.
[0070] In the embodiments of the present application, as Figure 5 shown, correcting the second type of mentioned word according to the second type of entity corresponding to the first type of mentioned word in the knowledge graph includes steps 502 to 508.
[0071] Step 502: Query and obtain the second type of entities that have the target path relationship with the first type of mentioned words in the knowledge graph.
[0072] In the above embodiment, in the process of error correction, in order to be able to analyze the user's intention from multiple angles, the second type of entity that has the target path relationship with the first type of mentioned word is first queried and obtained in the knowledge graph. For example, when the knowledge graph is a government knowledge graph, if the first type of mentioned word is the person's name "Li Si", then the second type of entity that has a target path relationship of "position" with the person's name "Li Si" is the position name "Director".
[0073] Step 504: Determine whether the second-category mentioned words and the second-category entities satisfy a preset node path relationship; if so, execute step 506; if not, determine that the second-category entities and the second-category mentioned words do not satisfy error correction conditions.
[0074] Specifically, determining whether the second-category mentioned word and the second-category entity satisfy a preset node path relationship includes:
[0075] Determine whether the second-category mentioned words and the second-category entities correspond to different path relationships with the same third-category entity in the knowledge graph. Specifically, the present application determines whether the second-category mentioned words in the triple to be detected and the second-category entities corresponding to the first-category mentioned words satisfy a preset node path relationship, that is, whether the second-category mentioned words and the second-category entities correspond to different path relationships with the same third-category entity in the knowledge graph, wherein the third-category entities generally refer to other entities in the knowledge graph excluding the second-category entities and the first-category entities.
[0076] In a specific application, the third type of entity can be the name of the administrative agency excluding the name of the person and the title of the position in the government knowledge graph, and the node path relationship can be the relationship path between the principal and the deputy position, for example, "person's name<—deputy position<—administrative agency name—>principal position—>person's name". At this time, when the first type of mentioned words is the person's name "Li Si", the second type of mentioned words is "Director of Bureau A", and the second type of entity is "Deputy Director of Bureau A", the second type of mentioned words and the second type of entities respectively correspond to the relationship path between the principal and the deputy position of the third type entity "Bureau A".
[0077] Step 506: construct an error correction triple based on the first category of mentioned words and the second category of entities and return it to the user.
[0078] In the above embodiments, if the second type of mentioned word and the second type of entity satisfy a preset node path relationship, it is presumed that there is a problem with the second type of mentioned word input by the user. Then, the second type of mentioned word is replaced according to the second type of entity, and an error correction triple corresponding to the triple to be detected is constructed based on the first type of mentioned word and the second type of entity and returned to the user for reference.
[0079] This application first corrects the second type of mentioned word according to the second type of entity corresponding to the first type of mentioned word in the knowledge graph, and corrects the second type of mentioned word using the preset node path relationship, so as to return the error correction result closest to the user's intention, realizing comprehensive consideration from multiple dimensions.
[0080] Step 206: In the case where the second type of entity and the second type of mentioned word do not meet the error correction condition, correct the first type of mentioned word according to the first type of entity corresponding to the second type of mentioned word in the knowledge graph.
[0081] In the embodiments of this application, as Figure 6 shown, the specific steps of step 206 include step 602 to step 608.
[0082] Step 602: Query and obtain the first type of entity in the knowledge graph where the second type of mentioned word has the target path relationship.
[0083] In the above embodiments, during the error correction process, in order to analyze the user's intention from multiple perspectives, in the case where the second type of mentioned word cannot be corrected according to the second type of entity corresponding to the first type of mentioned word in the knowledge graph, this application further queries and obtains the first type of entity in the knowledge graph where the second type of mentioned word has the target path relationship. For example, in the case where the knowledge graph is a government affairs knowledge graph, if the second type of mentioned word is the institution name "Director of Bureau A", the first type of entity with the target path relationship of "position" to the institution name "Director of Bureau A" is the person's name "Li Xiaosi".
[0084] Step 604: Determine the feature similarity between the first type of mentioned word and the first type of entity, and judge whether the feature similarity is less than a preset verification threshold; if so, execute step 606; if not, execute step 608.
[0085] Specifically, determining the feature similarity between the first type of mentioned word and the first type of entity includes:
[0086] Calculate the feature edit distance between the first type of mentioned words and the first type of entities, and determine the feature similarity according to the ratio of the feature edit distance to the number of words of the first type of entities. Specifically, calculate the minimum number of editing words required to replace the first type of mentioned words with the first type of entities as the feature edit distance, and then determine the first similarity according to the ratio of the feature edit distance to the number of words of the first mapped entity. For example, if the first type of mentioned word input by the user is the person's name "Li Si", and the first type of entity corresponding to the person's name "Li Si" in the knowledge graph is the person's name "Li Xiaosi", at this time, the feature edit distance is "1" and the number of words of the first type of entity is "3", then the feature similarity is 0.33.
[0087] Step 606: Construct an error correction triple based on the first type of entity and the second type of mentioned words and return it to the user.
[0088] In the above embodiment, when the feature similarity is less than the preset verification threshold, it is presumed that there is a problem with the first type of mentioned words input by the user. Then, replace the first type of mentioned words according to the first type of entity, and thus construct an error correction triple corresponding to the triple to be detected based on the first type of entity and the second type of mentioned words, and return it to the user for the user to refer to.
[0089] Step 608: Construct an error correction triple based on the first type of mentioned words and the second type of entities and return it to the user.
[0090] In the above embodiment, when the feature similarity is greater than or equal to the preset verification threshold, it is still presumed that there is a problem with the second type of mentioned words input by the user. Then, replace the second type of mentioned words according to the second type of entity, and thus construct an error correction triple corresponding to the triple to be detected based on the first type of mentioned words and the second type of entities, and return it to the user for the user to refer to.
[0091] Optionally, the verification threshold can be set to 0.25 to 0.5.
[0092] This application extracts the triple to be detected containing the first type of mentioned words and the second type of mentioned words from the text input by the user, uses the graph structure relationship of the knowledge graph to perform writing error correction on the extracted triple to be detected, and in the process of correcting the matching error of the triple to be detected, returns error correction results with higher text similarity based on the first type of mentioned words and the second type of mentioned words respectively according to the priority, so as to support the analysis of the user's potential intention from multiple perspectives, effectively increasing the accuracy and rationality of the error correction system.
[0093] Figure 7Illustrates an information error correction method based on a knowledge graph according to an embodiment of this specification. This information error correction method based on a knowledge graph is described by taking the triple to be detected (Li Si, position, Director of Bureau A) as an example, and includes steps 702 to 710.
[0094] Step 702: Extract the triple to be detected (Li Si, position, Director of Bureau A) from the text information input by the user.
[0095] Step 704: Determine that there is an error in the triple to be detected (Li Si, position, Director of Bureau A) through a preset knowledge graph.
[0096] Step 706: Obtain a second type of entity, "Deputy Director of Bureau A", from the knowledge graph that has a "position" target path relationship with the first type of mentioned word "Li Si" in the triple to be detected.
[0097] Step 708: A preset node path relationship "Li Si <-- Deputy Director <-- Bureau A --> Director --> Wang Wu" is satisfied between the second type of mentioned word "Director of Bureau A" in the triple to be detected and the second type of entity "Deputy Director of Bureau A".
[0098] Step 710: Construct an error correction triple (Li Si, position, Deputy Director of Bureau A) based on the first type of mentioned word and the second type of entity and return it to the user.
[0099] Figure 8 Illustrates an information error correction method based on a knowledge graph according to an embodiment of this specification. This information error correction method based on a knowledge graph is described by taking the triple to be detected (Li Si, position, Director of Bureau A) as an example, and includes steps 802 to 814.
[0100] Step 802: Extract the triple to be detected (Li Si, position, Director of Bureau A) from the text information input by the user.
[0101] Step 808: Determine that there is an error in the triple to be detected (Li Si, position, Director of Bureau A) through a preset knowledge graph.
[0102] Step 806: Obtain a second type of entity, "Secretary of Bureau B", from the knowledge graph that has a "position" target path relationship with the first type of mentioned word "Li Si" in the triple to be detected.
[0103] Step 808: The second type of mentioned word "Director of Bureau A" in the triple to be detected and the second type of entity "Secretary of Bureau B" do not satisfy the preset path rule "Li Si <-- Deputy Director <-- Bureau A --> Director --> Wang Wu".
[0104] Step 810: Obtain the first type of entity "Li Xiaosi" that has a "position" target path relationship with the second type of mention word "Director of Bureau A" in the to-be-detected triple from the knowledge graph.
[0105] Step 812: Calculate that the feature similarity between the first type of mention word "Li Si" and the first type of entity "Li Xiaosi" is 0.33 and less than the preset threshold.
[0106] Step 814: Construct a corrected triple (Li Xiaosi, position, Director of Bureau A) corresponding to the to-be-detected triple based on the first type of entity and the second type of mention word and return it to the user.
[0107] Figure 9 The information correction method based on the knowledge graph according to an embodiment of the present specification is shown. Taking the to-be-detected triple (Li Si, position, Director of Bureau A) as an example, the information correction method based on the knowledge graph includes steps 902 to 914.
[0108] Step 902: Extract the to-be-detected triple (Li Si, position, Director of Bureau A) from the text information input by the user.
[0109] Step 904: Determine that there is an error in the to-be-detected triple (Li Si, position, Director of Bureau A) through the preset knowledge graph.
[0110] Step 906: Obtain the second type of entity "Secretary of Bureau B" that has a "position" target path relationship with the first type of mention word "Li Si" in the to-be-detected triple from the knowledge graph.
[0111] Step 908: The second type of mention word "Director of Bureau A" in the to-be-detected triple does not satisfy the preset path rule "Li Si <-- Deputy Director <-- Bureau A --> Director --> Wang Wu" with the second type of entity "Secretary of Bureau B".
[0112] Step 910: Obtain the first type of entity "Zhang San" that has a "position" target path relationship with the second type of mention word "Director of Bureau A" in the to-be-detected triple from the knowledge graph.
[0113] Step 912: Calculate that the feature similarity between the first type of mention word "Li Si" and the first type of entity "Zhang San" is 1 and greater than the preset threshold.
[0114] Step 914: Construct a corrected triple (Li Si, position, Deputy Director of Bureau A) based on the first type of mention word and the second type of entity and return it to the user.
[0115] Figure 10The information error correction method based on a knowledge graph according to an embodiment of this specification is described by taking the triple to be detected (Li Si, position, director of Bureau A) as an example, and includes steps 1002 to 1008.
[0116] Step 1002: Extract the triple to be detected (Li Si, position, director of Bureau A) from the text information input by the user.
[0117] Step 1004: Determine that there is an error in the triple to be detected (Li Si, position, director of Bureau A) through a preset knowledge graph.
[0118] Step 1006: In the case where a second type of entity corresponding to the first type of mention word "Li Si" in the triple to be detected cannot be obtained from the knowledge graph, obtain a first type of entity "Li Xiaosi" that has a "position" target path relationship with the second type of mention word "director of Bureau A" in the triple to be detected from the knowledge graph.
[0119] Step 1008: Construct a corrected triple (Li Xiaosi, position, director of Bureau A) corresponding to the triple to be detected according to the first type of entity and the second type of mention word and return it to the user.
[0120] Corresponding to the above method embodiment, this specification also provides an embodiment of an information error correction device based on a knowledge graph. Figure 11 The structural schematic diagram of an information error correction device based on a knowledge graph according to an embodiment of this specification is shown. As Figure 11 shown, the device includes:
[0121] A triple acquisition module 1101, configured to acquire a triple to be detected including a first type of mention word and a second type of mention word;
[0122] A first error correction module 1102, configured to verify the triple to be detected according to a preset knowledge graph, and correct the second type of mention word according to the second type of entity corresponding to the first type of mention word in the knowledge graph in the case where the triple to be detected fails the verification;
[0123] A second error correction module 1103, configured to correct the first type of mention word according to the first type of entity corresponding to the second type of mention word in the knowledge graph in the case where the second type of entity and the second type of mention word do not meet the error correction condition.
[0124] Optionally, the triple acquisition module 1101 includes:
[0125] The entity recognition unit is configured to perform entity named recognition on the text statement input by the user to obtain the first type of mentioned words and the second type of mentioned words;
[0126] The triple construction unit is configured to determine the target path relationship between the first type of mentioned words and the second type of mentioned words through syntactic analysis, and construct the triple to be detected according to the path relationship.
[0127] Optionally, the first error correction module 1102 includes:
[0128] The entity mapping unit is configured to perform entity linking on the first type of mentioned words and the second type of mentioned words, and respectively obtain the first mapped entity of the first type of mentioned words mapped in the knowledge graph and the second mapped entity of the second type of mentioned words mapped in the knowledge graph;
[0129] The similarity calculation unit is configured to determine the first similarity between the first type of mentioned words and the first mapped entity and the second similarity between the second type of mentioned words and the second mapped entity;
[0130] The threshold verification unit is configured to determine whether both the first similarity and the second similarity are less than a preset verification threshold; if so, it is determined that the verification is passed and the execution of the first verification unit ends; if not, it is determined that the verification is not passed and it is confirmed that there is an error in the triple to be detected, and the second verification unit is executed.
[0131] Optionally, the first similarity calculation unit includes:
[0132] The edit distance calculation sub-unit is configured to calculate the first edit distance between the first type of mentioned words and the first mapped entity, and determine the first similarity according to the ratio of the first edit distance to the number of characters of the first mapped entity;
[0133] The similarity calculation sub-unit is configured to calculate the second edit distance between the second type of mentioned words and the second mapped entity, and determine the second similarity according to the ratio of the second edit distance to the number of characters of the second mapped entity.
[0134] Optionally, the first error correction module 1102 includes:
[0135] The first entity acquisition unit is configured to query and obtain the second type of entity in the knowledge graph that has the target path relationship with the first type of mentioned words;
[0136] The path relationship judgment unit is configured to determine whether the second type of mentioned words and the second type of entity satisfy a preset node path relationship; if so, execute the first error correction unit; if not, it is determined that the second type of entity and the second type of mentioned words do not satisfy the error correction condition;
[0137] The first error correction unit is configured to construct an error correction triple based on the first type of mentioned words and the second type of entities and return it to the user.
[0138] Optionally, the path relationship judgment unit includes:
[0139] A path relationship judgment subunit is configured to judge whether the second type of mentioned words and the second type of entities respectively correspond to different path relationships with the same third type of entity in the knowledge graph.
[0140] Optionally, the second error correction module 1103 includes:
[0141] A second entity acquisition unit is configured to query and acquire a first type of entity in the knowledge graph where the second type of mentioned words has the target path relationship;
[0142] A similarity comparison unit is configured to determine the feature similarity between the first type of mentioned words and the first type of entity, and judge whether the feature similarity is less than a preset verification threshold; if so, execute the second error correction unit; if not, execute the third error correction unit;
[0143] A second error correction unit is configured to construct an error correction triple based on the first type of entity and the second type of mentioned words and return it to the user;
[0144] A third error correction unit is configured to construct an error correction triple based on the first type of mentioned words and the second type of entities and return it to the user.
[0145] Optionally, the similarity comparison unit includes:
[0146] A similarity comparison subunit is configured to calculate the feature edit distance between the first type of mentioned words and the first type of entity, and determine the feature similarity according to the ratio of the feature edit distance to the number of characters of the first type of entity.
[0147] This application extracts person-name - position-name triples from the text input by the user, and uses the graph structure relationship of the knowledge graph to perform writing error correction on the extracted person-name - position-name triples. In the process of correcting the matching errors of the person-name - position-name triples, error correction results with higher text similarity are returned based on the person name and the position name respectively according to the priority, so as to support the analysis of the user's potential intention from multiple angles, effectively increasing the accuracy and rationality of the error correction system.
[0148] It should be noted that each component in the apparatus claim should be understood as a functional module that must be established to implement each step of the program flow or each step of the method. These functional modules are not actual functional partitions or separation limitations. The apparatus claim defined by such a set of functional modules should be understood as mainly implementing the functional module architecture of the solution through the computer program recorded in the specification, rather than mainly implementing the physical apparatus of the solution through hardware means.
[0149] An embodiment of the present application further provides a computing device, including a memory, a processor, and computer instructions stored on the memory and executable on the processor. When the processor executes the instructions, the following steps are implemented:
[0150] Obtain a to-be-detected triple containing a first type of mentioned word and a second type of mentioned word;
[0151] Verify the to-be-detected triple according to a preset knowledge graph. In the case where the to-be-detected triple fails the verification, correct the second type of mentioned word according to the second type of entity corresponding to the first type of mentioned word in the knowledge graph;
[0152] In the case where the second type of entity and the second type of mentioned word do not meet the error correction condition, correct the first type of mentioned word according to the first type of entity corresponding to the second type of mentioned word in the knowledge graph.
[0153] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions. When the instructions are executed by a processor, the steps of the information error correction method based on a knowledge graph as described above are implemented.
[0154] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this computer-readable storage medium and the technical solution of the above information error correction method based on a knowledge graph belong to the same concept. For the details not described in detail in the technical solution of the computer-readable storage medium, reference can be made to the description of the technical solution of the above information error correction method based on a knowledge graph.
[0155] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0156] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0157] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0158] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0159] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments to better explain the principle and practical application of the present application, so that those skilled in the art can understand and utilize the present application well. The present application is only limited by the claims and their full scope and equivalents.
Claims
1. An information error correction method based on a knowledge graph, characterized in that, Including: Obtain a triple to be detected that includes a first type of mention word, a target path relationship, and a second type of mention word; Verify the triple to be detected according to a preset knowledge graph. When the triple to be detected fails the verification, correct the second type of mention word according to the second type of entity corresponding to the first type of mention word in the knowledge graph. Among them, correcting the second type of mention word according to the second type of entity corresponding to the first type of mention word in the knowledge graph includes: querying and obtaining, in the knowledge graph, a second type of entity that has the target path relationship with the first type of mention word; determining whether the second type of mention word and the second type of entity satisfy a preset node path relationship; if so, construct a corrected triple according to the first type of mention word and the second type of entity and return it to the user. Determining whether the second type of mention word and the second type of entity satisfy a preset node path relationship includes: determining whether the second type of mention word and the second type of entity correspond to different path relationships with the same third type of entity in the knowledge graph; If not, determine that the second type of entity and the second type of mention word do not meet the correction condition. When the second type of entity and the second type of mention word do not meet the correction condition, correct the first type of mention word according to the first type of entity corresponding to the second type of mention word in the knowledge graph.
2. The method according to claim 1, wherein Obtaining a triple to be detected that includes a first type of mention word and a second type of mention word includes: Perform entity naming recognition on the text statement input by the user to obtain the first type of mention word and the second type of mention word; Determine the target path relationship between the first type of mention word and the second type of mention word through syntactic analysis, and construct the triple to be detected according to the path relationship. The triple to be detected includes a first type of mention word, a target path relationship, and a second type of mention word.
3. The method according to claim 1, characterized in that, Verifying the triple to be detected according to a preset knowledge graph includes: Perform entity linking on the first type of mention word and the second type of mention word, and respectively obtain a first mapped entity of the first type of mention word mapped in the knowledge graph and a second mapped entity of the second type of mention word mapped in the knowledge graph; Determine a first similarity between the first type of mention word and the first mapped entity and a second similarity between the second type of mention word and the second mapped entity; Judge whether both the first similarity and the second similarity are less than a preset verification threshold; If so, determine that the verification is passed and end; If not, determine that the verification fails and confirm that there is an error in the triple to be detected.
4. The method according to claim 3, characterized in that, Determining a first similarity between the first type of mention word and the first mapped entity and a second similarity between the second type of mention word and the second mapped entity includes: Calculate a first edit distance between the first type of mention word and the first mapped entity, and determine the first similarity according to the ratio of the first edit distance to the number of characters of the first mapped entity; Calculate a second edit distance between the second type of mentioned word and the second mapped entity, and determine the second similarity according to a ratio of the second edit distance to the number of characters of the second mapped entity.
5. The method according to claim 1, characterized in that Correct the first type of mentioned word according to the first type of entity corresponding to the second type of mentioned word in the knowledge graph, including: Query and obtain, in the knowledge graph, the first type of entity having the target path relationship with the second type of mentioned word; Determine a feature similarity between the first type of mentioned word and the first type of entity, and determine whether the feature similarity is less than a preset verification threshold; If so, construct a correction triple according to the first type of entity and the second type of mentioned word and return it to the user; If not, construct a correction triple according to the first type of mentioned word and the second type of entity and return it to the user.
6. The method according to claim 5, wherein Determine the feature similarity between the first type of mentioned word and the first type of entity, including: Calculate a feature edit distance between the first type of mentioned word and the first type of entity, and determine the feature similarity according to a ratio of the feature edit distance to the number of characters of the first type of entity.
7. An information error correction device based on a knowledge graph, characterized in that, Including: A triple acquisition module configured to acquire a triple to be detected including a first type of mentioned word, a target path relationship, and a second type of mentioned word; A first correction module configured to verify the triple to be detected according to a preset knowledge graph, and correct the second type of mentioned word according to a second type of entity corresponding to the first type of mentioned word in the knowledge graph when the triple to be detected fails the verification; The first correction module includes: a first entity acquisition unit configured to query and obtain, in the knowledge graph, a second type of entity having the target path relationship with the first type of mentioned word; a path relationship judgment unit configured to judge whether the second type of mentioned word and the second type of entity satisfy a preset node path relationship; if so, execute a first correction unit; if not, determine that the second type of entity and the second type of mentioned word do not satisfy the correction condition; a first correction unit configured to construct a correction triple according to the first type of mentioned word and the second type of entity and return it to the user; the path relationship judgment unit includes: a path relationship judgment subunit configured to judge whether the second type of mentioned word and the second type of entity respectively correspond to different path relationships with the same third type of entity in the knowledge graph; A second correction module configured to correct the first type of mentioned word according to a first type of entity corresponding to the second type of mentioned word in the knowledge graph when the second type of entity and the second type of mentioned word do not satisfy the correction condition.
8. A computing device, comprising a memory, a processor, and computer instructions stored on the memory and executable on the processor, characterized in that, When the processor executes the instructions, the steps of the method according to any one of claims 1-6 are implemented.
9. A computer-readable storage medium storing computer instructions, characterized in that, When the instructions are executed by the processor, the steps of the method according to any one of claims 1-6 are implemented.
Citation Information
Patent Citations
False news detection method, electronic device and computer readable storage medium
CN110275965A
Semantic error correction method, electronic equipment and storage medium
CN111291571A