A knowledge proposition error correction method and system based on knowledge graph optimization and upgrading

By constructing the initial knowledge graph based on the knowledge graph method, combining deep learning and graph neural network models to detect logical contradictions and errors in propositions, and using an integrated learning algorithm for error correction, the problem that traditional methods are difficult to deeply analyze semantic errors is solved, and efficient and accurate knowledge proposition error correction is achieved.

CN120373298BActive Publication Date: 2025-10-17网才科技(广州)集团股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510471403.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-10-17
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Traditional knowledge proposition error correction methods have difficulty in deeply analyzing semantic errors, logical contradictions, and deep-seated knowledge errors, and are unable to automatically make accurate error correction decisions.

Method used

Based on the knowledge graph optimization and upgrading method, the initial knowledge graph is constructed through entity linking, relationship extraction and attribute alignment. Combined with deep learning semantic parsing and graph neural network model, logical contradictions, knowledge errors and information missing in propositions are detected, and integrated learning algorithm is used to make error correction decisions.

Benefits of technology

It improves the quality and efficiency of knowledge proposition error correction, ensures the accuracy and reliability of knowledge, and avoids blind error correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373298B_ABST
    Figure CN120373298B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of knowledge processing, and particularly discloses a knowledge proposition error correction method and system based on knowledge graph optimization and upgrading, which comprises the following steps: effective information is extracted from multi-source knowledge data by using entity linking, relation extraction and attribute alignment methods to construct an initial knowledge graph; knowledge propositions are subjected to syntax, semantic and logical analysis by deep learning-based semantic analysis to generate proposition element triplets; the triplets are encoded by using a graph neural network to obtain proposition vectors and neighborhood feature vectors, based on which structural similarity is calculated, and rules and statistical methods are combined to detect proposition errors; an integrated learning algorithm is used to classify error sources, and error correction decisions are made in combination with multi-source evidence to obtain accurate error correction results, thereby improving the accuracy and reliability of knowledge propositions; the error correction process is more logical and reliable, blind error correction is avoided, and the quality and efficiency of knowledge proposition error correction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of knowledge processing, and in particular to a knowledge proposition error correction method and system based on knowledge graph optimization and upgrading. BACKGROUND

[0002] In today's era of information explosion, the dissemination and application of knowledge are extremely widespread. The accuracy of knowledge propositions, as an important carrier of knowledge, is crucial. Whether it is the compilation of teaching materials in the field of education, the writing of papers in academic research, or the application of knowledge-related scenarios such as intelligent question-answering systems and knowledge graph construction, all rely on accurate and error-free knowledge propositions. Any errors in knowledge propositions can lead to misinformation, misleading learners, researchers, and people who rely on these knowledge to make decisions. With the rapid development of artificial intelligence technology, knowledge graphs have emerged as a powerful tool for knowledge representation and organization. Knowledge graphs visually display the relationships between entities in a graphical structure, integrating a large amount of knowledge information and providing new ideas for the management and application of knowledge. Applying knowledge graph technology to knowledge proposition error correction is expected to break through the limitations of traditional methods and achieve more accurate and in-depth error correction.

[0003] Traditional knowledge proposition error correction methods are mainly based on simple grammar rule checking and limited lexical matching. With the increasing complexity and diversity of knowledge, these methods gradually fail to meet actual needs. For example, when faced with semantic errors, logical contradictions, and deep-seated knowledge errors, traditional methods often seem inadequate. Traditional methods currently cannot use advanced knowledge graph optimization and upgrading-based semantic analysis methods to deeply analyze the grammar, semantic roles, and logical structure of propositions, making the understanding of propositions superficial. They also cannot fully exploit context semantics and graph structure features, making it difficult to discover deep-seated errors in propositions. Moreover, when there are contradictions in detected errors, they cannot automatically make accurate error correction decisions.

[0004] Therefore, the present application proposes a knowledge proposition error correction method and system based on knowledge graph optimization and upgrading. SUMMARY

[0005] The present invention provides a knowledge proposition error correction method and system based on the optimization and upgrading of the knowledge graph. By extracting effective information from multi-source knowledge data through methods such as entity linking, relationship extraction and attribute alignment, an initial knowledge graph is constructed to lay a comprehensive knowledge foundation for error correction and avoid the limitation of a single data source. Deep learning semantic parsing is used to conduct in-depth analysis of the knowledge propositions to be corrected, and proposition element triplets are generated to accurately extract key information and clearly present the structure to facilitate problem discovery. With the help of a graph neural network model, the proposition semantics and knowledge graph structural information are integrated. By calculating the structural similarity, combined with preset logical conflict rules and statistical learning methods, the logical contradictions, knowledge errors and information missing of the proposition are comprehensively detected from multiple dimensions such as semantics and structure. An integrated learning algorithm is used to classify the error traceability, and scientific error correction decisions are made based on relevant evidence collected from multiple sources to avoid blind error correction, effectively improve the quality and efficiency of knowledge proposition error correction, and effectively guarantee the accuracy and reliability of knowledge.

[0006] The present invention provides a knowledge proposition error correction method based on knowledge graph optimization and upgrading, comprising:

[0007] S1: Extract valid entities, relationships, and attribute information from multi-source knowledge data based on entity linking, relationship extraction, and attribute alignment methods, and construct an initial knowledge graph based on the valid entities, relationships, and attribute information from multi-source knowledge data;

[0008] S2: Use a deep learning-based semantic parsing method to perform syntax analysis, semantic role labeling, and logical structure extraction on the knowledge proposition to be corrected, obtain the entity, relationship, and attribute information in the knowledge proposition, and generate a triple representation of the proposition elements based on the entity, relationship, and attribute information in the knowledge proposition;

[0009] S3: Use the graph neural network model to embed the triple representation of proposition elements to obtain the proposition vector containing contextual semantics and graph structure features, and extract the neighborhood feature vector of the associated entity in the initial knowledge graph;

[0010] S4: Based on the neighborhood feature vectors of the associated entities in the initial knowledge graph, the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph is calculated. In combination with the preset logical conflict rules and statistical learning methods, logical contradictions, knowledge errors, and information missing in the proposition are detected;

[0011] S5: Use an integrated learning algorithm to trace and classify the logical contradictions, knowledge errors, and information missing in the error-correcting knowledge propositions, and combine relevant evidence collected from multi-source knowledge data to make correction decisions on the error-correcting knowledge propositions and obtain correction results.

[0012] Optionally, it also includes:

[0013] S6: verifying the error correction result, and if the verification is passed, updating the initial knowledge graph based on relevant evidence, wherein the verification includes consistency verification, rationality verification, and validity verification.

[0014] Optionally, S1: extracting valid entities, relationships, and attribute information in the multi-source knowledge data based on an entity linking method, a relationship extraction method, and an attribute alignment method, and constructing an initial knowledge graph based on the valid entities, relationships, and attribute information in the multi-source knowledge data, including:

[0015] Removing HTML tags, special characters, and duplicate information in the multi-source knowledge data to obtain preprocessed multi-source knowledge data;

[0016] Identifying all entities in the preprocessed multi-source knowledge data based on predefined rules, and learning entity attribute features and context semantic features of each entity in the preprocessed multi-source knowledge data using a supervised learning algorithm;

[0017] Linking and disambiguating all entities in the preprocessed multi-source knowledge data based on a multi-dimensional matching linking method, entity attribute features, and context semantic features of all entities to obtain all valid entities in the preprocessed multi-source knowledge data;

[0018] Aligning the entity attribute features of all disambiguated entities in the preprocessed multi-source knowledge data to obtain attribute information of all valid entities;

[0019] Analyzing the coexistence context of each pair of valid entities in the preprocessed multi-source knowledge data based on a relationship extraction method to identify and obtain relationships between all valid entities;

[0020] Constructing an initial knowledge graph based on all valid entities and corresponding relationships and attribute information in the preprocessed multi-source knowledge data.

[0021] Optionally, S2: performing syntax analysis, semantic role labeling, and logical structure extraction on the knowledge proposition to be corrected using a deep learning-based semantic parsing method, obtaining entity, relationship, and attribute information in the knowledge proposition, and generating proposition element triple representation based on the entity, relationship, and attribute information in the knowledge proposition, including:

[0022] Constructing a semantic parsing model based on a deep learning method, and identifying the syntax structure, semantic role, and logical structure in the knowledge proposition to be corrected based on the semantic parsing model;

[0023] Determining the part-of-speech standard result and syntactic dependency analysis result of the knowledge proposition to be corrected based on the syntax structure of the knowledge proposition to be corrected;

[0024] Identifying the implicit relationship in the knowledge proposition to be corrected based on the syntactic dependency analysis result of the knowledge proposition to be corrected and the attention mechanism;

[0025] all entities and corresponding attribute information in the knowledge proposition to be corrected are recognized based on the part-of-speech tagging result and the semantic role labeling of the knowledge proposition to be corrected, and the relationships between all entities in the knowledge proposition to be corrected are recognized based on the implicit relationships and the logical structure;

[0026] The proposition element triple representation is generated based on all entities and corresponding attribute information and the relationships between all entities in the knowledge proposition to be corrected.

[0027] Optionally, S3: the proposition element triple representation is embedded and coded using a graph neural network model to obtain a proposition vector containing context semantics and graph structure features, and the neighborhood feature vector of the associated entity in the initial knowledge graph is extracted, including:

[0028] The semantic embedding of the entity and the relationship of the knowledge proposition to be corrected is obtained by analyzing the proposition element triple representation based on the pre-trained language model, and the initial embedding of the knowledge proposition to be corrected is generated in combination with the pre-defined structure features of all entities in the graph structure prior in the knowledge proposition to be corrected;

[0029] Based on the double-attention graph convolution layer and the initial embedding of the knowledge proposition to be corrected, the feature of each entity node corresponding to all entities in the graph structure in the knowledge proposition to be corrected and the feature of all node corresponding to all entities are aggregated to obtain the feature message vector of all entity nodes in the proposition subgraph;

[0030] The feature message vector of all entity nodes in the proposition subgraph is mean-pooled to obtain the proposition vector;

[0031] The K-order neighborhood features of all associated entities of all entities in the knowledge proposition to be corrected in the initial knowledge graph are extracted, and the K-order neighborhood features of all associated entities are weighted and aggregated by a multi-head attention mechanism to obtain the neighborhood feature vector of the associated entity in the initial knowledge graph.

[0032] Optionally, S4: based on the neighborhood feature vector of the associated entity in the initial knowledge graph, the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph is calculated, and the logical contradiction, knowledge error and information missing in the proposition are detected in combination with the preset logical conflict rule and the statistical learning method, including:

[0033] Based on the neighborhood feature of the associated entity in the initial knowledge graph and the proposition vector, the similarity score of the proposition vector and the single associated entity node in the corresponding subgraph in the knowledge graph is calculated:

[0034] ;

[0035] In the formula, is the proposition vector and the single associated entity node in the corresponding subgraph in the knowledge graph The similarity score of Used to convert the input numerical vector into a probability distribution vector. is the weight matrix used to map the concatenated features to similarity scores, For a single associated entity node The domain feature vector of is the vector concatenation operation, is the bias term;

[0036] Obtaining a set of neighboring entity nodes of the associated entity, and calculating a similarity score between the proposition vector and each entity node in the set of neighboring entity nodes of the associated entity based on a neighborhood feature vector of each entity node in the set of neighboring entity nodes of the associated entity;

[0037] Calculate the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph based on the similarity score between the proposition vector and a single associated entity node or each entity node in the set of neighborhood entity nodes in the corresponding subgraph in the knowledge graph;

[0038] Based on the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph, and combined with preset logical conflict rules and statistical learning methods, logical contradictions, knowledge errors, and information missing in the proposition are detected.

[0039] Optionally, based on the similarity score between the proposition vector and a single associated entity node in the corresponding subgraph in the knowledge graph or each entity node in a set of neighborhood entity nodes, the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph is calculated, including:

[0040] ;

[0041] Where, is the proposition vector Corresponding subgraph in the knowledge graph The structural similarity of is the total number of associated entity nodes in the corresponding subgraph, is the set of associated entity nodes in the subgraph, is a hyperparameter and , is the associated entity node in the corresponding subgraph The set of neighborhood entity nodes, is the proposition vector and the associated entity node in the corresponding subgraph Entity nodes in the neighborhood entity node set The similarity score of is the associated entity node in the corresponding subgraph The total number of entity nodes in the neighborhood entity node set.

[0042] Optionally, based on the structure similarity of the proposition vector and the corresponding subgraph in the knowledge graph, and combined with the preset logical conflict rule and the statistical learning method, logical contradictions, knowledge errors, and information missing in the proposition are detected, including:

[0043] All proposition element mutual exclusion relationships are screened out in the proposition element triple representation based on the preset logical conflict rule;

[0044] Logical contradictions in the proposition are detected based on all proposition element mutual exclusion relationships and the structure similarity of the proposition vector and the corresponding subgraph in the knowledge graph;

[0045] The structure similarity of the proposition vector and the corresponding subgraph in the knowledge graph is spliced with the proposition element triple to obtain a proposition feature vector;

[0046] Based on the proposition feature vector and the knowledge error detection classifier and the information missing detection classifier, knowledge errors and information missing in the proposition are detected.

[0047] Optionally, S5: using an ensemble learning algorithm to trace the source classification of logical contradictions, knowledge errors, and information missing in the knowledge proposition to be corrected, and combining related evidence collected from multi-source knowledge data to make a correction decision for the knowledge proposition to be corrected, and obtaining a correction result, including:

[0048] Based on multiple base learners, the logical contradictions, knowledge errors, and information missing in the proposition are traced and classified to obtain multiple error types of the knowledge proposition to be corrected;

[0049] Based on the ensemble learning algorithm, the multiple error types of the knowledge proposition to be corrected are integrated to obtain the final error type of the knowledge proposition to be corrected;

[0050] Based on the final error type of the knowledge proposition to be corrected, and combined with the related evidence collected from multi-source knowledge data, a correction decision is made for the knowledge proposition to be corrected, and a correction result is obtained.

[0051] In addition, a knowledge proposition correction system based on knowledge graph optimization and upgrading is also provided, including:

[0052] A knowledge graph construction module is used to extract effective entities, relationships, and attribute information in multi-source knowledge data based on entity linking methods, relationship extraction methods, and attribute alignment methods, and to construct an initial knowledge graph based on the effective entities, relationships, and attribute information in the multi-source knowledge data;

[0053] A triple generation module is used to perform syntax analysis, semantic role labeling, and logical structure extraction on the knowledge proposition to be corrected using a deep learning-based semantic parsing method, to obtain entity, relationship, and attribute information in the knowledge proposition, and to generate a proposition element triple representation based on the entity, relationship, and attribute information in the knowledge proposition.

[0054] The neighborhood feature extraction module is used for embedding coding of the proposition element triple representation by using a graph neural network model, to obtain a proposition vector containing context semantics and graph structure features, and simultaneously extract neighborhood feature vectors of associated entities in the initial knowledge graph;

[0055] The error detection module is used for calculating the structural similarity of the proposition vector and the corresponding subgraph in the knowledge graph based on the neighborhood feature vectors of the associated entities in the initial knowledge graph, and combining preset logical conflict rules and statistical learning methods to detect logical contradictions, knowledge errors and information omissions in the proposition.

[0056] The proposition error correction module is used for tracing and classifying logical contradictions, knowledge errors and information omissions in the knowledge proposition to be corrected by using an ensemble learning algorithm, and combining relevant evidence collected from multiple source knowledge data to make error correction decisions on the knowledge proposition to be corrected, to obtain error correction results.

[0057] The beneficial effects of the present application relative to the prior art are: effective information is extracted from multiple source knowledge data by entity linking, relation extraction and attribute alignment to construct an initial knowledge graph, laying a comprehensive knowledge foundation for error correction and avoiding the limitations of a single data source. Deep learning semantic analysis is used to deeply analyze the knowledge proposition to be corrected, to generate proposition element triple and accurately extract key information, and to clearly present the structure to facilitate problem discovery. The graph neural network model is used to fuse the proposition semantics and the knowledge graph structure information, to calculate the structural similarity, and to combine the preset logical conflict rules and the statistical learning method to comprehensively detect the logical contradictions, knowledge errors and information omissions of the proposition from multiple dimensions such as semantics and structure. The ensemble learning algorithm is used to classify the errors, and scientific error correction decisions are made according to the relevant evidence collected from multiple sources, to avoid blind error correction, effectively improve the quality and efficiency of knowledge proposition error correction, and effectively guarantee the accuracy and reliability of the knowledge.

[0058] Other features and advantages of the present application will be set forth in the following description of the application, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the application. The objects and other advantages of the present application can be realized and obtained by the structure particularly pointed out in the application document.

[0059] The technical solutions of the present application will be further described in detail below with the help of the drawings and examples. BRIEF DESCRIPTION OF DRAWINGS

[0060] The accompanying drawings are used to provide a further understanding of the present application, and constitute a part of the specification, together with the embodiments of the present application, to explain the present application, and do not constitute a limitation of the present application. In the drawings:

[0061] Figure 1An execution logic diagram of a knowledge proposition error correction method based on knowledge graph optimization and upgrading in an embodiment of the present application;

[0062] Figure 2 An execution logic diagram of a knowledge proposition error correction system based on knowledge graph optimization and upgrading in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0064] Reference Figure 1 The present application provides an embodiment of a knowledge proposition error correction method based on knowledge graph optimization and upgrading, which comprises:

[0065] S1: Based on the entity linking method, the relationship extraction method, and the attribute alignment method, the effective entities, relationships, and attribute information in the multi-source knowledge data are extracted, and the initial knowledge graph is constructed based on the effective entities, relationships, and attribute information in the multi-source knowledge data. The "multi-source knowledge data" here can be text materials of different formats, such as web articles, academic papers, and knowledge base documents. The "entity linking method" refers to associating different expressions that actually point to the same real thing in these data. For example, in some texts, "the Forbidden City" and "the Palace Museum" both refer to the same historical building, and through entity linking, it can be determined that they are the same entity. The "relationship extraction method" is used to discover the connection between entities, such as in the text describing the relationship between a person and a place, the relationship "Zhang San-lives in-Beijing" can be extracted. The "attribute alignment method" ensures that the attributes of the same entity obtained from different data sources are consistent. For example, for the entity "giant panda", the "diet" attribute of different materials should be unified as "mainly eating bamboo". Through these three methods, the effective entity, relationship, and attribute information are extracted, and then the initial knowledge graph is constructed like building a network, with entities as nodes, relationships as edges connecting nodes, and attributes as characteristics of nodes. This graph becomes the basis for subsequent error correction.

[0066] S2: Adopting a semantic parsing method based on deep learning to perform syntax analysis, semantic role labeling and logical structure extraction on the knowledge proposition to be corrected, to obtain entity, relationship and attribute information in the knowledge proposition, and to generate proposition element triple representation based on the entity, relationship and attribute information in the knowledge proposition; "syntax analysis" is to analyze the sentence structure of the proposition, such as the sentence "bird flies in the sky", syntax analysis can determine that "bird" is the subject, "fly" is the predicate, and "in the sky" is the adverbial. "Semantic role labeling" clearly defines the role each component plays in the semantics. In the above sentence, "bird" is the performer of the action "fly". "Logical structure extraction" sorts out the logical relationships between the parts of the proposition, such as cause and effect, parallelism, etc. Through these analyses, the entity (such as "bird"), relationship (such as "fly in the sky"), and attribute information (such as the color and size of the bird) in the proposition are obtained, and then they are sorted into proposition element triple representation, usually in the form of (entity, relationship, entity or attribute), such as (bird, in, sky), (bird, has color, green).

[0067] S3: Utilize the graph neural network model to embed and encode the proposition element triple representation, obtain the proposition vector containing the context semantics and graph structure features, and extract the neighborhood feature vector of the associated entity in the initial knowledge graph; "embedding and encoding" is simply to convert the information in the form of proposition element triple into a vector form that is easier for computers to process, while retaining the context semantics and graph structure features. The specific process is, first, based on the pre-trained language model, analyze the proposition element triple representation to obtain the semantic embedding of entities and relationships, and then combine the pre-defined structure features of all entities in the graph structure prior to generate the initial embedding. For example, for the triple of the proposition "Apple is a fruit" (Apple, is, Fruit), after analysis by the pre-trained language model, the representations of "Apple", "is", and "Fruit" in the semantic space are obtained, and then combined with the graph structure prior knowledge to generate the initial embedding. Then, through a double-attention graph convolution layer, the feature message vector of all entity nodes in the proposition subgraph is obtained by aggregating the corresponding entity node features and all node features of all entities in the graph structure. For example, in a simple knowledge graph subgraph, the "Apple" node and the related "Fruit", "Red", and other nodes are processed by the double-attention graph convolution layer to aggregate the node features and obtain the feature message vector. Taking the average of these feature message vectors, i.e., taking the average, obtains the proposition vector containing the context semantics and graph structure features. At the same time, the K-order neighborhood features of all associated entities of all entities in the initial knowledge graph are extracted. Here, "K-order neighborhood features" means the number of hops from the entity in the proposition to be corrected, such as K=1, which is the feature of the neighborhood entity directly connected to the entity; K=2, which is the feature of the next layer of neighborhood entities connected to the directly connected neighborhood entities. Through multi-head attention mechanism, the neighborhood features are weighted and aggregated to obtain the neighborhood feature vector of the associated entity in the initial knowledge graph, which is simply to combine the neighborhood features according to their importance.

[0068] S4: Based on the neighborhood feature vector of the associated entity in the initial knowledge graph, the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph is calculated, and the logical contradiction, knowledge error and information missing in the proposition are detected by combining the preset logical conflict rule and statistical learning method. Here, a formula is used to calculate the similarity score between the proposition vector and a single associated entity node in the corresponding subgraph in the knowledge graph. After calculating the similarity score with the single associated entity node by this formula, the neighborhood entity node set of the associated entity is obtained, and the similarity score between the proposition vector and each entity node in the set is calculated based on the neighborhood feature vector of each entity node in the set by using the above formula. Then another formula is used to calculate the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph. Through this formula, the structural similarity is obtained by comprehensively considering the similarity between each node in the subgraph and the proposition vector. Finally, the logical contradiction, knowledge error and information missing in the proposition are detected by combining the preset logical conflict rule and statistical learning method. For example, the preset rule stipulates that "the sum of the internal angles of a triangle is 180 degrees", and if the proposition appears "the sum of the internal angles of a certain triangle is 200 degrees", the logical contradiction can be detected by combining the structural similarity; for the detection of knowledge error and information missing, the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph is spliced into the proposition feature vector, and the knowledge error detection classifier and the information missing detection classifier are input into the pre-trained knowledge error detection classifier and the information missing detection classifier to determine whether there is knowledge error and information missing.

[0069] S5: The integrated learning algorithm is used to trace the classification of logical contradiction, knowledge error and information missing in the knowledge proposition to be corrected, and the related evidence collected from multiple sources of knowledge data is used to make correction decision for the knowledge proposition to be corrected, so as to obtain the correction result. The "integrated learning algorithm" is a method of integrating the results of multiple base learning algorithms to improve the classification accuracy. Here, multiple base learning algorithms are used to trace the classification of these errors in the proposition, and multiple error types of the knowledge proposition to be corrected are obtained. For example, a proposition may be judged by base learning algorithm A as a knowledge error, and the error type is concept confusion; base learning algorithm B judges it as a logical contradiction, and the error type is causal relationship error. Then, based on the integrated learning algorithm, the final error type of the knowledge proposition to be corrected is obtained. The related evidence collected from multiple sources of knowledge data is used to make correction decision for the knowledge proposition to be corrected, so as to obtain the correction result. For example, in order to judge the error of the proposition "whale is a fish", the classification information of whale collected from biological textbooks, academic research papers and other multiple sources of data is used as evidence, and according to the final error type (here, it is a knowledge error, and whale is actually a mammal), it is determined that the proposition is corrected to "whale is a mammal".

[0070] In an alternative embodiment, it further comprises:

[0071] S6: Verify the error correction results. If the verification passes, update the initial knowledge graph based on relevant evidence. Verification includes consistency verification, rationality verification, and validity verification.

[0072] Consistency verification checks whether the corrected proposition is consistent with the existing knowledge system, logical rules, and other relevant information in the knowledge graph. In other words, the corrected content cannot conflict with existing correct knowledge. Suppose the knowledge graph clearly states that "all mammals are viviparous and lactate." For the incorrect proposition "whales are fish," after correction, the result is "whales are mammals." At this point, consistency verification is performed to confirm that the new proposition "whales are mammals" does not conflict with the definition of mammals and the characteristics of other mammals in the knowledge graph. If the knowledge graph also records "dolphins are mammals and are viviparous," then the proposition "whales are mammals" aligns with existing information on key information such as the reproductive method of mammals, thus passing consistency verification.

[0073] Reasonability verification involves assessing the rationality of the correction results based on common sense, actual conditions, and prevailing knowledge within the field. Even if a proposition doesn't logically conflict with existing knowledge, it may not be reasonable from a common sense or practical perspective. For example, suppose the incorrect proposition "There is a large amount of liquid water on the moon" is corrected to "There is a very small amount of water in solid form on the moon." When conducting a rationality verification, based on our general understanding of the lunar environment, the lunar surface temperature, gravity, and other conditions make it difficult for liquid water to exist in large quantities, but solid water (ice) can exist in small amounts in certain areas. Therefore, this correction is reasonable from the perspective of actual conditions and prevailing knowledge and passes the rationality verification. However, if the correction were to become "There are large amounts of oil resources on the moon," given the moon's formation and geological conditions, this would be seriously inconsistent with our current understanding of the moon and would fail the rationality verification.

[0074] Validation is the process of confirming whether the correction effectively addresses the error in the original proposition, adds valuable and accurate information to the knowledge proposition, and improves its usability in practical applications. For example, the original proposition, "The sun revolves around the earth," is incorrect; after correction, it becomes "The earth revolves around the sun." Validation involves checking whether the corrected proposition truly corrects the original error and is accurate and applicable in practical applications such as astronomical teaching, popular science, and scientific research, helping people correctly understand the operational relationships of celestial bodies in the solar system. This verifies the validity of the proposition. Simply changing the proposition to "The sun does not revolve around the earth" corrects the original error, but fails to clarify the correct operational relationship. The information provided is insufficient for practical applications and may fail validation.

[0075] The verification of the error correction result can ensure the accuracy and reliability of the error correction, and avoid introducing new errors or unreasonable content.

[0076] In an alternative embodiment, S1: based on the entity linking method, the relationship extraction method, and the attribute alignment method, effective entities, relationships, and attribute information in the multi-source knowledge data are extracted, and an initial knowledge graph is constructed based on the effective entities, relationships, and attribute information in the multi-source knowledge data, including:

[0077] HTML tags, special characters, and duplicate information in the multi-source knowledge data are removed to obtain preprocessed multi-source knowledge data. The source knowledge data is widely sourced and can contain a large amount of useless information. The HTML tag is a mark used for web display format, such as in a description "The Apples are rich in vitamins " in a description "The Tags have no substantial effect on extracting knowledge and need to be removed. Special characters such as some symbols used for separation and decoration but not carrying core knowledge also need to be removed. Repetitive information will cause resource waste and information redundancy and also need to be removed. After these processes, the pre-processed multi-source knowledge data that is more concise and only retains core knowledge content is obtained.

[0078] All entities in the pre-processed multi-source knowledge data are identified based on predefined rules, and entity attribute features and context semantic features of each entity in the pre-processed multi-source knowledge data are learned by using a supervised learning algorithm. These predefined rules are usually set according to morphological and syntactic rules. For example, a noun phrase can be an entity, and entities such as "Eiffel Tower in Paris" and "Huawei Technology Co., Ltd." can be identified. After the entities are identified, the supervised learning algorithm is used to learn the entity attribute features and context semantic features of each entity. Taking the entity "car" as an example, the supervised learning algorithm is used to learn its attribute features such as color (which can be black, white, etc.) and function (used for transportation) from a large amount of text containing "car", and the context semantic features are obtained by analyzing the context relationship of "car" in the sentence, such as "I go to work by car every day", which can understand the relationship between "car" and "work" as a means of transportation and clarify the semantic information of "car" as a means of transportation.

[0079] All entities in the pre-processed multi-source knowledge data are linked and disambiguated based on multi-dimensional matching and linking methods, entity attribute features and context semantic features of all entities, and all valid entities in the pre-processed multi-source knowledge data are obtained. Multi-dimensional matching and linking involves name matching, attribute matching, semantic matching and other methods. For example, the word "apple" is ambiguous and can refer to the fruit "apple" or "Apple Inc.". By analyzing the entity attribute features, the fruit apple has the attribute of being edible, while Apple Inc. has the attribute of commercial operation; combined with the context semantic features, such as "I bought several apples at the supermarket", it can be determined that "apple" here refers to the fruit according to the context. Through these methods, entities with the same meaning are linked and disambiguated, so that all valid entities in the pre-processed multi-source knowledge data are obtained.

[0080] The entity attribute features of all disambiguated entities in the pre-processed multi-source knowledge data are aligned, and the attribute information of all valid entities is obtained. The entity attribute features of all disambiguated entities in the pre-processed multi-source knowledge data are aligned to obtain accurate and consistent attribute information of all valid entities. For example, different data sources may have different expressions or accuracies for the radius attribute of the entity "Earth". Through attribute alignment, the description and standard of these attributes are unified to ensure that the attribute information of each valid entity is accurate and consistent.

[0081] Based on the relationship extraction method, the co-occurrence context of each pair of effective entities in the pre-processed multi-source knowledge data is analyzed, and the relationship between all effective entities is recognized and obtained; for example, in the sentence "Xiaoming is learning in school", "Xiaoming" and "school" are two effective entities, and by analyzing the context in which they co-occur, it can be recognized that there is a relationship of "learning in" between "Xiaoming" and "school".

[0082] Based on all effective entities and corresponding relationship and attribute information in the pre-processed multi-source knowledge data, an initial knowledge graph is constructed. Taking entities as nodes, such as "Xiaoming", "school" as different nodes, taking the relationship between entities as edges, i.e. "learning in" as the edge connecting "Xiaoming" and "school", and taking the attribute information of each entity as the characteristics of the node, such as the age, gender, etc. attributes of "Xiaoming", the address, scale, etc. attributes of "school", an initial knowledge graph is constructed, which can intuitively display the relationship and attributes between entities, and provide a rich knowledge base for subsequent knowledge proposition error correction.

[0083] In an alternative embodiment, S2: using a deep learning-based semantic parsing method to perform syntax analysis, semantic role labeling and logical structure extraction on the knowledge proposition to be corrected, obtaining the entity, relationship and attribute information in the knowledge proposition, and generating proposition element triple representation based on the entity, relationship and attribute information in the knowledge proposition, including:

[0084] Based on the deep learning method, a semantic parsing model is constructed, and based on the semantic parsing model, the syntax structure, semantic role, and logical structure in the knowledge proposition to be corrected are recognized; for example, using a neural network architecture, a large amount of text data containing various syntax structures, semantic relationships and logic are trained, and the model learns the rules and patterns of language expression. Based on this trained semantic parsing model, the syntax structure, semantic role and logical structure in the knowledge proposition to be corrected can be recognized. Syntax structure involves the composition of a sentence, such as the arrangement of subject-predicate-object, state-determining-complement, etc.; semantic role represents the role of each component in the sentence at the semantic level, such as agent, patient, etc.; logical structure relates to the logical relationship within the proposition, such as cause and effect, parallelism, transition, etc.

[0085] The part-of-speech standard result of the knowledge proposition to be corrected and the syntactic dependency analysis result of the knowledge proposition to be corrected are determined based on the syntactic structure in the knowledge proposition to be corrected. The part-of-speech standard result clearly indicates the part of speech of each word, such as noun, verb, adjective, etc. Taking the sentence "The bird is happily flying in the sky" as an example, through analysis of the syntactic structure, "bird" is a noun and "fly" is a verb. The syntactic dependency analysis result shows the dependency relationship between the words in the sentence, indicating which word depends on which word and the relationship type between them. For example, in the above sentence, the action "fly" depends on "bird", there is a "subject-predicate" dependency relationship, and "in the sky" indicates the location of "fly", there is a "dynamic state" dependency relationship.

[0086] Based on the syntactic dependency analysis result of the knowledge proposition to be corrected and the attention mechanism, the implicit relationship in the knowledge proposition to be corrected is identified. Based on the syntactic dependency analysis result, the attention mechanism helps the model focus on key information, thereby discovering relationships that are not directly expressed but implied in the sentence. For example, in the sentence "Xiaoming gives a book to Xiaohong, Xiaohong is very happy", through syntactic dependency analysis, it can be known that the action "give" involves "Xiaoming", "Xiaohong" and "book", and the attention mechanism will guide the model to pay attention to the possible causal relationship between "Xiaohong is very happy" and "Xiaoming gives a book to Xiaohong", that is, because Xiaoming gave a book to Xiaohong, Xiaohong is very happy, and this causal relationship is an implicit relationship.

[0087] All entities and corresponding attribute information in the knowledge proposition to be corrected are identified based on the part-of-speech tagging result and semantic role, and the relationship between all entities in the knowledge proposition to be corrected is identified based on the implicit relationship and logical structure. From part-of-speech tagging and semantic role, nouns can usually be used as entities, for example, in "beautiful flowers emit charming fragrance", "flowers" is an entity. At the same time, adjectives and other modifiers of nouns can be used as attribute information, here "beautiful" and "charming" are attributes of "flowers". Combined with the previously identified implicit relationship and logical structure, the relationship between entities is clear, such as the relationship between "flowers" and "fragrance" exists "emitting".

[0088] Based on all entities and corresponding attribute information in the knowledge proposition to be corrected and the relationship between all entities, a proposition element triple representation is generated. The proposition element triple is generally in the form of (entity, relationship, entity / attribute). Taking the above example, proposition element triples such as (flowers, has attribute, beautiful) and (flowers, emits, fragrance) can be generated, clearly and concisely extracting the key information in the knowledge proposition, providing convenience for subsequent knowledge proposition correction processing.

[0089] In an alternative implementation, S3: embedding coding is performed on the proposition element triple representation using a graph neural network model to obtain a proposition vector containing context semantics and graph structure features, and the neighborhood feature vector of the associated entity in the initial knowledge graph is extracted, including:

[0090] Based on the pre-trained language model, the semantic embedding of the entity and the relationship of the knowledge proposition to be corrected is obtained by analyzing the proposition element triple representation, and the initial embedding of the knowledge proposition to be corrected is generated by combining the pre-defined structure features of all entities in the graph structure prior in the knowledge proposition to be corrected. For example, the common BERT model can perform semantic analysis on each element in the proposition element triple, such as entities and relationships, to obtain their embedding representation in the semantic space, that is, to convert entities and relationships into numerical vectors, which can reflect their semantic features. In combination with the pre-defined structure features of all entities in the graph structure prior in the knowledge proposition to be corrected, the initial embedding of the knowledge proposition to be corrected is generated. The graph structure prior here refers to some feature information about the entity in the graph structure that is set in advance, such as the position of the entity in the graph, the connection mode with other entities, etc. The semantic embedding obtained by the pre-trained language model is combined with these pre-defined structure features to generate an initial embedding containing semantic and structural information. For example, for the proposition "apple is fruit", the pre-trained language model analyzes the semantic embedding of "apple" and "fruit", and then combines the structure features such as the positions of "apple" and "fruit" in the graph in the graph structure prior to obtain the initial embedding of the proposition.

[0091] Based on the double-attention graph convolution layer and the initial embedding of the knowledge proposition to be corrected, the feature message vector of all entity nodes in the proposition subgraph is obtained by aggregating each entity node feature corresponding to all entities in the graph structure and all node features corresponding to all entities. The double-attention graph convolution layer aggregates each entity node feature corresponding to all entities in the graph structure and all node features corresponding to all entities. For example, in a simple knowledge graph structure, the entity node "apple" not only has its own features, but also is connected to other nodes (such as "red", "round", etc. representing attribute nodes, and "fruit" representing category relationship nodes), the double-attention graph convolution layer will simultaneously focus on the features of each entity node itself and the features of all nodes connected thereto, and aggregate these features together to obtain the feature message vector of all entity nodes in the proposition subgraph. This feature message vector integrates the information of the entity node and its surrounding nodes, and more comprehensively reflects the semantic and structural relationship of the entity in the graph.

[0092] The feature message vectors of all entity nodes in the proposition subgraph are mean-pooled to obtain a proposition vector; that is, by calculating the average of the feature message vectors, multiple feature values are compressed into one value to obtain the proposition vector. The purpose of this is to further refine the complex feature information of all entity nodes in the proposition subgraph to form a vector that can represent the entire proposition. This proposition vector contains context semantics and graph structure features, facilitating subsequent comparison and analysis with the knowledge graph.

[0093] The K-order neighborhood features of all associated entities of the entities in the to-be-corrected knowledge proposition in the initial knowledge graph are extracted, and the K-order neighborhood features of all associated entities are weighted and aggregated by a multi-head attention mechanism to obtain a neighborhood feature vector of the associated entities in the initial knowledge graph.

[0094] The K-order neighborhood feature here refers to the feature of the neighborhood entity with a distance of K steps from the entity in the to-be-corrected knowledge proposition. For example, when K = 1, it is the feature of the neighborhood entity directly connected to the entity; when K = 2, it is the feature of the next layer of neighborhood entities connected to the directly connected neighborhood entities.

[0095] The multi-head attention mechanism is a mechanism that can simultaneously focus on different aspects of information. It assigns different weights to each K-order neighborhood feature, and then aggregates these weighted neighborhood features to obtain a neighborhood feature vector of the associated entities in the initial knowledge graph. The neighborhood feature vector obtained in this way comprehensively considers the features of neighborhood entities at different distances and is weighted according to their importance, more accurately reflecting the local structure and semantic information of the associated entities in the knowledge graph, providing an important basis for subsequent detection of errors in the proposition.

[0096] In an alternative embodiment, S4: based on the neighborhood feature vector of the associated entities in the initial knowledge graph, the structure similarity between the proposition vector and the corresponding subgraph in the knowledge graph is calculated, and combined with the preset logical conflict rules and statistical learning methods, the logical contradictions, knowledge errors, and information omissions in the proposition are detected, including:

[0097] Based on the neighborhood features of the associated entities in the initial knowledge graph and the proposition vector, the similarity score of the proposition vector and a single associated entity node in the corresponding subgraph in the knowledge graph is calculated:

[0098] ;

[0099] In the formula, is the similarity score of the proposition vector and a single associated entity node in the corresponding subgraph in the knowledge graph, is used to convert the input numerical vector into a probability distribution vector, is a weight matrix, which is used to map the concatenated feature to a similarity score, is a single related entity node is a domain feature vector of the single related entity node, is a vector concatenation operation, is a bias term, which is used to add a fixed value to the calculation result, which helps to adjust the range and accuracy of the similarity score, for example, take 0.1;

[0100] At the beginning, the elements of the weight matrix W can be randomly sampled from a certain distribution (such as normal distribution, uniform distribution). For example, randomly generate matrix elements from a normal distribution with mean 0 and standard deviation 0.1. This is done because at the beginning of model training, we know little about the internal relationship of the data, and random initialization can allow the model to explore the parameter space widely.

[0101] In the subsequent training process, the weight matrix is adjusted according to the gradient of the loss function through the back propagation algorithm. The loss function measures the difference between the model's predicted results (such as the similarity score of the proposition vector and the subgraph of the knowledge graph) and the true situation (if there is a correct similarity score marked). In our previous example of "cat is a mammal", if the similarity score calculated by the model deviates greatly from the expected correct score, the back propagation algorithm will calculate the contribution of each weight to the loss function (i.e. gradient). For example, if a certain weight makes the similarity score deviate from the correct value, then according to the gradient information, the weight is adjusted using an optimizer (such as stochastic gradient descent SGD, Adagrad, Adam, etc.) to make the similarity score closer to the correct value. After multiple iterations of training, the weight matrix is gradually optimized to better adapt to the characteristics of the data.

[0102] Obtain the neighborhood entity node set of the related entity, and based on the neighborhood feature vector of each entity node in the neighborhood entity node set of the related entity, calculate the similarity score of the proposition vector and each entity node in the neighborhood entity node set of the related entity using the above formula;

[0103] Based on the similarity score of the proposition vector and each entity node in the single related entity node or neighborhood entity node set in the corresponding subgraph in the knowledge graph, calculate the structural similarity of the proposition vector and the corresponding subgraph in the knowledge graph, and obtain a value that reflects the similarity degree of the proposition vector and the entire subgraph structure.

[0104] Based on the proposition vector and the structural similarity of the corresponding subgraph in the knowledge graph, and combined with the preset logical conflict rules and statistical learning methods, detect the logical contradiction, knowledge error, and information missing in the proposition.

[0105] Preset logical conflict rules are pre-set logical judgment criteria. For example, consider the universal cognitive logic of "all birds can fly." If a proposition states, "Ostriches can't fly, and ostriches are birds, but all birds can fly," the logical contradiction can be detected by comparing the structural similarity with the knowledge graph and applying this pre-set rule. Statistical learning methods utilize a large amount of existing correct proposition data for statistical analysis, building models to determine whether the proposition contains knowledge errors or missing information. For example, statistics show that most propositions about fruit mention some common attributes of fruit. If a proposition about fruit fails to include these common attributes, it can be detected as missing information.

[0106] The above scheme uses the initial knowledge graph and proposition vector to detect various errors in propositions by calculating similarity and combining specific rules and methods.

[0107] In an alternative embodiment, the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph is calculated based on the similarity score between the proposition vector and a single associated entity node or each entity node in a set of neighborhood entity nodes in the corresponding subgraph in the knowledge graph, including:

[0108] ;

[0109] Where, is the proposition vector Corresponding subgraph in the knowledge graph The structural similarity of is the total number of associated entity nodes in the corresponding subgraph, is the set of associated entity nodes in the subgraph, is a hyperparameter and , is the associated entity node in the corresponding subgraph The set of neighborhood entity nodes, is the proposition vector and the associated entity node in the corresponding subgraph Entity nodes in the neighborhood entity node set The similarity score of is the associated entity node in the corresponding subgraph The total number of entity nodes in the neighborhood entity node set.

[0110] Through such calculations, we obtain the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph. This value can help us further judge the degree of match between the proposition and the knowledge graph structure, and then detect logical contradictions, knowledge errors, information missing and other problems in the proposition.

[0111] In an alternative embodiment, based on the structural similarity of the proposition vector and the corresponding subgraph in the knowledge graph, and combined with the preset logical conflict rules and statistical learning methods, logical contradictions, knowledge errors, and information missing in the proposition are detected, including:

[0112] Based on the preset logical conflict rules, all proposition element mutual exclusion relationships are screened out in the proposition element triple representation; for example, in the animal knowledge field, we know that "penguins cannot fly" is a general cognition, so in the preset logical conflict rules, it is possible to set that there is an exclusion relationship between "penguins" and "can fly". For the given proposition "penguins can fly, penguins are birds", it is converted into proposition element triple representation, which may be (penguins, can, fly), (penguins, is, birds). Based on the preset logical conflict rules, it is screened out that the element (penguins, can, fly) has an exclusion relationship, because according to the preset, penguins should not be able to fly.

[0113] Based on all proposition element mutual exclusion relationships and the structural similarity of the proposition vector and the corresponding subgraph in the knowledge graph, logical contradictions in the proposition are detected; for example: by comprehensively judging the confidence of all proposition element mutual exclusion relationships and the structural similarity of the corresponding subgraph in the knowledge graph, whether all proposition element mutual exclusion relationships are established is determined, and then logical contradictions in the proposition are determined.

[0114] The proposition vector and the structural similarity of the corresponding subgraph in the knowledge graph are spliced with the proposition element triple to obtain a proposition feature vector;

[0115] Based on the proposition feature vector and the knowledge error detection classifier and the information missing detection classifier, knowledge errors and information missing in the proposition are detected.

[0116] The knowledge error detection classifier and the information missing detection classifier are models trained based on a large number of labeled correct and incorrect proposition data through statistical learning methods.

[0117] The constructed proposition feature vector is input into the knowledge error detection classifier, which judges whether the proposition has a knowledge error according to the pattern learned by it. For example, if in the training data, propositions similar to "penguins can fly" are mostly labeled as incorrect, and the proposition feature vector has similarity with the feature vectors of these incorrect propositions, the classifier may judge that the proposition has a knowledge error.

[0118] For information missing detection, it is assumed that in the training data, propositions that completely describe birds usually contain features corresponding to key information such as "have feathers" and "oviparity". Since the current proposition feature vector does not reflect the features of these key information, the information missing detection classifier may judge that the proposition has an information missing problem, such as the description "penguins are birds" does not mention some typical feature information of birds.

[0119] Through the above steps, the combination of the proposition vector, the knowledge graph subgraph structure similarity, the preset logical conflict rule, and the classifier trained by the statistical learning method can comprehensively detect logical contradictions, knowledge errors, and information missing in the proposition.

[0120] In an alternative embodiment, S5: using an ensemble learning algorithm to trace the classification of logical contradictions, knowledge errors, and information missing in the knowledge proposition to be corrected, and combining the relevant evidence collected from multiple sources of knowledge data to make a correction decision for the knowledge proposition to be corrected, to obtain a correction result, including:

[0121] Based on multiple base learners, the logical contradictions, knowledge errors, and information missing in the proposition are classified to obtain multiple error types of the knowledge proposition to be corrected. Multiple base learners can be understood as multiple different and relatively simple learning models, and each base learner is good at analyzing problems from different angles. For example, the first base learner may focus on the chronological order of historical events, the second base learner may focus on the correspondence between characters and events, and the third base learner may focus on the matching degree of historical facts and general cognition. The knowledge proposition to be corrected "Tang Taizong Li Shimin is the founding emperor of the Tang Dynasty, and the Tang Dynasty was established in 960 AD” is input into these base learners.

[0122] The base learner focusing on the chronological order, through learning and analysis of the historical timeline, finds that the Tang Dynasty was actually established in 618 AD, not 960 AD, so it judges that the proposition has a knowledge error, and the error type is "time error”.

[0123] The base learner focusing on the correspondence between characters and events knows that the founding emperor of the Tang Dynasty is Li Yuan, not Li Shimin, according to its learning of historical figures and events in the Tang Dynasty, so it judges that the proposition has a logical contradiction, and the error type is "characters and events do not correspond”.

[0124] The base learner focusing on the matching degree of historical facts and general cognition also identifies the above two errors, which are classified as "time knowledge error” and "character identity logical contradiction” respectively.

[0125] Through the analysis of multiple base learners, we obtain multiple error types of the knowledge proposition to be corrected, such as "time error”, "characters and events do not correspond” and the like.

[0126] An ensemble learning algorithm is used to perform ensemble learning on multiple error types in the knowledge proposition to be corrected, obtaining the final error type for the knowledge proposition to be corrected. This algorithm integrates the judgment results of multiple base learners. Common ensemble learning methods include voting and averaging. For example, suppose there are five base learners. Three of them believe that the part "Emperor Taizong of Tang, Li Shimin, is the founding emperor of the Tang Dynasty" is a "personality logical contradiction" error, while two base learners believe it is a related error type. Based on the voting results, "personality logical contradiction" receives the most votes. After integration by the ensemble learning algorithm, the final error types for the knowledge proposition to be corrected are "personality logical contradiction" and "time knowledge error." This approach, combining the judgments of multiple base learners, improves the accuracy and reliability of error type judgments and avoids the potential misjudgment by a single base learner.

[0127] Based on the final error type of the knowledge proposition to be corrected and combined with relevant evidence collected from multi-source knowledge data, a correction decision is made on the knowledge proposition to be corrected to obtain the correction result.

[0128] Evidence related to the proposition is collected from multi-source knowledge data, such as authoritative historical books, academic research papers, and professional historical databases. These materials all indicate that the founding emperor of the Tang Dynasty was Li Yuan, and that the Tang Dynasty was established in 618 AD. Based on the final error type and the collected evidence, the knowledge proposition to be corrected is corrected. For "Emperor Taizong of Tang, Li Shimin, is the founding emperor of the Tang Dynasty," it is corrected to "Li Yuan is the founding emperor of the Tang Dynasty" based on the evidence; for "The Tang Dynasty was established in 960 AD," it is corrected to "The Tang Dynasty was established in 618 AD." After correction, the correct knowledge proposition is obtained as "Li Yuan is the founding emperor of the Tang Dynasty, and the Tang Dynasty was established in 618 AD," completing the correction process for the knowledge proposition to be corrected.

[0129] Through the above three steps, the use of integrated learning algorithms and the combination of evidence from multi-source knowledge data can accurately trace and classify the logical contradictions, knowledge errors, and information missing in the error correction knowledge propositions, make reasonable error correction decisions, obtain accurate error correction results, and improve the accuracy of knowledge propositions.

[0130] refer to Figure 2 The present invention provides an implementation of a knowledge proposition correction system based on knowledge graph optimization and upgrading, including:

[0131] The knowledge graph construction module is used to extract valid entities, relationships, and attribute information from multi-source knowledge data based on entity linking methods, relationship extraction methods, and attribute alignment methods, and to construct an initial knowledge graph based on the valid entities, relationships, and attribute information in the multi-source knowledge data;

[0132] The triple generation module is configured to perform syntax analysis, semantic role labeling and logical structure extraction on the knowledge proposition to be corrected by using a deep learning-based semantic parsing method, to obtain entity, relationship and attribute information in the knowledge proposition, and to generate a proposition element triple representation based on the entity, relationship and attribute information in the knowledge proposition;

[0133] The neighborhood feature extraction module is configured to embed and encode the proposition element triple representation by using a graph neural network model to obtain a proposition vector containing context semantics and graph structure features, and to extract a neighborhood feature vector of an associated entity in the initial knowledge graph;

[0134] The error detection module is configured to calculate a structural similarity between the proposition vector and a corresponding subgraph in the knowledge graph based on the neighborhood feature vector of the associated entity in the initial knowledge graph, and to detect logical contradictions, knowledge errors and information omissions in the proposition by combining a preset logical conflict rule and a statistical learning method.

[0135] The proposition correction module is configured to trace and classify logical contradictions, knowledge errors and information omissions in the knowledge proposition to be corrected by using an ensemble learning algorithm, and to make a correction decision on the knowledge proposition to be corrected by combining relevant evidence collected from multiple sources of knowledge data, to obtain a correction result.

[0136] The knowledge graph construction module in the system extracts effective information from multi-source knowledge data and constructs an initial knowledge graph by means of entity linking, relation extraction and attribute alignment, which makes the system have a rich and comprehensive knowledge base and can provide extensive data support for subsequent error correction, reducing the limitations caused by single data. The triple generation module uses a deep learning-based semantic parsing method to perform deep analysis of the syntax, semantics and logical structure of the knowledge proposition to be corrected, and then generates a triple representation of the proposition elements, accurately extracts the key information of the proposition, and clearly presents the internal structure of the proposition, providing a good start for accurately finding problems. The neighborhood feature extraction module uses a graph neural network model to embed and encode the proposition element triples, obtains a proposition vector that fuses the context semantic and graph structure features, and extracts the neighborhood feature vector of the associated entities in the knowledge graph. This fusion processing effectively combines the proposition semantic and knowledge graph structure information, significantly improving the accuracy of error detection. The error detection module calculates the structural similarity between the proposition vector and the corresponding subgraph of the knowledge graph based on the above neighborhood feature vector, and combines pre-set logical conflict rules and statistical learning methods to comprehensively and meticulously detect various problems such as logical contradictions, knowledge errors and information omissions in the proposition from multiple dimensions. The final proposition correction module uses an ensemble learning algorithm to trace and classify the detected problems, identifies the root cause of the error, and makes scientific and reasonable correction decisions based on the relevant evidence collected from multi-source knowledge data, finally obtaining accurate correction results, effectively improving the accuracy and reliability of the knowledge proposition, and ensuring the high-quality presentation and application of knowledge.

[0137] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalent technology, the present application also intends to include these modifications and variations.

Claims

1. A knowledge proposition error correction method based on knowledge graph optimization and upgrading, characterized in that: include: S1: Extract valid entities, relationships, and attribute information from multi-source knowledge data based on entity linking, relationship extraction, and attribute alignment methods, and construct an initial knowledge graph based on the valid entities, relationships, and attribute information from multi-source knowledge data; S2: Use a deep learning-based semantic parsing method to perform grammatical analysis, semantic role labeling, and logical structure extraction on the knowledge proposition to be corrected, obtain the entity, relationship, and attribute information in the knowledge proposition, and generate a triple representation of the proposition elements based on the entity, relationship, and attribute information in the knowledge proposition, including: Build a semantic parsing model based on deep learning methods, and identify the grammatical structure, semantic role, and logical structure of the knowledge proposition to be corrected based on the semantic parsing model; Determine the part of speech standard result and syntactic dependency analysis result of the knowledge proposition to be corrected based on the grammatical structure in the knowledge proposition to be corrected; Based on the syntactic dependency analysis results and attention mechanism of the knowledge proposition to be corrected, the implicit relationship in the knowledge proposition to be corrected is identified; Based on the part-of-speech tagging results and semantic roles of the knowledge proposition to be corrected, all entities and corresponding attribute information in the knowledge proposition to be corrected are identified, and based on implicit relations and logical structures, the relationships between all entities in the knowledge proposition to be corrected are identified; Generate proposition element triple representation based on all entities and corresponding attribute information in the knowledge proposition to be corrected and the relationship between all entities; S3: Use the graph neural network model to embed the triple representation of proposition elements to obtain the proposition vector containing contextual semantics and graph structure features, and extract the neighborhood feature vector of the associated entity in the initial knowledge graph; S4: Based on the neighborhood feature vectors of the associated entities in the initial knowledge graph, the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph is calculated. In combination with the preset logical conflict rules and statistical learning methods, logical contradictions, knowledge errors, and information missing in the proposition are detected; S5: Use an integrated learning algorithm to trace and classify the logical contradictions, knowledge errors, and information missing in the error-correcting knowledge propositions, and combine relevant evidence collected from multi-source knowledge data to make correction decisions on the error-correcting knowledge propositions and obtain correction results.

2. The knowledge proposition error correction method based on knowledge graph optimization and upgrading according to claim 1 is characterized in that: Also includes: S6: Verify the error correction results. If the verification passes, update the initial knowledge graph based on relevant evidence. Verification includes consistency verification, rationality verification, and validity verification.

3. The knowledge proposition error correction method based on knowledge graph optimization and upgrading according to claim 1 is characterized in that: S1: Extract valid entities, relationships, and attribute information from multi-source knowledge data based on entity linking, relationship extraction, and attribute alignment methods, and construct an initial knowledge graph based on the valid entities, relationships, and attribute information from multi-source knowledge data, including: Remove HTML tags, special words, and repeated information from multi-source knowledge data to obtain pre-processed multi-source knowledge data; Identify all entities in the preprocessed multi-source knowledge data based on predefined rules, and use supervised learning algorithms to learn entity attribute features and contextual semantic features of each entity in the preprocessed multi-source knowledge data; Based on the multi-dimensional matching linking method, the entity attribute characteristics and contextual semantic characteristics of all entities, all entities in the pre-processed multi-source knowledge data are linked and disambiguated to obtain all valid entities in the pre-processed multi-source knowledge data; Perform attribute alignment on the entity attribute features of all disambiguated entities in the preprocessed multi-source knowledge data to obtain the attribute information of all valid entities; Based on the relationship extraction method, the coexistence context of two valid entities in the pre-processed multi-source knowledge data is analyzed to identify the relationship between all valid entities; Construct the initial knowledge graph based on all valid entities, corresponding relationships and attribute information in the preprocessed multi-source knowledge data.

4. The knowledge proposition error correction method based on knowledge graph optimization and upgrading according to claim 1 is characterized in that: S3: Use the graph neural network model to embed the triple representation of proposition elements to obtain a proposition vector that contains contextual semantics and graph structure features. At the same time, extract the neighborhood feature vectors of related entities in the initial knowledge graph, including: Based on the pre-trained language model, the triple representation of proposition elements is analyzed to obtain the semantic embedding of the entities and relations of the knowledge proposition to be corrected. The initial embedding of the knowledge proposition to be corrected is generated by combining the predefined structural features of all entities in the knowledge proposition to be corrected in the graph structure prior. Based on the dual-attention graph convolution layer and the initial embedding of the knowledge proposition to be corrected, the corresponding entity node features and all corresponding node features of all entities in the knowledge proposition to be corrected in the graph structure are aggregated to obtain the feature message vectors of all entity nodes in the proposition subgraph; Perform mean pooling on the feature message vectors of all entity nodes in the proposition subgraph to obtain the proposition vector; The K-order neighborhood features of all associated entities in the knowledge proposition to be corrected in the initial knowledge graph are extracted, and the K-order neighborhood features corresponding to all associated entities are weightedly aggregated through the multi-head attention mechanism to obtain the neighborhood feature vector of the associated entities in the initial knowledge graph.

5. The knowledge proposition error correction method based on knowledge graph optimization and upgrading according to claim 1 is characterized in that: S4: Based on the neighborhood feature vectors of the associated entities in the initial knowledge graph, the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph is calculated. Combined with the preset logical conflict rules and statistical learning methods, logical contradictions, knowledge errors, and information missing in the proposition are detected, including: Based on the neighborhood features and proposition vectors of the associated entities in the initial knowledge graph, the similarity score between the proposition vector and the single associated entity node in the corresponding subgraph in the knowledge graph is calculated: ; Where, is the proposition vector A single associated entity node in the corresponding subgraph in the knowledge graph The similarity score of Used to convert the input numerical vector into a probability distribution vector. is the weight matrix used to map the concatenated features to similarity scores, For a single associated entity node The domain feature vector of is the vector concatenation operation, is the bias term; Obtaining a set of neighboring entity nodes of the associated entity, and calculating a similarity score between the proposition vector and each entity node in the set of neighboring entity nodes of the associated entity based on a neighborhood feature vector of each entity node in the set of neighboring entity nodes of the associated entity; Calculate the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph based on the similarity score between the proposition vector and a single associated entity node or each entity node in the set of neighborhood entity nodes in the corresponding subgraph in the knowledge graph; Based on the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph, and combined with preset logical conflict rules and statistical learning methods, logical contradictions, knowledge errors, and information missing in the proposition are detected.

6. The knowledge proposition error correction method based on knowledge graph optimization and upgrading according to claim 5 is characterized in that: Based on the similarity score between the proposition vector and a single associated entity node in the corresponding subgraph in the knowledge graph or each entity node in the set of neighborhood entity nodes, the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph is calculated, including: ; Where, is the proposition vector Corresponding subgraph in the knowledge graph The structural similarity of is the total number of associated entity nodes in the corresponding subgraph, is the set of associated entity nodes in the corresponding subgraph, is a hyperparameter and , is the associated entity node in the corresponding subgraph The set of neighborhood entity nodes, is the proposition vector and the associated entity node in the corresponding subgraph Entity nodes in the neighborhood entity node set The similarity score of is the associated entity node in the corresponding subgraph The total number of entity nodes in the neighborhood entity node set.

7. The knowledge proposition error correction method based on knowledge graph optimization and upgrading according to claim 5 is characterized in that: Based on the structural similarity between proposition vectors and corresponding subgraphs in the knowledge graph, combined with preset logical conflict rules and statistical learning methods, logical contradictions, knowledge errors, and information missing in propositions are detected, including: Based on the preset logical conflict rules, all mutually exclusive relationships between proposition elements are screened out in the proposition element triple representation; Detect logical contradictions in propositions based on the mutual exclusion relationship between all proposition elements and the structural similarity between proposition vectors and corresponding subgraphs in the knowledge graph; The proposition feature vector is obtained by concatenating the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph with the proposition element triples; Based on the proposition feature vector, the knowledge error detection classifier and the information missing detection classifier, the knowledge error and information missing in the proposition are detected.

8. The knowledge proposition error correction method based on knowledge graph optimization and upgrading according to claim 1, The feature is that S5: an integrated learning algorithm is used to trace and classify the logical contradictions, knowledge errors, and information missing in the error-correcting knowledge proposition, and combined with relevant evidence collected from multi-source knowledge data, an error correction decision is made on the error-correcting knowledge proposition to obtain an error correction result, including: Based on multiple base learners, the logical contradictions, knowledge errors, and information missing in the propositions are traced and classified to obtain multiple error types of the knowledge propositions to be corrected; Based on the ensemble learning algorithm, multiple error types of the knowledge proposition to be corrected are subjected to ensemble learning to obtain the final error type of the knowledge proposition to be corrected; Based on the final error type of the knowledge proposition to be corrected and combined with relevant evidence collected from multi-source knowledge data, a correction decision is made on the knowledge proposition to be corrected to obtain the correction result.

9. A knowledge proposition correction system based on knowledge graph optimization and upgrading, characterized by: The method for correcting knowledge proposition errors based on knowledge graph optimization and upgrading according to any one of claims 1 to 8 comprises: The knowledge graph construction module is used to extract valid entities, relationships, and attribute information from multi-source knowledge data based on entity linking methods, relationship extraction methods, and attribute alignment methods, and to construct an initial knowledge graph based on the valid entities, relationships, and attribute information in the multi-source knowledge data; The triple generation module is used to perform grammatical analysis, semantic role labeling, and logical structure extraction on the knowledge proposition to be corrected using a deep learning-based semantic parsing method, obtain the entity, relationship, and attribute information in the knowledge proposition, and generate a triple representation of the proposition elements based on the entity, relationship, and attribute information in the knowledge proposition; The neighborhood feature extraction module is used to embed the triple representation of proposition elements using a graph neural network model to obtain a proposition vector containing contextual semantics and graph structure features, and to extract the neighborhood feature vectors of related entities in the initial knowledge graph; The error detection module is used to calculate the structural similarity between the proposition vector and the corresponding subgraph in the knowledge graph based on the neighborhood feature vectors of the associated entities in the initial knowledge graph. It also combines preset logical conflict rules and statistical learning methods to detect logical contradictions, knowledge errors, and information missing in the proposition. The proposition correction module is used to use an integrated learning algorithm to trace and classify the logical contradictions, knowledge errors, and information missing in the correction knowledge propositions, and combine the relevant evidence collected from multi-source knowledge data to make correction decisions on the correction knowledge propositions and obtain correction results.

Citation Information

Patent Citations

  • Multi-modal knowledge graph construction method

    CN112200317A

  • Knowledge graph completion method based on large language model

    CN117634604A