Knowledge graph proposition error correction method and system based on artificial intelligence

By introducing multimodal feature extraction, two-layer semantic verification and adaptive error correction fusion technology into the knowledge graph error correction method, the problem of insufficient recognition and correction capabilities of complex semantic errors in the existing technology is solved, and a more efficient and accurate knowledge graph error correction effect is achieved.

CN119940511AActive Publication Date: 2025-05-06网才科技(广州)集团股份有限公司

Patent Information

Application Number
CN202510130825.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-06
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The existing knowledge graph error correction methods lack comprehensive analysis when dealing with multi-level semantic associations, making it difficult to accurately identify and correct complex semantic errors, and are inefficient in computing efficiency and are not robust to noise data.

Method used

Using the knowledge graph proposition error correction method based on artificial intelligence, multimodal feature extraction, two-layer semantic verification and adaptive error correction fusion are adopted, including feature extraction, local semantic verification, global semantic verification and semantic error correction fusion.

Benefits of technology

It significantly improves the accuracy and usability of knowledge graphs, can more accurately identify and correct complex semantic errors, and improves the computing efficiency of processing large-scale knowledge graphs and its robustness to noise data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940511A_ABST
    Figure CN119940511A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph proposition error correction method and system based on artificial intelligence, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the feature extraction of knowledge graph data, the knowledge graph data comprises text data and structure data, and obtaining multi-modal feature data; constructing a local semantic verification model based on a graph attention network, performing local semantic consistency detection on the multi-modal feature data, and outputting a local semantic verification result; constructing a global semantic verification model based on a knowledge graph embedding technology, performing global semantic consistency detection on the multi-modal feature data, and outputting a global semantic verification result; and inputting the local semantic verification result and the global semantic verification result into a conditional random field model to generate a semantic error correction result of the knowledge graph. According to the method, the accuracy and availability of the knowledge graph are remarkably improved, reliable technical support is provided for maintenance and optimization of the knowledge graph, and the method has important practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for correcting errors in knowledge graph propositions based on artificial intelligence. Background Art

[0002] With the rapid development of artificial intelligence technology, knowledge graphs, as a structured knowledge representation method, play an important role in intelligent question answering, recommendation systems, information retrieval and other fields. However, due to factors such as noise in the knowledge acquisition process, the dynamic nature of knowledge evolution and the complexity of knowledge representation, semantic errors are inevitable in knowledge graphs. Traditional knowledge graph error correction methods mainly rely on rule matching and statistical analysis. These methods often only focus on single-dimensional features, such as only considering text semantics or graph structure features, resulting in limited error correction performance. At the same time, most existing error correction technologies adopt independent local or global verification strategies, lack comprehensive analysis of multi-level semantic associations, and are difficult to accurately identify and correct complex semantic errors in knowledge graphs. In addition, when processing large-scale knowledge graphs, traditional methods have low computational efficiency and insufficient robustness to noisy data.

[0003] With the advancement of deep learning technology, knowledge graph processing methods based on neural networks have shown good performance. However, existing knowledge graph error correction methods based on deep learning still have some problems: first, there is a lack of effective fusion mechanism for multimodal features of knowledge graphs, making it difficult to fully utilize text and structural information; second, in the process of semantic verification, existing methods often ignore the interactive relationship between local context and global knowledge constraints, resulting in insufficient accuracy and reliability of verification results; third, there is a lack of adaptive error correction strategies, making it difficult to provide accurate correction suggestions based on different types of semantic errors. Therefore, how to design a knowledge graph error correction method that can comprehensively utilize multimodal features and integrate local and global semantic verification has become an important topic of current research.

[0004] Existing knowledge graph error correction technologies have technical problems such as insufficient feature utilization, single verification strategy, and inaccurate correction suggestions. The present invention provides a knowledge graph proposition error correction method and system based on artificial intelligence, aiming to improve the accuracy and practicality of knowledge graph error correction through technical means such as multimodal feature extraction, two-layer semantic verification, and adaptive error correction fusion. Summary of the invention

[0005] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0006] In view of the above existing problems, the present invention is proposed.

[0007] Therefore, the present invention provides a knowledge graph proposition error correction method and system based on artificial intelligence, which can solve the problems mentioned in the background technology.

[0008] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a knowledge graph proposition correction method based on artificial intelligence, which comprises extracting features from knowledge graph data, wherein the knowledge graph data comprises text data and structure data, and obtaining multimodal feature data; Building a local semantic verification model based on the graph attention network, performing local semantic consistency detection on the multimodal feature data, and outputting a local semantic verification result; Building a global semantic verification model based on knowledge graph embedding technology, performing global semantic consistency detection on the multimodal feature data, and outputting a global semantic verification result; The local semantic verification results and the global semantic verification results are input into the conditional random field model to generate semantic error correction results of the knowledge graph.

[0009] As a preferred solution of the artificial intelligence-based knowledge graph proposition correction method of the present invention, wherein: the feature extraction includes: The text data in the knowledge graph is divided into head entity text, relationship text and tail entity text according to the triple structure. The calculation formula is as follows: ; Among them, Triple represents a triple in the knowledge graph, h represents the head entity, r represents the relationship, and t represents the tail entity; The head entity text, relation text and tail entity text are encoded respectively to obtain the text feature vector. The encoding formula is as follows: ; ; ; in, represents the encoding function, represents the pre-trained language model encoder; Constructing a multi-head self-attention matrix based on attention weights, calculating the semantic relevance of the text feature vector, and obtaining a semantically enhanced feature vector; The structural data in the knowledge graph is converted into an adjacency matrix, and a convolution operation is performed on the adjacency matrix based on a graph neural network to obtain a graph structure feature vector; A multi-layer perceptron is used to perform spatial mapping transformation on the graph structure feature vector to obtain a structure enhancement feature vector; Constructing a feature fusion layer to splice the semantic enhancement feature vector and the structural enhancement feature vector into a fused feature matrix; A residual connection structure is used to reorganize the features of the fused feature matrix and output multimodal feature data.

[0010] As a preferred solution of the artificial intelligence-based knowledge graph proposition correction method of the present invention, the calculation formula for obtaining the semantic enhancement feature vector is as follows: ; ; Among them, Attention represents the attention function, Q represents the query matrix, K represents the key matrix, V represents the value matrix, d represents the feature dimension, MultiHead represents the multi-head attention function, and head i represents the i-th attention head, W O Represents the output projection matrix, Concat represents the concatenation operation; The calculation formula for obtaining the structure enhancement feature vector is as follows: ; Where MLP( ) represents the multi-layer perceptron function, x represents the input feature, and represents the weight matrix, and represents the bias vector, ReLU() represents the activation function; The calculation formula of the feature recombination is as follows: ; ; in, represents the fusion feature, Semantic features, Represents structural features, || represents feature concatenation operation, LayerNorm() represents layer normalization function, SubLayer() represents sublayer network, and Output represents the final output features.

[0011] As a preferred solution of the artificial intelligence-based knowledge graph proposition correction method of the present invention, the local semantic consistency detection includes: Extracting a feature representation of a current triple to be verified from the multimodal feature data, and simultaneously obtaining a set of neighbor triples directly connected to the triple; A computing unit of a graph attention network is constructed, and an initial attention weight coefficient is assigned to each triple in the set of neighbor triples. The calculation formula for assigning the initial attention weight to the set of neighbor triples is as follows: ; in, Represents a slave node To Node The initial attention coefficient, represents the feature transformation weight matrix, and Respectively represent nodes and nodes The characteristic vector of represents the attention vector parameter, || represents the vector concatenation operation, Representation Node The set of neighbor nodes, LeakyReLU represents the activation function; Calculating, for each triple in the neighbor triple set, a semantic similarity matrix with the current triple to be verified, and generating a triple comparison score; Based on the triple comparison scores, the attention weight coefficient in the graph attention network is updated to screen out key neighbor triplets whose semantic relevance is higher than a preset threshold; Performing local semantic consistency calculation on the features of the key neighbor triplet and the features of the current triplet to be verified to obtain a local consistency feature vector; A bidirectional gated recurrent unit network is constructed, and the local consistency feature vector is input into the network for temporal dependency analysis to obtain a local semantic verification result.

[0012] As a preferred solution of the artificial intelligence-based knowledge graph proposition correction method of the present invention, the temporal dependency analysis includes: Construct a bidirectional GRU network to process local consistency features; The verification result is obtained by fusing the bidirectional features. The calculation formula is as follows:

[0013] in, represents the final local semantic verification result vector, represents the output layer weight matrix, represents the final hidden state of the forward GRU, represents the final hidden state of the backward GRU, represents the bias term, represents the vector concatenation operation, and softmax represents the normalized activation function.

[0014] As a preferred solution of the artificial intelligence-based knowledge graph proposition correction method of the present invention, the global semantic consistency detection includes: Constructing a multidimensional tensor decomposition model to map the multimodal feature data into a unified semantic space to obtain an initial semantic embedding vector; Establish a global semantic constraint rule matrix of the knowledge graph, which includes entity type constraints, relationship transitivity constraints, and mutual exclusion constraints; The global semantic constraint rule matrix for building the knowledge graph includes: Constructing the entity type constraint matrix , defines type compatibility between entities; Constructing the relation transitivity constraint matrix , describing the logical reasoning rules between relations; Generate a comprehensive constraint matrix, the calculation formula is as follows: ; in, represents the final global constraint matrix, represents the type constraint matrix, represents the transitivity constraint matrix, represents the mutual exclusion constraint matrix, Represents the weight coefficient of each constraint; Performing a tensor operation on the initial semantic embedding vector and the global semantic constraint rule matrix to generate a constraint enhancement vector; A scoring model based on an energy function is constructed to calculate the semantic association score between the constraint enhancement vector and the existing triples of the knowledge graph. The calculation formula for constructing the energy function scoring model is as follows: ; in, represents the energy score of the triple (h, r, t), represents the constraint weight, || ||2 represents the L2 norm; A negative sampling technique is used to generate a comparison sample set, and the semantic association score is evaluated with the score of the comparison sample to form a global consistency score; The global consistency scores are grouped based on a hierarchical clustering algorithm, potential semantic conflict points are identified, global semantic verification results are output, and the consistency clustering criteria are defined. The calculation formula is as follows: ; in, Represents a triple Verification results Represents a triple The consistency score, represents a high confidence threshold, Indicates a low confidence threshold.

[0015] As a preferred solution of the artificial intelligence-based knowledge graph proposition correction method of the present invention, wherein: the semantic error correction result of the generated knowledge graph includes: Assigning weight coefficients to the local semantic verification result and the global semantic verification result respectively, and obtaining a fused feature vector by weighted average calculation; A directed graph is used to represent the association relationship between triplets. The fused feature vector is used as the initial state value of the node in the graph. The state transition probability between nodes is calculated using the normalized adjacency matrix to generate an error state probability distribution. Performing triple clustering on the error state probability distribution according to a preset probability threshold and to obtain a high confidence error cluster, a low confidence error cluster and a correct cluster; Retrieving triples having the same head entity or tail entity as the triples in the high-confidence error cluster from the knowledge graph, and constructing a set of correction candidate items; For each candidate item in the modified candidate item set, calculate its conformity score with the type constraints, logical rules and global consistency in the original knowledge graph, and generate a candidate item scoring list; According to the scores of the candidate scoring list, an error correction result report including the error type, the suggested correction triplet and its score is output.

[0016] In a second aspect, the present invention provides a knowledge graph proposition error correction system based on artificial intelligence, which includes: a multimodal feature extraction module, a local semantic verification module, a global semantic verification module and a semantic error correction fusion module; The multimodal feature extraction module is used to extract features from the knowledge graph data, where the knowledge graph data includes text data and structure data, to obtain multimodal feature data; The local semantic verification module is used to build a local semantic verification model based on the graph attention network, perform local semantic consistency detection on the multimodal feature data, and output a local semantic verification result; The global semantic verification module is used to build a global semantic verification model based on the knowledge graph embedding technology, perform global semantic consistency detection on the multimodal feature data, and output a global semantic verification result; The semantic error correction fusion module is used to input the local semantic verification result and the global semantic verification result into the conditional random field model to generate a semantic error correction result of the knowledge graph.

[0017] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the processor executes the computer program, the steps of a knowledge graph proposition correction method based on artificial intelligence are implemented.

[0018] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, the steps of a method for correcting knowledge graph propositions based on artificial intelligence are implemented.

[0019] Compared with the prior art, the beneficial effects of the present invention are that the knowledge graph data is decomposed into text data and structure data through the feature extraction step, and double enhancement is performed by combining the multi-head self-attention mechanism and the graph neural network, thereby achieving the comprehensiveness and accuracy of feature expression; the local semantic verification model constructed by the graph attention network utilizes dynamic attention weight allocation and a bidirectional gated recurrent unit network to achieve accurate judgment of semantic consistency between neighboring nodes; the global semantic verification model constructed by the knowledge graph embedding technology integrates multi-dimensional tensor decomposition and multiple semantic constraint rules to achieve systematic verification of the overall semantic consistency of the knowledge graph; finally, the local and global verification results are integrated through the conditional random field model, combined with differentiated error correction strategies and candidate scoring mechanisms, to form a complete error correction closed loop. This method not only significantly improves the accuracy and availability of the knowledge graph, but also provides reliable technical support for the maintenance and optimization of the knowledge graph, and has important practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 A method flow chart of a knowledge graph proposition error correction method and system based on artificial intelligence provided by one embodiment of the present invention; Figure 2 An internal structural diagram of a computer device of an artificial intelligence-based knowledge graph proposition correction system provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the above-mentioned purposes, features and advantages of the present invention more understandable, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0025] Example 1, reference Figure 1 , which is the first embodiment of the present invention, and provides a knowledge graph proposition error correction method based on artificial intelligence, comprising: Before describing the embodiments of the present application in detail, some related concepts are first explained for the sake of clarity.

[0026] Multi-head self-attention mechanism: a feature processing method in deep learning. By projecting the input features into multiple subspaces in parallel and calculating the attention weights independently, each "head" can focus on different aspects of the features. Finally, the results of multiple heads are combined to enhance the model's ability to understand the input data from multiple angles. It is particularly suitable for processing complex feature data that need to consider multiple correlations.

[0027] DBSCAN clustering algorithm: a density-based clustering method that divides density-connected data points into the same cluster by setting two parameters: neighborhood radius and minimum number of samples. It does not require the number of clusters to be specified in advance. It can discover clusters of any shape and has strong anti-interference ability for noise data. It is particularly suitable for processing non-uniformly distributed data sets.

[0028] Bidirectional Gated Recurrent Unit Network: An improved recurrent neural network structure that includes two processing directions, forward and reverse. It controls the flow of information by updating gates and resetting gates. It can simultaneously consider the historical and future information of sequence data, overcoming the gradient vanishing problem of traditional recurrent neural networks when processing long sequences. It is particularly suitable for sequence data processing tasks that need to consider contextual dependencies.

[0029] Conditional random field model: a probabilistic graphical model that can consider the dependencies and contextual information between elements in the input sequence, achieve global optimal sequence prediction by modeling the transition probability between sequence labels, support joint inference of multiple features, and is widely used in sequence labeling and structured prediction tasks.

[0030] Figure 1The method flow chart of the knowledge graph proposition error correction method and system based on artificial intelligence is shown, including: S1: extracting features from knowledge graph data, where the knowledge graph data includes text data and structure data, to obtain multimodal feature data; Furthermore, feature extraction includes, The text data in the knowledge graph is divided into head entity text, relationship text and tail entity text according to the triple structure, specifically: ; Among them, Triple represents a triple in the knowledge graph, h represents the head entity, r represents the relationship, and t represents the tail entity; The head entity text, relation text and tail entity text are encoded respectively to obtain the text feature vector. The encoding formula is as follows: ; ; ; in, represents the encoding function, represents the pre-trained language model encoder; Constructing a multi-head self-attention matrix based on attention weights, calculating the semantic relevance of the text feature vector, and obtaining a semantically enhanced feature vector; ; ; Among them, Attention represents the attention function, Q represents the query matrix, K represents the key matrix, V represents the value matrix, d represents the feature dimension, MultiHead represents the multi-head attention function, and head i represents the i-th attention head, W O Represents the output projection matrix, Concat represents the concatenation operation; The structural data in the knowledge graph is converted into an adjacency matrix, and a convolution operation is performed on the adjacency matrix based on a graph neural network to obtain a graph structure feature vector; A multi-layer perceptron is used to perform spatial mapping transformation on the graph structure feature vector to obtain a structure enhancement feature vector. The calculation formula is as follows: ; Where MLP( ) represents the multi-layer perceptron function, x represents the input feature, and represents the weight matrix, and represents the bias vector, ReLU() represents the activation function; Constructing a feature fusion layer to splice the semantic enhancement feature vector and the structural enhancement feature vector into a fused feature matrix; The residual connection structure is used to reorganize the fusion feature matrix and output multimodal feature data. The calculation formula is as follows: ; ; in, represents the fusion feature, Semantic features, Represents structural features, || represents feature concatenation operation, LayerNorm() represents layer normalization function, SubLayer() represents sublayer network, and Output represents the final output features.

[0031] S2: construct a local semantic verification model based on the graph attention network, perform local semantic consistency detection on the multimodal feature data, and output a local semantic verification result; Further, local semantic consistency detection includes, Extracting a feature representation of a current triple to be verified from the multimodal feature data, and simultaneously obtaining a set of neighbor triples directly connected to the triple; A computing unit of a graph attention network is constructed, and an initial attention weight coefficient is assigned to each triple in the set of neighbor triples. The calculation formula for assigning the initial attention weight to the set of neighbor triples is as follows: ; in, Represents a slave node To Node The initial attention coefficient, represents the feature transformation weight matrix, and Respectively represent nodes and nodes The characteristic vector of represents the attention vector parameter, || represents the vector concatenation operation, Representation Node The set of neighbor nodes, LeakyReLU represents the activation function; Calculating, for each triple in the neighbor triple set, a semantic similarity matrix with the current triple to be verified, and generating a triple comparison score; Based on the triple comparison scores, the attention weight coefficient in the graph attention network is updated to screen out key neighbor triplets whose semantic relevance is higher than a preset threshold; Performing local semantic consistency calculation on the features of the key neighbor triplet and the features of the current triplet to be verified to obtain a local consistency feature vector; A bidirectional gated recurrent unit network is constructed, and the local consistency feature vector is input into the network for temporal dependency analysis to obtain a local semantic verification result.

[0032] Furthermore, timing dependency analysis includes: Construct a bidirectional GRU network to process local consistency features; The verification result is obtained by fusing the bidirectional features. The calculation formula is as follows: ; in, represents the final local semantic verification result vector, represents the output layer weight matrix, represents the final hidden state of the forward GRU, represents the final hidden state of the backward GRU, represents the bias term, represents the vector concatenation operation, and softmax represents the normalized activation function.

[0033] S3: construct a global semantic verification model based on knowledge graph embedding technology, perform global semantic consistency detection on the multimodal feature data, and output a global semantic verification result; Further, global semantic consistency detection includes, Constructing a multidimensional tensor decomposition model to map the multimodal feature data into a unified semantic space to obtain an initial semantic embedding vector; Establish a global semantic constraint rule matrix of the knowledge graph, which includes entity type constraints, relationship transitivity constraints, and mutual exclusion constraints; The global semantic constraint rule matrix for building the knowledge graph includes: Constructing the entity type constraint matrix , defines type compatibility between entities; Constructing the relation transitivity constraint matrix , describing the logical reasoning rules between relations; Generate a comprehensive constraint matrix, the calculation formula is as follows: ; in, represents the final global constraint matrix, represents the type constraint matrix, represents the transitivity constraint matrix, represents the mutual exclusion constraint matrix, Represents the weight coefficient of each constraint; Performing a tensor operation on the initial semantic embedding vector and the global semantic constraint rule matrix to generate a constraint enhancement vector; A scoring model based on an energy function is constructed to calculate the semantic association score between the constraint enhancement vector and the existing triples of the knowledge graph. The calculation formula for constructing the energy function scoring model is as follows: ; in, represents the energy score of the triple (h, r, t), represents the constraint weight, || ||2 represents the L2 norm; A negative sampling technique is used to generate a comparison sample set, and the semantic association score is evaluated with the score of the comparison sample to form a global consistency score; The global consistency scores are grouped based on a hierarchical clustering algorithm, potential semantic conflict points are identified, global semantic verification results are output, and the consistency clustering criteria are defined. The calculation formula is as follows: ; in, Represents a triple Verification results Represents a triple The consistency score, represents a high confidence threshold, Indicates a low confidence threshold.

[0034] S4: Input the local semantic verification result and the global semantic verification result into the conditional random field model to generate a semantic error correction result of the knowledge graph.

[0035] Furthermore, the semantic error correction results of the generated knowledge graph include: The local semantic verification result and the global semantic verification result are respectively assigned weight coefficients α and β, and a fused feature vector F is obtained by weighted average calculation. The weight coefficients α and β are dynamically adjusted based on the confidence score of the verification result, and α+β=1 is satisfied. It should be noted that the present invention first sets the initial weight coefficients α=0.5 and β=0.5, and normalizes the confidence scores in the local semantic verification result and the global semantic verification result. When the confidence of the local semantic verification result is greater than 0.7, the α value is increased by 0.1, and the β value is correspondingly decreased by 0.1. Conversely, when the confidence of the global semantic verification result is greater than 0.7, the β value is increased by 0.1, and the α value is correspondingly decreased by 0.1. Finally, the fused feature vector is calculated by weighted summation, and the feature values ​​with top 50% confidence in the fused feature vector are retained. A directed graph is used to represent the association relationship between triplets, the fused feature vector is used as the initial state value of the node in the graph, and the state transition probability between nodes is calculated using the normalized adjacency matrix W to generate the error state probability distribution P; specifically, a directed graph G is constructed, the nodes in the graph represent the triplets in the knowledge graph, and the fused feature vector is used as the initial state value of the node; if two triplets share a head entity or a tail entity, a directed edge is established between the corresponding nodes, the weight of the edge is obtained by calculating the number of shared entities between the nodes and dividing it by the total number of entities, and a normalized adjacency matrix is ​​constructed; a random walk algorithm is used to perform 5 rounds of iterative calculations to update the error state probability value of the node until the node state probability distribution tends to be stable; Based on the DBSCAN clustering algorithm, the error state probability distribution P is clustered into triples according to preset probability thresholds θ1 and θ2 to obtain a high confidence error cluster, a low confidence error cluster and a correct cluster; Retrieve triples with the same head entity or tail entity as the triples in the high-confidence error cluster from the knowledge graph, and construct a set of modified candidate items M; specifically, for each triple in the high-confidence error cluster, retrieve all triples containing the same head entity or tail entity in the original knowledge graph; retain candidates with the same relationship type as the original triples to form a set of modified candidate items; annotate each candidate item with its overlapping entity information with the original triple; For each candidate in the modified candidate set M, calculate its conformity score with the type constraints, logical rules and global consistency in the original knowledge graph, and generate a candidate scoring list S; specifically, construct scoring criteria for three types of constraints: entity type constraints check whether the correspondence between entity types and relationships in the candidate is reasonable, and score 1 if satisfied, and score 0 if not satisfied; relationship logic constraints verify whether the candidate violates the transitivity and symmetry of the relationship, and score 1 if no violation, and score 0 if violation exists; global consistency constraints count the number of supporting evidences for the candidate in the knowledge graph, and assign points according to the proportion of the number of evidences; the weighted average of the scores of the three types of constraints is used as the final score of the candidate; According to the scores of the candidate scoring list S, an error correction result report R including the error type, the suggested correction triples and their scores is output.

[0036] Example 2, reference Figure 2 , which is the second embodiment of the present invention, this embodiment also provides a knowledge graph proposition error correction system based on artificial intelligence, including: a multimodal feature extraction module, a local semantic verification module, a global semantic verification module and a semantic error correction fusion module; The multimodal feature extraction module is used to extract features from the knowledge graph data, where the knowledge graph data includes text data and structure data, to obtain multimodal feature data; The local semantic verification module is used to build a local semantic verification model based on the graph attention network, perform local semantic consistency detection on the multimodal feature data, and output a local semantic verification result; The global semantic verification module is used to build a global semantic verification model based on the knowledge graph embedding technology, perform global semantic consistency detection on the multimodal feature data, and output a global semantic verification result; The semantic error correction fusion module is used to input the local semantic verification results and the global semantic verification results into the conditional random field model to generate the semantic error correction results of the knowledge graph. This embodiment also provides a computer device, which may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 2 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a knowledge graph proposition correction method based on artificial intelligence is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0037] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: extracting features from knowledge graph data, where the knowledge graph data includes text data and structure data, to obtain multimodal feature data; Building a local semantic verification model based on the graph attention network, performing local semantic consistency detection on the multimodal feature data, and outputting a local semantic verification result; Building a global semantic verification model based on knowledge graph embedding technology, performing global semantic consistency detection on the multimodal feature data, and outputting a global semantic verification result; The local semantic verification results and the global semantic verification results are input into the conditional random field model to generate semantic error correction results of the knowledge graph.

[0038] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

[0039] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present application may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.

[0040] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0041] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0042] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1A step that specifies a function in one or more boxes.

[0043] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0044] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A knowledge graph proposition error correction method based on artificial intelligence, characterized by: Including, extracting features from knowledge graph data, where the knowledge graph data includes text data and structure data, to obtain multimodal feature data; Building a local semantic verification model based on the graph attention network, performing local semantic consistency detection on the multimodal feature data, and outputting a local semantic verification result; Building a global semantic verification model based on knowledge graph embedding technology, performing global semantic consistency detection on the multimodal feature data, and outputting a global semantic verification result; The local semantic verification results and the global semantic verification results are input into the conditional random field model to generate semantic error correction results of the knowledge graph.

2. The method for correcting questions in a knowledge graph based on artificial intelligence as claimed in claim 1, characterized in that: The feature extraction includes: The text data in the knowledge graph is divided into head entity text, relationship text and tail entity text according to the triple structure. The calculation formula is as follows: ; Among them, Triple represents a triple in the knowledge graph, h represents the head entity, r represents the relationship, and t represents the tail entity; The head entity text, relation text and tail entity text are encoded respectively to obtain the text feature vector. The encoding formula is as follows: ; ; ; in, represents the encoding function, represents the pre-trained language model encoder; Constructing a multi-head self-attention matrix based on attention weights, calculating the semantic relevance of the text feature vector, and obtaining a semantically enhanced feature vector; The structural data in the knowledge graph is converted into an adjacency matrix, and a convolution operation is performed on the adjacency matrix based on a graph neural network to obtain a graph structure feature vector; A multi-layer perceptron is used to perform spatial mapping transformation on the graph structure feature vector to obtain a structure enhancement feature vector; Constructing a feature fusion layer to splice the semantic enhancement feature vector and the structural enhancement feature vector into a fused feature matrix; A residual connection structure is used to reorganize the features of the fused feature matrix and output multimodal feature data.

3. The method for correcting questions in a knowledge graph based on artificial intelligence as claimed in claim 2, characterized in that: The calculation formula for obtaining the semantic enhancement feature vector is as follows: ; ; Among them, Attention represents the attention function, Q represents the query matrix, K represents the key matrix, V represents the value matrix, d represents the feature dimension, MultiHead represents the multi-head attention function, and head i represents the i-th attention head, W O Represents the output projection matrix, Concat represents the concatenation operation; The calculation formula for obtaining the structure enhancement feature vector is as follows: ; Where MLP( ) represents the multi-layer perceptron function, x represents the input feature, and represents the weight matrix, and represents the bias vector, ReLU() represents the activation function; The calculation formula of the feature recombination is as follows: ; ; in, represents the fusion feature, Semantic features, Represents structural features, || represents feature concatenation operation, LayerNorm() represents layer normalization function, SubLayer() represents sublayer network, and Output represents the final output features.

4. The method for correcting questions in a knowledge graph based on artificial intelligence as claimed in claim 3, characterized in that: The local semantic consistency detection includes: Extracting a feature representation of a current triple to be verified from the multimodal feature data, and simultaneously obtaining a set of neighbor triples directly connected to the triple; A computing unit of a graph attention network is constructed, and an initial attention weight coefficient is assigned to each triple in the set of neighbor triples. The calculation formula for assigning the initial attention weight to the set of neighbor triples is as follows: ; in, Represents a slave node To Node The initial attention coefficient, represents the feature transformation weight matrix, and Respectively represent nodes and nodes The characteristic vector of represents the attention vector parameter, || represents the vector concatenation operation, Representation Node The set of neighbor nodes, LeakyReLU represents the activation function; Calculating, for each triple in the neighbor triple set, a semantic similarity matrix with the current triple to be verified, and generating a triple comparison score; Based on the triple comparison scores, the attention weight coefficient in the graph attention network is updated to screen out key neighbor triplets whose semantic relevance is higher than a preset threshold; Performing local semantic consistency calculation on the features of the key neighbor triplet and the features of the current triplet to be verified to obtain a local consistency feature vector; A bidirectional gated recurrent unit network is constructed, and the local consistency feature vector is input into the network for temporal dependency analysis to obtain a local semantic verification result.

5. The method for correcting questions in a knowledge graph based on artificial intelligence as claimed in claim 4, characterized in that: The timing dependency analysis includes: Construct a bidirectional GRU network to process local consistency features; The verification result is obtained by fusing the bidirectional features. The calculation formula is as follows: ; in, represents the final local semantic verification result vector, represents the output layer weight matrix, represents the final hidden state of the forward GRU, represents the final hidden state of the backward GRU, represents the bias term, represents the vector concatenation operation, and softmax represents the normalized activation function.

6. The method for correcting questions in a knowledge graph based on artificial intelligence as claimed in claim 5, characterized in that: The global semantic consistency detection includes: Constructing a multidimensional tensor decomposition model to map the multimodal feature data into a unified semantic space to obtain an initial semantic embedding vector; Establish a global semantic constraint rule matrix of the knowledge graph, which includes entity type constraints, relationship transitivity constraints, and mutual exclusion constraints; The global semantic constraint rule matrix for building the knowledge graph includes: Constructing the entity type constraint matrix , defines type compatibility between entities; Constructing the relation transitivity constraint matrix , describing the logical reasoning rules between relations; Generate a comprehensive constraint matrix, the calculation formula is as follows: ; in, represents the final global constraint matrix, represents the type constraint matrix, represents the transitivity constraint matrix, represents the mutual exclusion constraint matrix, Represents the weight coefficient of each constraint; Performing a tensor operation on the initial semantic embedding vector and the global semantic constraint rule matrix to generate a constraint enhancement vector; A scoring model based on an energy function is constructed to calculate the semantic association score between the constraint enhancement vector and the existing triples of the knowledge graph. The calculation formula for constructing the energy function scoring model is as follows: ; in, represents the energy score of the triple (h, r, t), represents the constraint weight, || ||2 represents the L2 norm; A negative sampling technique is used to generate a comparison sample set, and the semantic association score is evaluated with the score of the comparison sample to form a global consistency score; The global consistency scores are grouped based on a hierarchical clustering algorithm, potential semantic conflict points are identified, global semantic verification results are output, and the consistency clustering criteria are defined. The calculation formula is as follows: ; in, Represents a triple Verification results Represents a triple The consistency score, represents a high confidence threshold, Indicates a low confidence threshold.

7. The method for correcting questions in a knowledge graph based on artificial intelligence as claimed in claim 6, characterized in that: The semantic error correction results of the generated knowledge graph include: Assigning weight coefficients to the local semantic verification result and the global semantic verification result respectively, and obtaining a fused feature vector by weighted average calculation; A directed graph is used to represent the association relationship between triplets. The fused feature vector is used as the initial state value of the node in the graph. The state transition probability between nodes is calculated using the normalized adjacency matrix to generate an error state probability distribution. Performing triple clustering on the error state probability distribution according to a preset probability threshold and to obtain a high confidence error cluster, a low confidence error cluster and a correct cluster; Retrieving triples having the same head entity or tail entity as the triples in the high-confidence error cluster from the knowledge graph, and constructing a set of correction candidate items; For each candidate item in the modified candidate item set, calculate its conformity score with the type constraints, logical rules and global consistency in the original knowledge graph, and generate a candidate item scoring list; According to the scores of the candidate scoring list, an error correction result report including the error type, the suggested correction triplet and its score is output.

8. A knowledge graph proposition correction system based on artificial intelligence, based on the knowledge graph proposition correction method based on artificial intelligence according to any one of claims 1 to 7, characterized in that: It includes a multimodal feature extraction module, a local semantic verification module, a global semantic verification module, and a semantic error correction fusion module; The multimodal feature extraction module is used to extract features from the knowledge graph data, where the knowledge graph data includes text data and structure data, to obtain multimodal feature data; The local semantic verification module is used to build a local semantic verification model based on the graph attention network, perform local semantic consistency detection on the multimodal feature data, and output a local semantic verification result; The global semantic verification module is used to build a global semantic verification model based on the knowledge graph embedding technology, perform global semantic consistency detection on the multimodal feature data, and output a global semantic verification result; The semantic error correction fusion module is used to input the local semantic verification result and the global semantic verification result into the conditional random field model to generate a semantic error correction result of the knowledge graph.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the knowledge graph proposition correction method based on artificial intelligence according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the knowledge graph proposition correction method based on artificial intelligence described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Knowledge graph error detection method, system and device

    CN118277584A

  • Characteristic representation-enhanced proper noun named entity recognition method

    CN118313381A

  • Domain knowledge graph construction method based on GlobalPointer joint extraction

    CN118396102A

  • Power equipment knowledge graph error detection method and system based on knowledge reconstruction

    CN119046474A

  • Generating answers to multi-hop constraint-based questions from knowledge graphs

    US20230169361A1

Cited By

  • Disease science popularization error correction method and system based on artificial intelligence

    CN120108774A

  • Bid invitation file error content optimization method and system based on artificial intelligence

    CN120430297A

  • Enterprise application complex service node optimization method and system based on graph neural network

    CN120892607A

  • Environmental anomaly broadcasting method and system based on semantic analysis and knowledge graph

    CN121072551A

  • Online automatic thematic map making method and system based on artificial intelligence

    CN121616699A