Training method and quality inspection method for a strongly robust knowledge graph triple quality inspection network model based on a noisy dataset

By constructing a noise-free dataset and using a network model to extract triplet features, the problem of complex knowledge graph quality inspection algorithms in the existing technology ignore implicit triplets, and more accurate knowledge graph quality inspection and noise processing are achieved.

CN116150401BActive Publication Date: 2025-06-27DALIAN OCEAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310121294.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-16
Publication Date
2025-06-27
Estimated Expiration
2043-02-16

AI Technical Summary

Technical Problem

When using complex knowledge graph quality inspection algorithms, only using isolated triplets as positive samples, will greatly weaken the knowledge contained in the knowledge graph and lack effective methods for building noise data sets.

Method used

By constructing a noise-bound dataset, including source triplets, implicit triplets and noise triplets, the initial features, static features and internal correlation features of triplets are extracted using the network model, the fusion features are aggregated to obtain the fusion features, and the correlation relationship between entities is trained through a multi-label classification algorithm.

Benefits of technology

A more accurate confidence calculation of implicit triplets is achieved, the knowledge mining ability in the knowledge graph is improved, the problem of no transfer relationship between entities in the noise triplets is avoided, and the robustness and accuracy of quality inspection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116150401B_ABST
    Figure CN116150401B_ABST
Patent Text Reader

Abstract

A strong and robust knowledge graph triple quality inspection network model training method and quality inspection method based on a noisy data set belong to the field of knowledge graph triple quality inspection. In order to solve the problem that only isolated triples are used as positive samples, which will greatly weaken the knowledge contained in the knowledge graph, a data set is constructed, which includes source triples; implicit triples consisting of a transitive relationship between head and tail entities are constructed; noise triples are constructed; the confidence of the triples is obtained; triple fusion features are obtained through network model aggregation; the network model uses a multi-label classification algorithm to train the association relationship between entities to distinguish triplets with no association relationship from triplets with an entity association relationship; the model parameters are optimized through the entity association relationship loss and the binary cross entropy loss in the feature modeling process, the effect is to improve the knowledge contained in the knowledge graph and more accurately mine the implicit semantic relationship between knowledge graph nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of quality inspection of knowledge graph triples, and specifically relates to a training method and a quality inspection method for a strongly robust knowledge graph triple quality inspection network model based on a noisy data set. Background Art

[0002] The basic storage unit of a knowledge graph is a triple, which consists of a head entity, a relation, and a tail entity. Triples are linked together by relations to form a huge directed graph. Large knowledge bases such as DBpedia and NELL are all crawled from multiple websites, cleaned and made. Their complex knowledge structures are often difficult to effectively quality inspect and analyze. In the process of making a knowledge graph, some noisy data are often introduced, such as false relations, wrong entities, and even non-existent triples out of thin air. Due to the inevitable introduction of noisy triples in the process of making a knowledge graph, these triples damage the network structure of the knowledge graph, making it difficult for knowledge to be effectively displayed, and fatal errors will occur in knowledge recommendation and search based on the knowledge graph.

[0003] To effectively quality inspect a knowledge graph, Ruobing Xie et al. proposed a triple confidence algorithm. The confidence of a triple can be calculated before and after graph construction. The result of confidence calculation implies the internal features of the knowledge graph and the implicit information between triples. Shengbin Jia et al. integrated the internal semantic features of triples, the global semantic dependence information of nodes, and the credibility between the components of triples based on a deep learning model to construct a strongly robust noisy triple quality inspection algorithm, whose performance far exceeds traditional algorithms such as TransE and TransR. Yu Zhao et al. expanded the vector representation of head and tail node entities to some extent, mainly considering that entities already contain rich semantic information. Shengbin Jia and Yu Zhao et al. both used the Trans series algorithms as the basic algorithms and integrated entity and relation vectors at multiple levels to achieve good results. However, there are the following problems in the current quality inspection of knowledge graphs: 1) Most scholars design quality inspection algorithms based on common open-source knowledge graphs, artificially construct noisy data sets, and convert the quality inspection of knowledge graphs into common classification tasks. There is a lack of effective methods for constructing noisy data sets; 2) Complex knowledge graphs such as the FB15K-237 knowledge graph contain 237 kinds of relations, and there are complex relation transmissions between triples. Using only isolated triples as positive samples will greatly weaken the knowledge contained in the knowledge graph. Summary of the Invention

[0004] To solve the problem that using only isolated triples as positive samples will greatly weaken the knowledge contained in the knowledge graph, in a first aspect, a training method for a strongly robust knowledge graph triple quality inspection network model based on a noisy data set according to some embodiments of the present application includes:

[0005] Construct a data set, where the data set includes source triples formed by a direct relationship between a head entity and a tail entity;

[0006] Construct implicit triples formed by a transitive relationship between the head entity and the tail entity according to the data set;

[0007] Construct noise triples according to the source triples in the data set;

[0008] Obtain the confidence levels of the source triples, implicit triples, and noise triples;

[0009] Extract the initial features, static features, and internal association features of the source triples, implicit triples, and noise triples through a network model, and aggregate them to obtain the fusion features of the source triples, implicit triples, and noise triples;

[0010] According to the fusion features of the source triples, implicit triples, and noise triples, the network model uses a multi-label classification algorithm to train the association relationship between entities to distinguish triples with no association relationship between entities from triples with an association relationship between entities;

[0011] Optimize the model parameters through the entity association relationship loss and the total loss in the feature modeling process.

[0012] According to the training method for a strongly robust knowledge graph triple quality inspection network model based on a noisy data set according to some embodiments of the present application, a method for constructing implicit triples with a transitive relationship between the head entity and the tail entity according to the data set includes:

[0013] Use the entities in the data set as search starting points, search for the longest directed paths starting from these entities, traverse all entities in the data set, and obtain the longest directed paths of each entity and the search paths of each entity;

[0014] Delete the included sub-paths from the search paths to obtain all non-inclusive search paths;

[0015] Construct an entity-relationship matrix E through all non-inclusive search paths, and use the entity-relationship matrix E to construct implicit triples using the relationship transfer direction, where the entity-relationship matrix E is represented by the following formula:

[0016]

[0017] where, sig i,j={0,1}, where D is the number of distinct entities in the dataset, and sig i,j is the relationship between entity En i and En j . sig i,j =0 indicates that there is no association between these two entities, and sig i,j =1 indicates that there is an association between these two entities. For the triple <En i ,sig i,j ,En j >, in the entity-relationship matrix E, the relationship sig i,j =1, and the unit group composed of the entity En i,j corresponding to the relationship sig i,j and the entity En i and En j is an implicit triple.

[0018] According to the training method of the strongly robust knowledge graph triple quality inspection network model based on a noisy dataset in some embodiments of the present application, the method for obtaining the confidence of implicit triples in the obtaining of the confidence of triples includes

[0019] Traverse all search paths, restore any search paths through the entity-relationship matrix E, and obtain the longest search path;

[0020] Based on the identification of the longest search path, calculate the confidence matrix of each entity triple on the longest search path;

[0021] Calculate the confidence of the constructed implicit triples through the confidence matrix of each entity triple on the longest search path, and each longest search path is independent of each other:

[0022] The confidence is represented by formula (3):

[0023]

[0024] where r represents the confidence, ← represents the pointing direction, F refers to the number of the longest search paths containing the triple <En i ,sig i,j ,En j >, d k refers to the search depth of the current triple in the current triple it belongs to, p k is the total length of the current search path, that is, the number of triples included, L is the maximum length of all the longest search paths, and all the confidences are normalized by the parameter L. D is the number of distinct entities in the dataset.

[0025] The method for training a strong and robust knowledge graph triple quality inspection network model based on a noisy data set according to some embodiments of the present application, and the method for constructing noisy triples based on the source triples and implicit triples of the data set, including that any one of replacing the head entity <?, r, t>, replacing the relationship <h,?, t>, and replacing the tail entity <h, r,?> in the triples results in a noisy triple, and the source triples, implicit triples, and noisy triples are retained in the data set.

[0026] The method for training a strong and robust knowledge graph triple quality inspection network model based on a noisy data set according to some embodiments of the present application, the network model includes a TransR network, a residual network, and a BiLSTM network, and the method for extracting the initial features, static features, and internal association features of the triples through the network model includes

[0027] Obtaining the initial features of the source triples, implicit triples, and noisy triples through the TransR network;

[0028] Extracting the static features of the source triples, implicit triples, and noisy triples through the residual network;

[0029] Extracting the internal association features of the source triples, implicit triples, and noisy triples through a multi-layer BiLSTM network.

[0030] The method for pre-training the source triples, implicit triples, and noisy triples by the TransR model according to some embodiments of the present application, including taking the inner product of the embeddings of the source triples, implicit triples, and noisy triples with the confidence of the triples to obtain a weighted feature vector, and the weighted feature vector is the initial feature of the triples.

[0031] According to some embodiments of the present application, in the feature modeling process, the entity association relationship loss and the total loss are respectively represented by formulas (7) and (8):

[0032]

[0033]

[0034] where, L EP represents the entity association relationship loss, B represents the input batch size of the current training, a is the association depth of all batch samples, y i represents the entity association relationship label, p i represents the entity association relationship prediction probability. L represents the total loss, y - represents the triple quality inspection label p-j Represents the quality inspection classification probability of each triple by the neural network; y j Represents the entity association relationship label in the feature modeling process, p j Represents the prediction probability of the neural network for each entity association relationship.

[0035] In a second aspect, a strong robustness knowledge graph triple quality inspection method based on a noisy dataset according to some embodiments of the present application includes

[0036] Inputting the dataset to be quality inspected into the network model with the optimized model parameters obtained by the training method:

[0037] Extracting the initial features, static features, and internal association features of the triples in the dataset to be quality inspected through the network model, and aggregating to obtain the fusion features of the triples;

[0038] According to the fusion features of the triples, the network model predicts the association relationship between entities through a multi-label classification algorithm, and distinguishes the triples with no association relationship between entities from the triples with an association relationship between entities.

[0039] According to some embodiments of the present application, a method for a strong robustness knowledge graph triple quality inspection method based on a noisy dataset, the method for extracting the initial features, static features, and internal association features of the triples in the dataset to be quality inspected through the network model, includes

[0040] Obtaining the initial features of the triples through the TransR network;

[0041] Extracting the static features of the triples through the residual network;

[0042] Extracting the internal association features of the triples through a multi-layer BiLSTM network.

[0043] According to some embodiments of the present application, a method for pre-training the TransR model on source triples, implicit triples, and noisy triples in a strong robustness knowledge graph triple quality inspection method based on a noisy dataset includes taking the inner product of the embeddings of the source triples, implicit triples, and noisy triples with the confidence of the triples to obtain a weighted feature vector, and the weighted feature vector is the initial feature of the triples.

[0044] Advantages of the present invention:

[0045] 1) The present invention assigns a preset weight to each triple, represents the confidence that the triple is true, and proposes a more accurate method for calculating the confidence of implicit triples.

[0046] 2) Construct implicit triples for complex knowledge graphs and use them in the training of quality inspection models. The triples distinguished by the quality inspection models used will no longer ignore the implicit triples with indirect relationships in the test dataset, improving the knowledge contained in the knowledge graph. To more accurately mine the implicit semantic relationships between nodes in the knowledge graph, a method for representing the strength of relationships based on search depth is also proposed. Nodes in complex knowledge graphs are linked through relationships, based on the link depth. The present invention uses a depth-first search algorithm based on a directed graph to search for all possible paths and constructs new implicit triples based on the search paths to expand the scale of the source triples;

[0047] 3) Construct noise triples based on the expanded triples. There are three types of noise triples constructed in the present invention, namely replacing the head entity <?, r, t>, replacing the relationship <h,?, t>, and replacing the tail entity <h, r,?>. Since the present invention has greatly expanded the source triples, it can greatly avoid the situation where there is no implicit transitive relationship between any pair of entities in the constructed noise triples;

[0048] 4) The present invention uses TransR to pre-train the expanded real triples to obtain the initial expressions of entities and relationships, and then uses a variety of deep learning algorithms to model the triples, and finally completes quality inspection through feature fusion. Description of the Drawings

[0049] Figure 1 Basic framework diagram.

[0050] Figure 2 Graphs of experimental results for Accuracy, F-Score, Precision, and Recall, Figure 2 A of which is the graph of Recall experimental results, Figure 2 B of which is the graph of Accuracy experimental results, Figure 2 C of which is the graph of F1 experimental results, Figure 2 D of which is the graph of Precision experimental results.

[0051] Figure 3 Graphs of comparative experimental results for 5% noise samples, Figure 3 A of which is the graph of Recall experimental results, Figure 3 B of which is the graph of Precision experimental results.

[0052] Figure 4 Graphs of comparative experimental results for 3% noise samples, Figure 4 A of which is the graph of Recall experimental results, Figure 4 B of which is the graph of Precision experimental results. Detailed Implementation Modes

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only possible technical implementations of the present invention, not all possible implementations. Those skilled in the art can completely combine the embodiments of the present invention to obtain other embodiments without creative labor, and these embodiments are also within the protection scope of the present invention.

[0054] Definition 1: An implicit triple refers to a situation in a complex knowledge graph where the relationship transmission between entities results in an indirect relationship between entities. The newly constructed triple based on the relationship transmission is called an implicit triple.

[0055] Definition 2: A source triple refers to a triple in a knowledge graph that is formed by a direct relationship between the head entity and the tail entity.

[0056] Definition 3: A positive triple refers to a triple in a knowledge graph that is formed by a direct or transitive relationship between the head entity and the tail entity. There are two sources of positive triples: 1) the source triples provided by the training set, 2) the implicit triples described in Definition 1).

[0057] The inventor found that the common open-source knowledge graph design quality inspection model training usually uses the original triples provided by the training set, and there is a direct relationship between the head entity and the tail entity of the original triples. However, for a complex knowledge graph, there is often a transitive relationship between the head entity and the tail entity. The triples formed by this transitive relationship reflect the indirect relationship between the head entity and the tail entity. The prior art does not use implicit triples to participate in the training in the quality inspection model training, and the triples distinguished by the used quality inspection model ignore the implicit triples with indirect relationships in the test dataset, which will greatly weaken the knowledge contained in the knowledge graph. To more accurately mine the implicit semantic relationships between the nodes of the knowledge graph, the present invention first preprocesses the dataset to obtain an implicit triple dataset, expands the source triples in the source graph, then constructs noise triples, and uses the source triples, implicit triples, and noise triples to train the network model.

[0058] Specifically, the present invention is a method for training a strongly robust knowledge graph triple quality inspection network model based on a noisy dataset, including

[0059] The method for constructing implicit triples in the present invention includes the following steps:

[0060] S101. Construct a dataset, where the dataset includes source triples formed by a direct relationship between the head and tail entities. Among them, in step S101, Neo4J datasets are respectively constructed based on the FB15K-237 and WN18RR datasets.

[0061] S102. Construct implicit triples in which there is a transitive relationship between the head and tail entities based on the said data set. Moreover, in this step, a method for calculating the confidence of implicit triples is specifically recorded, while for source triples and noisy triples, existing confidence calculation methods can be used. Among them, in step S102, taking the entities in the data set as the search starting point, search for the longest directed path starting from this entity. Traverse all entities in the data set to obtain all search paths, then delete the included sub-paths, and finally obtain all non-inclusive paths and construct an entity-relationship matrix E. Based on the entity-relationship matrix, use the relationship transmission direction to construct implicit triples. The entity-relationship matrix E is shown in formula 1.

[0062]

[0063] Among them, sig i,j = {0, 1}, D is the number of non-repeated entities in the data set, sig i,j is the relationship between entity En i and En j . sig i,j = 0 indicates that there is no association between these two entities, and sig i,j = 1 indicates that there is an association between these two entities. For the triple <En i , sig i,j , En j >, in the entity-relationship matrix E, the relationship sig i,j = 1, and the unit group composed of the relationship sig i,j and the relationship sig i,j and the corresponding entities En i and En j is an implicit triple.

[0064] Among them, Entity→R D means, and Entity↓ means both are entities.

[0065] Since the entity-relationship matrix E is obtained based on a directed graph search, the triple <En i , sig i,j , En j > and the triple <En j , sig i,j , En i > are considered different triples. Any search path can be restored based on matrix E, and the search path is represented as shown in formula 2.

[0066] DPath←<En i , 1, En j >∪<En j , 1, En k>∪…<En m ,1,En n > Formula 2

[0067] Since each search path requires a directed edge between adjacent nodes and points from the head entity to the tail entity, the present invention constructs a triple confidence matrix based on the search depth based on the directed search path. This confidence matrix is used to identify the strength of the association between the head and tail entities in each triple. Considering that some entities may be included in multiple search paths and the confidence calculation is chaotic due to different depths, to solve this problem, the present invention calculates the confidence of the constructed implicit triples only based on the longest search path identified by matrix E, and each longest search path is independent of each other. The confidence calculation method is shown in Formula 3.

[0068]

[0069] Among them, r represents the confidence, ← represents the pointing direction, F refers to the triple <En i ,sig i,j ,En j > the number of the longest search paths, d k refers to the search depth of the current triple in the currently affiliated triple, p k is the total length of the current search path, that is, the number of triples included, L is the maximum length of all the longest search paths, and all confidences are normalized through the parameter L, and D is the number of non-repeating entities in the dataset.

[0070] S103. Construct noise triples according to the source triples of the dataset. Among them,

[0071] Noise triples refer to false triples that have no intersection with positive triples and are not included in the extended knowledge graph. To fully test the quality inspection effect of the algorithm of the present invention on the knowledge graph, the present invention constructs 3 sets of noise datasets for each original dataset, namely HR_FAKE_T, H_FAKER_T, and FAKEH_R_T. HR_FAKE_T randomly replaces the tail entity based on the positive triple, H_FAKER_T randomly replaces the relationship based on the positive triple, and FAKEH_R_T randomly replaces the head entity based on the positive triple. The construction process of the 3 sets of noise datasets is shown in Algorithm 1.

[0072] Algorithm 1 Noise Dataset Construction

[0073]

[0074] In Algorithm 1, the Check function respectively realizes the selection of 3 types of noise triples, and the pseudocode is shown in Algorithm 2.

[0075] Algorithm 2 Check (Selecting Noisy Triples)

[0076]

[0077] Algorithm 1 and Algorithm 2 implement the selection and filtering of three types of noisy datasets. The filtering conditions include two: 1) The newly generated noisy triples should not appear in the set of extended positive triples; 2) The newly generated noisy triples should not appear in the entity-relationship association matrix E. Through the above two filtering methods, it is possible to greatly avoid the absence of transitive relationships between the head and tail entities of noisy triples. The positive triples and noisy triples are merged to obtain a new dataset.

[0078] S104. Obtain the confidence levels of the source triples, implicit triples, and noisy triples. As described above, the method for calculating the confidence level of implicit triples is recorded in step S102, and existing confidence level calculation methods can be used for source triples and noisy triples.

[0079] S105. Extract the initial features, static features, and internal association features of the source triples, implicit triples, and noisy triples through a network model, and aggregate them to obtain the fusion features of the source triples, implicit triples, and noisy triples. Among them, since there are a large number of 1:N and N:N relationships in the FB15K-237 and WN18RR datasets, the present invention trains positive triples based on the TransR algorithm to obtain vector representations of entities and relationships, and then traverses the noisy triples in the three datasets, and initializes all the noisy triples with the model parameters trained by TransR. The embedding of all positive triples is inner-producted with their confidence levels to obtain a weighted feature vector, that is, the initial features are obtained.

[0080] According to Figure 1 shown, Po-TransR represents positive triples initialized based on the TransR algorithm, and N-Random represents noisy triples. Both noisy triples and positive triples are initialized with vectors of the same dimension. DeepPath is a search path constructed based on the entity-relationship matrix.

[0081]

[0082] The present invention extracts the static features of triples through a residual network.

[0083] Considering that a certain scale of directed search paths has been obtained during the in-depth preprocessing of the knowledge graph in the present invention, the spatio-temporal semantic associations between entities are of certain significance for the deep representation of entity vectors. The prior art uses TransE to train triples to obtain vector representations of triples, and directly solves the local features, global features, and path features containing semantics of triples based on the vector distribution of triples and directed subgraphs. In the present invention, a multi-layer BiLSTM is used to model the spatial semantic relationships of the original input and learn the local association relationships between entities; then BiLSTM is used to extract the internal association features of triples. The initial features, static features, and internal association features of triples are aggregated to obtain the fusion features of triples.

[0084] S106. According to the fusion features of the source triples, implied triples, and noise triples, the network model distinguishes the triples with no association between entities from the triples with association between entities through training the association relationships between entities by a multi-label classification algorithm, and optimizes the model parameters through the entity association relationship loss and binary cross-entropy loss during the feature modeling process.

[0085] The initial features, static features, and internal association features of triples are aggregated to obtain the fusion features of triples, and the input of feature modeling is shown in Formula 4.

[0086]

[0087] Where B refers to BatchSize, that is, the size of the input batch for the current training, a is the association depth of all batch samples, and a ≤ B. The target output label of feature modeling is shown in Formula 5, and the label meaning is shown in Formula 6.

[0088]

[0089] Symbol represents that there is no association between entity En i and entity En j , and their association label is 0. The symbol → represents that there is an association between entity En i and entity En j , and their association label is 1.

[0090] In the algorithm of the present invention, during the feature modeling process, through the training and prediction of the association relationships between entities by a multi-label classification algorithm, entities with no association are distinguished, and the quality inspection of true and false triples is realized in a binary classification manner. The losses of the two are aggregated to jointly optimize the network parameters. The entity association relationship loss during the feature modeling process is shown in Formula 7.

[0091]

[0092] The triple quality inspection uses the common binary cross - entropy loss. After merging with Equation 7, the total loss is obtained, as shown in Equation 8.

[0093]

[0094] Among them, L EP represents the entity association relationship loss, B represents the input batch size of the current training, a is the association depth of all batch samples, y i represents the entity association relationship label, p i represents the entity association relationship prediction probability. L represents the total loss, y - represents the triple quality inspection label p - j represents the neural network's classification probability for each triple quality inspection; y j represents the entity association relationship label in the feature modeling process, p j represents the neural network's prediction probability for each entity association relationship.

[0095] The difficulty of knowledge graph triple quality inspection is to distinguish real triples and noise triples. Common open - source knowledge graphs do not contain noise triples. Currently, few existing triple quality inspection algorithms consider the influence of a large number of implicit triples in the knowledge graph due to relationship transitivity on the quality inspection effect, and do not effectively utilize the spatial semantic association between entities, resulting in insufficient extraction of entity features. To address the above problems, a strongly robust implicit triple quality inspection algorithm with a noisy dataset (Implied triplet quality inspection, ITQI) is proposed. First, a Neo4J knowledge graph is made based on an open - source dataset; then, all possible search paths are searched based on the longest path search algorithm of a directed graph, and triples with implicit relationships are constructed according to the relationship transitivity of the knowledge graph. Expanding the source triples can greatly increase the number of valid triples; finally, three types of noise triples are constructed, namely <h,r,?>, <h,?,t>, <?,r,t>, where? represents a missing value and is obtained by random sampling. The scale of these three types of noise triples is the same as that of the expanded real triples. The initial features of the expanded real triples are obtained through TransR pre - training, and then the residual network is used to extract the static features of the triples, and the multi - layer BiLSTM is used to extract the internal association features of the triples. The above three types of features are aggregated to obtain the fusion features of the triples for binary classification of the triples to achieve the purpose of triple quality inspection. The algorithm of the present invention is experimented on two datasets, FB15K and WN18RR. The experimental results show that the quality inspection effect of the algorithm of the present invention on three types of noise data reaches the best, and the robustness is the strongest.

[0096] Experimental example

[0097] ITQI algorithm comparison experiment

[0098] Experimental environment

[0099] The datasets used in this invention are FB15K-237 and WN18RR. These two datasets will be introduced later. The ITQI algorithm proposed in this invention can be quickly deployed and run on the GPU, and comparative experiments are conducted with other algorithms on the CPU. The configurations of the comparative experiments are shown in Table 1. The basic settings of the experiments are shown in Table 2.

[0100] Table 1 Experimental hardware conditions

[0101]

[0102] Table 2 Experimental condition settings

[0103]

[0104] Dataset

[0105] The ITQI algorithm and the comparative algorithms are compared on multiple datasets. The basic information of the datasets used in this invention is shown in Table 3.

[0106] Table 3 Basic information of the experimental datasets

[0107]

[0108] In Section 2.2 of this invention, the directed longest path search algorithm is used to map the presence or absence of the relationship between all entities to the entity relationship association matrix E. Entities with direct or indirect relationships are considered to be able to construct positive triples. Based on matrix E, the original positive triples are greatly expanded. The data scale of the expanded training set is shown in Table 4.

[0109] Table 4 Basic information of the positive triples in the training set

[0110]

[0111] The noisy triples are constructed according to Algorithm 1 and Algorithm 2, and their triple scales are basically the same as the scales of the training set, test set, and validation set of each dataset.

[0112] The comparative algorithms used in the experiments of the present invention are shown in Table 5. The evaluation metrics are: Accuracy, Precision, Recall-Score, F1-Score, and Quality. The calculation formulas for these 4 evaluation metrics directly call the calculation formulas encapsulated in Sklearn.metrics to calculate the values of these 4 metrics. The Quality metric is an evaluation metric for measuring the quality of triple quality inspection. The present invention draws on the formula for calculating the Quality metric proposed by Shengbin Jia et al., and takes 0.5 as the dividing line for triple quality inspection. That is, if the probability of a triple predicted as positive is less than 0.5, it is considered a prediction error, and if the probability of a triple predicted as positive is greater than 0.5, it is also considered a prediction error.

[0113] Table 5 Comparative Algorithms and Evaluation Metrics

[0114]

[0115] FB15K-237 Dataset Comparative Experiment

[0116] The algorithm of the present invention first conducts quality inspection experiments on the FB15K-237 dataset, and the experimental objects are as follows:

[0117] 1) Positive triples + HR_FAKE_T;

[0118] 2) Positive triples + H_FAKER_T;

[0119] 3) Positive triples + FAKEH_R_T.

[0120] Among them, the creation of the 3 noise datasets of HR_FAKE_T, H_FAKER_T, and FAKEH_R_T has been introduced in detail above. The evaluation metrics for the 3 groups of experiments are Accuracy, F-Score, Precision, and Recall. The experimental results are as Figure 2 shown, and the experimental results are summarized in Table 6.

[0121] Table 6 Experimental Results on Three Datasets

[0122]

[0123] It can be seen from the experimental results of the algorithm of the present invention on the 3 datasets that the experiments of the present invention have good robustness, and the experimental results under the 4 evaluation metrics are all relatively high. The present invention uses the two evaluation metrics of Recall and Quality to conduct comparative experiments with the comparative algorithms respectively, and the experimental results are shown in Table 7.

[0124] Table 7 Comparative Experimental Results

[0125]

[0126] As can be seen from Table 7, the experimental results of the proposed algorithm ITQI of the present invention on the three extended sets of the FB15K-237 dataset are better than the experimental results of other algorithms on the original dataset. The summary of the improvement of the evaluation metrics on the three datasets is shown in Table 8. Compared with the average recall rate and the average quality inspection quality of other comparison algorithms, the maximum improvement in the recall rate of the algorithm of the present invention on the three extended sets is 6.09% and the minimum improvement is 2.92%; the maximum improvement under the Quality index is 15.09% and the minimum improvement is 12.09%; and KGTtm - 、PTransE - and TransR - In comparison, the maximum improvement in the recall rate of the algorithm of the present invention is 7.275% and the minimum improvement is 0.201%; the maximum improvement in the Quality index is 14.98% and the minimum improvement is 1.251%. Based on the above comparison results, the average improvement rate of the algorithm of the present invention on the two comparison metrics and the improvement rate on a single comparison algorithm are both positive. The experiment shows that the algorithm of the present invention has certain advantages.

[0127] Table 8 Comparison and improvement results with other algorithms

[0128]

[0129] WN18RR dataset comparative experiment

[0130] The comparative experiment of the present invention verifies the effect of each algorithm on the quality inspection of triples under different proportions of noise and conflict samples. The experimental dataset is WN18RR, and the experimental objects are the same as those in Section 3.3.1. The experimental evaluation metrics of the algorithm of the present invention on these three groups of datasets are Precision and Recall. The calculation formulas of Precision and Recall are the same as those proposed by Qinggang Zhang et al. The experimental results under the injection condition of 5% noise samples are as Figure 3 shown, and the experimental results under the injection condition of 3% noise samples are as Figure 4 shown. The summary of the experimental results is shown in Table 9.

[0131] Table 9 Summary of experimental results

[0132]

[0133] The present invention uses two evaluation metrics, Recall and Precision, to conduct comparative experiments with the comparison algorithms respectively. The experimental results are shown in Table 10.

[0134] Table 10 Comparative experiment results

[0135]

[0136] As can be seen from Table 10, the average experimental results of the ITQI algorithm proposed in the present invention on the three extended sets of the WN18RR dataset are better than the experimental results of other algorithms on the original dataset. The summary of the improvement of the evaluation metrics on the three datasets is shown in Table 11. Compared with the average recall rate and the average quality inspection quality of other comparison algorithms, the maximum improvement in the average recall rate of the algorithm of the present invention on the three extended sets is 58.92%, and the minimum improvement is 20.55%; under the Precision index, the maximum improvement is 58.68%, and the minimum improvement is 24.14%; and KGTtm - 、KGIst - and CAGED - Compared, the maximum improvement is 73.88%, and the minimum improvement is 3.17%; under the Precision index, the maximum improvement is 73.61%, and the minimum improvement is 6.33%. Based on the above comparison results, the average improvement rate of the algorithm of the present invention on the two comparison metrics and the improvement rate on a single comparison algorithm are both positive. The experiment shows that the algorithm of the present invention has certain advantages.

[0137] Table 11 Comparison and improvement results with other algorithms

[0138]

[0139] Ablation experiment

[0140] To verify the influence of each module of the algorithm of the present invention on the algorithm effect, the ablation experiment algorithm set in Table 5 is set in the present invention. All comparison algorithms do not include the DeepPath part in the algorithm framework Figure 1 For the convenience of analysis, the present invention only uses Recall as the evaluation index. The ablation experiment results are shown in Table 12.

[0141] Table 12 Ablation experiment results

[0142]

[0143] From the ablation experiment results, compared with the comparison algorithm, the maximum average improvement rate of the algorithm of the present invention is 2.84%, and the minimum is 1.40%. Moreover, the more feature extraction modules are added, the more significant the effect is. When the DeepPath module is not added to the comparison algorithms, the recall rates are all lower than those of the algorithm of the present invention. The experimental results show that the DeepPath structure has a certain improvement effect on triple quality inspection.

[0144] Aiming at the fact that existing triple quality inspection algorithms rarely consider the impact of a large number of implicit triple pairs in the knowledge graph due to relationship transmission on the quality inspection effect, the present invention proposes a strong robustness implicit triple quality inspection algorithm ITQI based on a noisy data set. First, the FB15K-237 and WN18RR data sets are respectively expanded to obtain a larger-scale triple, and three groups of noisy data sets are generated using Algorithm 1 and Algorithm 2 respectively. Experiments on the data set by the algorithm of the present invention and the comparative algorithm show that the algorithm of the present invention has a higher accuracy rate and is superior to other algorithms. From the comparison results of the evaluation indicators, the algorithm of the present invention has a higher recall rate on data sets such as positive triples + FAKEH_R_T, and the quality inspection quality of the triples is higher. From the results of the ablation experiment, the relationship dependence features between entities can contribute to the modeling of noisy triples and can help distinguish noisy samples.

Claims

1. A training method for a strongly robust knowledge graph triple quality inspection network model based on a noisy dataset, characterized in that For knowledge recommendation and search, including: Construct a dataset, which includes source triples formed by direct relationships between head and tail entities, and the dataset is WN18RR; Construct implicit triples with transitive relationships between head and tail entities according to the dataset; Construct noise triples according to the source triples in the dataset; Obtain the confidence levels of the source triples, implicit triples, and noise triples; Extract the initial features, static features, and internal association features of the source triples, implicit triples, and noise triples through a network model, and aggregate them to obtain the fusion features of the source triples, implicit triples, and noise triples; According to the fusion features of the source triples, implicit triples, and noise triples, the network model uses a multi-label classification algorithm to train the association relationships between entities to distinguish triples with no association relationships between entities from triples with association relationships between entities; Optimize the model parameters through the entity association relationship loss and total loss during feature modeling.

2. The training method of the strong robust knowledge graph triple quality inspection network model based on the noisy data set according to claim 1, characterized in that, The method for constructing triples with transitive relationships between head and tail entities according to the dataset includes: Use the entities in the dataset as the search starting point, search for the longest directed path starting from the entity, traverse all entities in the dataset, and obtain the longest directed path of each entity and the search path of each entity; Delete the included sub-paths from the search paths to obtain all non-inclusive search paths; Construct an entity-relationship matrix E through all non-inclusive search paths, and use the entity-relationship matrix E to construct implicit triples using the relationship transfer direction, where the entity-relationship matrix E is represented by the following formula: Among them, sig i,j = {0, 1}, D is the number of distinct entities in the dataset, sig i,j is the relationship between entity En i and En j . sig i,j = 0 indicates that there is no association between these two entities, sig i,j = 1 indicates that there is an association between these two entities. For the triple <En i , sig i,j , En j >, in the entity-relationship matrix E, the relationship sig i,j = 1, and the unit group composed of the entity En i,j corresponding to the relationship sig i,j and the entity En i and En j is an implicit triple.

3. The training method of the strong robust knowledge graph triple quality inspection network model based on the noisy data set according to claim 2, characterized in that The method for obtaining the confidence level of the implicit triples in obtaining the confidence levels of the triples includes Traverse all search paths, restore any search path through the entity-relationship matrix E, and obtain the longest search path; Based on the identification of the longest search path, calculate the confidence matrix of each entity triple on the longest search path; Calculate the confidence level of the constructed implicit triples through the confidence matrix of each entity triple on the longest search path, and each longest search path is independent of each other; The confidence level is represented by formula (3): Among them, r represents the confidence level, ← represents the pointing direction, and F refers to the number of the longest search paths including the triple <En i ,sig i,j, En j >, d k refers to the search depth of the current triple in the current affiliated triple, p k is the total length of the current search path, that is, the number of triples included. L is the maximum length of all the longest search paths. All confidence levels are normalized by the parameter L. D is the number of non-repeating entities in the dataset.

4. The training method of the strongly robust knowledge graph triple quality inspection network model based on a noisy data set according to claim 2, characterized in that, The method for constructing noise triples according to the source triples and implicit triples in the dataset includes that any triple obtained by randomly replacing the head entity <?,r,t>, replacing the relationship <h,?,t>, or replacing the tail entity <h,r,?> in the triple is a noise triple, and the source triples, implicit triples, and noise triples are retained in the dataset.

5. The training method of the strongly robust knowledge graph triple quality inspection network model based on the noisy data set according to claim 2, wherein The network model includes a TransR network, a residual network, and a BiLSTM network. The extraction of the initial features, static features, and internal association features of the triples through the network model includes Obtain the initial features of the source triples, implicit triples, and noise triples through the TransR network; Extract the static features of the source triples, implicit triples, and noise triples through the residual network; Extract the internal association features of the source triples, implicit triples, and noise triples through a multi-layer BiLSTM network.

6. The training method of the strongly robust knowledge graph triple quality inspection network model based on the noisy data set according to claim 5, characterized in that The method for pre-training the TransR model on source triples, implicit triples, and noisy triples includes taking the inner product of the embeddings of the source triples, implicit triples, and noisy triples with the confidence of the triples to obtain weighted feature vectors, where the weighted feature vectors are the initial features of the triples.

7. The training method of the strong robust knowledge graph triple quality inspection network model based on the noisy data set according to claim 2, characterized in that During feature modeling, the entity association relationship loss and the total loss are represented by formula (7) and formula (8) respectively: Among them, L EP represents the loss of entity association relationship, B represents the input batch size of the current training, a is the association depth of all batch samples, and y i represents the entity association relationship label, and p i represents the entity association relationship prediction probability, L represents the total loss, and y - represents the triple quality inspection label p- j represents the neural network's classification probability for each triple quality inspection; y j represents the entity association relationship label during the feature modeling process, and p j represents the neural network's prediction probability for each entity association relationship.

8. A method for quality inspection of strong-robust knowledge graph triples based on a noisy dataset, including Inputting the dataset to be quality-inspected into the network model with optimized model parameters obtained by the training method described in claims 1-7: Extracting the initial features, static features, and internal association features of the triples in the dataset to be quality-inspected through the network model, and aggregating them to obtain the fused features of the triples; According to the fused features of the triples, the network model predicts the association relationship between entities through a multi-label classification algorithm, and distinguishes the triples with no association relationship between entities from the triples with an association relationship between entities.

9. The strong robust knowledge graph triple quality inspection method based on a noisy data set according to claim 8, characterized in that The method for extracting the initial features, static features, and internal association features of the triples in the dataset to be quality-inspected through the network model includes Obtaining the initial features of the triples through the TransR network; Extracting the static features of the triples through the residual network; Extracting the internal association features of the triples through the multi-layer BiLSTM network.

10. The strong robustness knowledge graph triple quality inspection method based on a noisy dataset according to claim 9, characterized in that, The method for pre-training the TransR model on source triples, implicit triples, and noisy triples includes taking the inner product of the embeddings of the source triples, implicit triples, and noisy triples with the confidence of the triples to obtain weighted feature vectors, where the weighted feature vectors are the initial features of the triples.

Citation Information

Patent Citations

  • Knowledge graph triple reliability evaluation method, system and device and medium

    CN115238582A

  • Method for pre-training knowledge graph on the basis of structured context information

    WO2022057669A1