A Knowledge Graph Completion Method and System Based on Multi-View Contrastive Learning
The multi-view contrastive learning approach enhances knowledge graph completion by optimizing models with data preprocessing and graph enhancement, addressing cold-start issues and negative sampling biases to improve predictive accuracy and robustness.
Patent Information
- Application Number
- CN202411356982.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-09-27
AI Technical Summary
The existing knowledge graph completion system is difficult to obtain timely and effective user feedback during the cold start stage, resulting in limited model generalization ability and recommendation accuracy. The negative sampling deviation problem affects the exposure of new entities and relationships, and affects the completeness and balance of the graph.
Using multi-view comparison learning method, the total loss function is constructed for model training through data cleaning, graph enhancement, multi-view mask coding and low-dimensional vector mapping, combined with graph convolutional network and self-supervised learning, and the decoder is used to complete knowledge graphs.
It improves the prediction accuracy and generalization ability of the model, improves the completeness and balance of the knowledge graph, and enhances the overall performance and robustness of the system.
Smart Images

Figure CN119312907B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a knowledge graph completion method and system based on multi-view contrast learning. Background Art
[0002] With the advent of the big data era, as an important data organization form, knowledge graphs have been widely applied in fields such as search engines, intelligent question-answering systems, and recommendation systems. However, during the construction of knowledge graphs, due to the diversity and complexity of data sources, there are often a large number of missing information, namely the so-called "knowledge graph completion" problem. In recent years, research on this problem has gradually emerged and achieved certain research results, but there are still the following challenges: On the one hand, in the cold start stage, due to the lack of interaction data for new entities or new relationships, knowledge graph completion systems often cannot obtain timely and effective user feedback to form a closed-loop optimization. This means that the model is difficult to correct and improve the prediction results based on the actual behavior of users, which may limit the generalization ability and recommendation accuracy of the model. On the other hand, the negative sampling bias problem is manifested as that when constructing and training a knowledge graph completion model, unobserved relationships are usually used as negative samples. However, in the cold start scenario, this approach is likely to cause the model to misjudge "cold" entities or newly generated relationships, further exacerbating the tendency towards popular entities and relationships, which is not conducive to the exposure and recommendation of new entities or relationships, and affects the completeness and balance of the overall graph. Summary of the Invention
[0003] The main purpose of the present invention is to overcome the disadvantages and deficiencies of the prior art, and provide a knowledge graph completion method and system based on multi-view contrast learning, which can improve the completeness and balance of the graph, enhance the prediction accuracy and generalization ability of the model, and improve the overall performance and robustness of the entire system.
[0004] According to one aspect of the present invention, the present invention provides a knowledge graph completion method based on multi-view contrast learning, and the method includes the following steps:
[0005] S1: Receive the original knowledge graph data, and perform data cleaning and standardization processing, graph augmentation processing, multi-view mask encoding and masking processing, and low-dimensional vector mapping and feature regularization processing on the original knowledge graph data;
[0006] S2: Construct the total loss function of the model, and perform model training and optimization according to the total loss function to obtain a knowledge graph completion model;
[0007] S3: Invoke the knowledge graph completion model, and use a decoder to complete the knowledge graph.
[0008] Preferably, the data cleaning and standardization processing includes:
[0009] Deduplicate, detect and correct errors in the knowledge graph data in the dataset containing entities, relationships, and connection information between the two, ensure the accuracy of entities and relationships, and standardize the naming of entities and the types of relationships.
[0010] Preferably, the graph enhancement processing includes:
[0011] Perform image enhancement processing through three graph enhancement strategies to obtain enhanced graphs from three different perspectives; the three graph enhancement strategies are: node dropout processing, feature noise injection processing, and subgraph sampling processing; among them, in node dropout processing, randomly delete some nodes and their connecting edges, and automatically adjust the deletion ratio according to the number of relationships in the graph; in feature noise addition processing, add different degrees of noise to the relationship features and entity features of the nodes; in subgraph sampling processing, select some connected subgraphs from the original graph, and select neighbor nodes to join the subgraph according to the preset priority sampling strategy until the predetermined subgraph size is reached.
[0012] Preferably, the multi-view masked encoding and masking processing includes:
[0013] Apply the masked graph autoencoder MaskedGAE to each independent view for encoding, extract high-dimensional feature vector representations, and update the nodes using the propagation rules of the path-by-bit random masking strategy.
[0014] Preferably, the low-dimensional vector mapping and feature regularization processing includes:
[0015] Convert the output of the final layer of the encoder into a low-dimensional vector through a projection layer, perform feature compression and regularization, and use activation functions and clipping functions for non-linear transformation to convert the output of the final layer into the low-dimensional vector space.
[0016] Preferably, the total loss function of the constructed model includes:
[0017] Define independent contrastive loss functions for the three enhanced views respectively to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs at the same time; combine the three contrastive loss functions through weighted summation to obtain the total adversarial loss function; based on the difference between the actual labels and the predicted probabilities, construct the cross-entropy loss function for the relationship prediction task, and balance according to the adversarial loss function and the cross-entropy loss function of the prediction task using hyperparameters to obtain the total loss function.
[0018] Preferably, the knowledge graph completion model obtained by training and optimizing the model according to the total loss function includes:
[0019] The gradient of the total loss function with respect to the model parameters is calculated using automatic differentiation, and the Adam optimization algorithm is used to adjust the model parameters according to the gradient and the current learning rate to achieve adaptive parameter update; the performance of the model is evaluated using a validation set, and when the performance of the model meets the preset conditions, the training of the model is stopped to obtain a knowledge graph completion model.
[0020] Preferably, the calling of the knowledge graph completion model and using the decoder to complete the knowledge graph includes:
[0021] The topological structure of the graph is restored through the structure decoder in the MaskedGAE decoder, and the existence probability of edges is predicted using the interaction between node embeddings; the degree decoder in the MaskedGAE decoder predicts the degree of nodes through a regression method to supplement the graph structure information, and all the predicted new triple data are integrated back into the original knowledge graph to complete the original knowledge graph.
[0022] Preferably, the using the decoder to complete the knowledge graph includes:
[0023] A MaskedGAE decoder is constructed by a structure decoder and a degree decoder; among them, the structure decoder directly calculates the inner product of two node embedding vectors and converts it into an edge connection probability through a Sigmoid function:
[0024]
[0025] z i 、z j represent the embedding vectors of two nodes, and T represents transpose;
[0026] The degree decoder uses the Softmax function to output a discrete distribution representing the probability distribution of different values of the degree of node v:
[0027] g φ (z v ) = Softmax(MLP(z v ))
[0028] where z v is the encoded vector of node v, and MLP represents a multi-layer perceptron.
[0029] According to another aspect of the present invention, the present invention also provides a knowledge graph completion system based on multi-view contrastive learning, and the system includes:
[0030] A processing module, configured to receive the original knowledge graph data, and perform data cleaning and standardization processing, graph augmentation processing, multi-view mask encoding and masking processing, and low-dimensional vector mapping and feature regularization processing on the original knowledge graph data;
[0031] A training module for constructing a total loss function of a model, training and optimizing the model according to the total loss function to obtain a knowledge graph completion model;
[0032] A completion module for calling the knowledge graph completion model and using a decoder to complete the knowledge graph.
[0033] Beneficial effects: The knowledge graph completion method and system based on multi-view contrast learning proposed by the present invention combines the ideas of graph convolutional network (GCN) and multi-view contrast learning. By improving the existing knowledge graph completion model through multi-view contrast learning, the prediction accuracy and generalization ability of the model can be further enhanced, thereby improving the overall performance and robustness of the entire system. This method combines the ideas of deep learning and graph neural networks, combines graph convolutional networks with self-supervised learning, and realizes effective prediction of missing information. The present invention can be extended to various practical application scenarios, such as search engines, intelligent question answering systems, etc., thereby promoting the wide application and development of knowledge graph technology.
[0034] The features and advantages of the present invention will become clear by referring to the following drawings and the detailed description of the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flowchart of a knowledge graph completion method based on multi-view contrast learning;
[0036] Figure 2 is a flowchart of the operation of a knowledge graph completion method based on multi-view contrast learning;
[0037] Figure 3 is a schematic structural diagram of a knowledge graph completion system based on multi-view contrast learning. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] Embodiment 1
[0040] Figure 1 is a flowchart of a knowledge graph completion method based on multi-view contrast learning. Refer to Figure 1 and Figure 2 , this embodiment provides a knowledge graph completion method based on multi-view contrast learning, and the method includes the following steps:
[0041] S1: Receive the original knowledge graph data, and perform data cleaning and standardization, graph augmentation, multi-view masked encoding and masking, and low-dimensional vector mapping and feature regularization on the original knowledge graph data;
[0042] S2: Construct the total loss function of the model, and perform model training and optimization according to the total loss function to obtain a knowledge graph completion model;
[0043] S3: Invoke the knowledge graph completion model and use the decoder to complete the knowledge graph.
[0044] This method can improve the completeness and balance of the graph, enhance the prediction accuracy and generalization ability of the model, and improve the overall performance and robustness of the entire system.
[0045] Preferably, the data cleaning and standardization include:
[0046] Deduplicate, detect and correct errors in the knowledge graph data in the dataset containing entities, relationships, and connection information between the two, ensure the accuracy of entities and relationships, and standardize the naming of entities and the types of relationships.
[0047] Specifically, the important terms and constraints of this embodiment are as follows:
[0048] Graph augmentation: mainly through the augmentation of graph-structured data, create a richer and more expressive graph representation to improve the performance of graph neural networks (GNNs) or other graph-based learning algorithms. Graph augmentation includes various strategies such as adding edges, node feature enhancement, and subgraph sampling. Its focus is on how to make the graph structure and feature information more conducive to the learning of downstream tasks through graph construction or transformation. This kind of augmentation may be rule-based or learned.
[0049] Masked Graph Autoencoder (MaskedGAE): Masked Graph Autoencoder (MaskGAE) is a graph neural network (GNN) architecture that combines the idea of autoencoders and masking mechanisms to process graph data. The core idea of MaskGAE is to introduce masking in the graph autoencoding task as a self-supervised learning strategy, aiming to improve the model's representation ability on graph-structured data. Different from traditional graph autoencoders, MaskGAE introduces masking in the input graph, randomly selecting a part of nodes or edges for "hiding" or "masking". This means that the encoder can only learn the global graph representation based on the partially observable graph structure, and the task of the decoder is to predict or reconstruct the entire graph or the structure of the masked part based on this limited information.
[0050] Contrastive Learning: Contrastive Learning is a machine learning method, especially suitable for unsupervised or self-supervised learning scenarios. Its core idea is to compare data samples (usually presented in pairs), enabling the model to learn how to distinguish different samples in a high-dimensional feature space, thereby learning useful representations of the data. This method does not rely on traditional label information but utilizes the structure and similarity of the data itself. In contrastive learning, a loss function is calculated, that is, the similarity between positive example pairs is calculated and maximized, while the similarity between negative example pairs is minimized. This process ensures that the feature representations learned by the model can make similar samples close in the feature space and different samples far apart.
[0051] Automatic Differentiation: Automatic Differentiation (AD), also known as automatic derivative calculation, is a technique that uses computer algorithms to efficiently and accurately calculate the derivative values of differentiable functions at a certain point. This technique is widely used in scientific computing, engineering problems, economics, as well as machine learning and deep learning, especially in situations where gradient information is required for optimization, such as minimizing loss functions.
[0052] Adam Optimizer: The Adam Optimizer (Adaptive Moment Estimation) is an efficient gradient descent algorithm widely used in deep learning and other machine learning fields, which can effectively combine the advantages of momentum gradient descent and adaptive learning rate.
[0053] In this step, the original knowledge graph data is received. This data set contains entities, relationships, and the connection information between them. Therefore, the original graph can be represented as G=(V, E), where V is the set of nodes and E is the set of edges. Each node v∈V has a feature vector Deduplicate, detect and correct errors in the knowledge graph data to ensure the accuracy of entities and relationships. At the same time, standardize entity naming and relationship types for unified processing.
[0054] Preferably, the graph enhancement processing includes:
[0055] Image enhancement processing is performed through three graph enhancement strategies to obtain enhanced graphs from three different perspectives; the three graph enhancement strategies are: node discard processing, feature noise injection processing, and subgraph sampling processing; among them, in node discard processing, some nodes and their connecting edges are randomly deleted, and the deletion ratio is automatically adjusted according to the number of relationships in the graph; in feature noise addition processing, different degrees of noise are added to the relationship features and entity features of the nodes; in subgraph sampling processing, some connected subgraphs are selected from the original graph, and neighbor nodes are selected to join the subgraph according to a preset priority sampling strategy until the predetermined subgraph size is reached.
[0056] Specifically, for the organized graph data, this model proposes a fusion method that integrates three graph enhancement strategies: node discard, feature noise injection, and subgraph sampling. In node discard, some nodes and their connecting edges are randomly removed to mimic the situation of incomplete information, and the deletion ratio is automatically adjusted according to the number of relationships in the graph. Feature noise addition adds random variations to the features of the nodes, especially adding a small amount of noise to the relationship features to ensure the clarity of the relationships, but adding more noise to the entity features. Subgraph sampling selects a part of the connected subgraphs from the original graph and focuses on protecting the edges that contain important or rare relationships. After this series of enhancement steps, three enhanced graph spectra from different angles are obtained, providing a rich multi-angle view for the original knowledge graph. After graph enhancement processing, enhanced graphs from three different perspectives are obtained.
[0057] For the processed data, this model uses graph enhancement strategies of node discard, feature noise injection, and subgraph sampling to create graph representations from multiple different perspectives. This model integrates the three strategies through a new method: node discard, feature noise injection, and subgraph sampling. While enhancing the relationships, it creates a multi-perspective description of the graph spectra, specifically as follows:
[0058] Node discard processing: Traverse the entire knowledge graph and count the number of occurrences freq(r) of each relationship type r;
[0059] Set a threshold T to divide rare relationships and common relationships, which is set as the median of the relationship frequency distribution, and relationships below the median are regarded as rare relationships.
[0060] For each relationship type ro, determine different discard ratios according to whether it is rare Among them, the discard ratio of common relationships is defined as The discard ratio of rare relationships Generally, it should be set lower than the discard ratio of common relationships, that is;
[0061]
[0062] For each node v, the set of relationships R it participates inv For each relationship r in R v , determine whether to discard the node v and its associated edges according to the discard ratio .
[0063] Retention probability
[0064] Feature noise injection processing: Identify which of the node features represent entity attributes and which represent relationship features. For entity features, add Gaussian noise with a relatively large scale where σ e is a relatively large standard deviation. For relationship features, add a small amount of noise ensuring a relatively small standard deviation σ r < σ e to protect the clarity of the relationship, obtaining:
[0065] Entity feature noise: where x e represents the original relationship feature vector, x e ' represents the new vector after adding noise to the original relationship feature vector x e , and ∈ e represents the noise term added to the relationship feature vector.
[0066] Relationship feature noise: where x e represents the original relationship feature vector, x e ' represents the new vector after adding noise to the original relationship feature vector x e , and ∈ e represents the noise term added to the relationship feature vector.
[0067] Subgraph sampling processing: Apply the PageRank algorithm to calculate the PageRank (node importance) value PR i = 1 / N, where N is the total number of nodes in the graph. Update the PageRank value of each node according to the iterative formula of PageRank until convergence or the maximum number of iterations is reached:
[0068]
[0069] where d is the damping factor, N(i) is the set of neighbor nodes of node i, and L(j) is the out-degree of node j.
[0070] Select the top m nodes with the highest PageRank values as seed nodes, where m is determined according to the proportion of the required subgraph size. Add all the seed nodes and their directly connected edges to the subgraph, and use the degree-first sampling strategy to select their neighbor nodes to join the subgraph until the predetermined subgraph size is reached.
[0071] Preferably, the multi-view mask encoding and masking processing includes:
[0072] Apply the Masked Graph Autoencoder (MaskedGAE) to each independent view for encoding to extract high-dimensional feature vector representations, and update the nodes using the propagation rules of the path-wise random masking strategy.
[0073] Specifically, by applying MaskedGAE to each independent view for encoding, high-dimensional feature vector representations are extracted. This step uses a masking strategy to mask part of the edge and node attributes on the graph data before inputting it to the GAE, simulating the situation where the relationship information in the graph data is incomplete, converting the graph of each view into a partially observable scenario. MaskedGAE endeavors to capture and reconstruct the key features of the original graph and update the node representations on this basis, considering the influence of the masked neighbor nodes, and updating the nodes using the propagation rules of the path-wise random masking strategy.
[0074] First, according to the set path-wise random masking strategy, part of the edge or node attributes need to be "hidden" in the graph data.
[0075] Assume that the probability transition matrix of the random walk is P. For each step t, the transition probability from node i to node j is P ij .
[0076] For each root node r, perform n walk independent first-order or second-order Markov chain random walks, and the length of each walk is l walk steps, which can be denoted as:
[0077]
[0078] Construct the masked edge set E mask :
[0079]
[0080] During the propagation process, the corresponding adjustment needs to be made to the neighbor node set, only considering the unmasked neighbor nodes and their corresponding edges.
[0081] The propagation rule of each layer of the graph neural network can be expressed by the following formula:
[0082]
[0083] Wherein, is the hidden state of node v at the l-th layer, W l is the weight matrix of this layer, N(v) represents the neighbor set of node v, C uv is the normalization coefficient (such as the inverse of the degree matrix), σ is the non-linear activation function, b l is the bias term of the l-th layer.
[0084] Therefore, for the update of the hidden state of node v at the (l + 1)-th layer after applying the masking strategy, it can be modified as:
[0085]
[0086] Wherein, N M (v) represents the actual neighbor set of node v after applying the masking strategy, including those neighbor nodes that are not masked during the random walk process.
[0087] Preferably, the low-dimensional vector mapping and feature regularization processing include:
[0088] Converting the output of the final layer of the encoder into a low-dimensional vector through a projection layer, implementing feature compression and regularization, and performing non-linear transformation using an activation function and a clipping function to convert the output of the final layer into the low-dimensional vector space.
[0089] Specifically, the final output of the encoder is converted into a low-dimensional vector z v , implementing feature compression and regularization, and performing non-linear transformation using the ReLU activation function and the clip clipping function to convert the output of the final layer into the low-dimensional vector space.
[0090] Mapping the output of the last layer of the GCN encoder to the low-dimensional vector space:
[0091]
[0092] Where L represents the number of layers of the GCN, and Project(·) is a dimensionality reduction operation of a fully connected layer. In order to map the output of the last layer of the GCN encoder to the low-dimensional vector space and achieve certain feature compression and regularization, the projection layer (Projector) adopts the following form:
[0093]
[0094] Wherein, W project is a weight matrix for converting the high-dimensional Mapped to a low-dimensional space; b project is the bias term; norm(·, p) represents the norm normalization operation, and here the L p norm is used for feature scaling; the max(·, ∈) operation is optional and is used to ensure that the projected vector value will not be too small (close to zero).
[0095] To limit the excessive results, by adding the activation function ReLU to limit the maximum value of the output, combined with norm normalization, ensuring non-negativity while limiting the maximum value, so norm normalization is performed first and the maximum-minimum limit is applied before using ReLU, and the clipping function (clip) is used to process the projection z v for processing:
[0096]
[0097] where the clip(x, lower, upper) function is defined as:
[0098]
[0099] Thus, the clip(., ∈, C) operation limits the output range between ∈ and C, where ∈ is the lower limit of the minimum value and C is the upper limit of the maximum value, which is used to directly "discard" or limit the excessive data, and the projected result z is obtained through processing v .
[0100] Preferably, the total loss function of the constructed model includes:
[0101] Separate contrastive loss functions are defined for the three enhanced views to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs at the same time; the three contrastive loss functions are combined by weighted summation to obtain the total adversarial loss function; based on the difference between the actual label and the predicted probability, a cross-entropy loss function for the relationship prediction task is constructed, and according to the adversarial loss function and the cross-entropy loss function of the prediction task, it is balanced by hyperparameters to obtain the total loss function
[0102] Specifically, the loss form adopted by this model is:
[0103]
[0104] where z v is the projected representation of the target node, represents the positive sample, is the set of negative samples, is the representation of the negative sample node, and τ is the temperature parameter
[0105] Since the model has contrastive losses for multiple views, independent adversarial losses are defined for each dimension here, and they are finally combined to obtain the total adversarial loss. For these three dimensions, denoted as A, B, and C respectively, the corresponding projection representations are known The corresponding representations of positive samples are and The set of negative samples is and where the negative sample node representations within each set are respectively known
[0106] At this time, the adversarial losses for the three dimensions can be defined respectively as follows:
[0107]
[0108] Then, the total adversarial loss can be used to combine the losses of each dimension by weighted summation or simple addition:
[0109]
[0110] where λ A , λ B and λ C are the weight factors for the corresponding dimensions, used to balance the importance of losses in different dimensions.
[0111] By complementing and training the relationships between nodes, a prediction task loss is constructed, and the cross-entropy loss of the relationship completion task is calculated:
[0112]
[0113] where (i, j) represents a pair of nodes in the graph, and V is the set of all considered relationship pairs. y ij represents the true label of this pair of nodes. If it is a binary variable, then y ij = 1 indicates that there is a relationship between node i and node j, and y ij = 0 indicates that there is no relationship. p ij is the probability that the model predicts there is a relationship between node i and node j, ranging from 0 to 1.
[0114] The total loss includes the total adversarial loss and the prediction task loss. The weights of the encoder are updated through backpropagation, and the hyperparameter λ is used to balance the weights of the two loss terms:
[0115]
[0116] Preferably, the model training and optimization according to the total loss function to obtain the knowledge graph completion model includes:
[0117] The gradient of the total loss function with respect to the model parameters is calculated using automatic differentiation, and the Adam optimization algorithm is used to adjust the model parameters according to the gradient and the current learning rate, achieving adaptive parameter update: the performance of the model is evaluated using the validation set, and when the performance of the model meets the preset conditions, the training of the model is stopped to obtain the knowledge graph completion model.
[0118] Specifically, the gradient of the total loss with respect to the model parameters is calculated using automatic differentiation, and the Adam optimization algorithm is used to adjust the model parameters according to the gradient and the current learning rate, achieving efficient and adaptive parameter update. According to the calculated total loss function, backpropagation is performed to update the network parameters. At the end of each epoch, the validation set is used to evaluate the model performance, and the performance metric (accuracy) on the validation set is monitored. When the metric no longer improves significantly, the training is stopped.
[0119] When performing backpropagation, automatic differentiation is used to calculate the gradient of the total loss L with respect to all trainable parameters (mainly the encoder weights W enc and possibly the decoder weights W dec of the gradient), that is where W represents the set of all parameters, and the Adam (Adaptive Moment Estimation) optimization algorithm is used to adaptively adjust the parameter weights, and the model parameters are updated according to the calculated gradient and the current learning rate, η is the learning rate.
[0120] Thus, according to the calculated total loss function, backpropagation is performed to update the network parameters. At the end of each epoch, the validation set is used to evaluate the model performance to ensure the generalization of the model. The performance metric (accuracy) on the validation set is monitored. When the metric no longer improves significantly, the training is stopped, enabling the model to gradually learn to distinguish positive examples from negative examples and optimize the quality of the feature representation to obtain the trained MaskedGAE model.
[0121] The trained MaskedGAE model is used to process the test set. Specifically, the data in the test set is input into the model. In particular, with the help of the MaskedGAE decoder part, this component is responsible for predicting and generating the missing entity-relationship pairs based on the patterns and rules learned by the model. This process is equivalent to using the model to fill in the blank links in the knowledge graph, thereby generating a more complete and information-rich version of the knowledge graph compared to the original test set. Through this step, the ability of the model to complete the knowledge graph on unknown data is effectively evaluated, and the effect of the model in completing the knowledge graph can be intuitively observed.
[0122] Preferably, the invoking the knowledge graph completion model and using the decoder to complete the knowledge graph includes:
[0123] Restore the topology of the graph through the structure decoder in the MaskedGAE decoder, and predict the existence probability of edges by leveraging the interaction between node embeddings; predict the degree of nodes through regression using the degree decoder in the MaskedGAE decoder to supplement the graph structure information, and integrate all the predicted new triple data back into the original knowledge graph to complete the original knowledge graph.
[0124] Preferably, the use of the decoder for knowledge graph completion includes:
[0125] Construct a MaskedGAE decoder from the structure decoder and the degree decoder; among them, the structure decoder directly calculates the inner product of two node embedding vectors and transforms it into an edge connection probability through the Sigmoid function:
[0126]
[0127] z i 、z j represent the embedding vectors of two nodes, and T represents the transpose;
[0128] The degree decoder uses the Softmax function to output a discrete distribution representing the probability distribution of different values of the degree of node v:
[0129] g φ (z v ) = Softmax(MLP(z v ))
[0130] where z v is the encoded vector of node v, and MLP represents a multi-layer perceptron.
[0131] Specifically, construct the decoder of MaskedGAE from the structure decoder (StructureDecoder) and the degree decoder (DegreeDecoder).
[0132] The structure decoder (Structure Decoder) adopts the form of direct inner product plus the Sigmoid function:
[0133]
[0134] This form directly calculates the inner product of two node embedding vectors and transforms it into an edge connection probability through the Sigmoid function.
[0135] The degree decoder (DegreeDecoder) adopts the following form:
[0136] g φ (z v ) = Softmax(MLP(zv ))
[0137] where z v is the encoded vector of node v. After extracting features through a multi-layer perceptron (MLP), the Softmax function is used to output a discrete distribution, representing the probability distribution of different values of the degree of node v. However, in this model, considering that the degree is often a numerical value rather than a category, a regression form is also adopted:
[0138]
[0139] where is the predicted degree of node v, which is obtained by passing the node embedding through an MLP and then through a linear layer plus a bias term b.
[0140] Through the decoder, a list of candidate entities is generated based on their proximity in the embedding space or other complex patterns, and new triple entities are constructed. All the newly predicted triples are integrated back into the original knowledge graph, and these prediction results are evaluated through a validation set or an independent test set to complete the original knowledge graph.
[0141] The knowledge graph completion method based on multi-view contrast learning proposed in this embodiment combines the ideas of graph convolutional network (GCN) and multi-view contrast learning. By improving the existing knowledge graph completion model through multi-view contrast learning, the prediction accuracy and generalization ability of the model can be further enhanced, thereby improving the overall performance and robustness of the entire system. This method combines the ideas of deep learning and graph neural networks, combines graph convolutional networks with self-supervised learning, and realizes effective prediction of missing information. The method of this embodiment can be extended to various practical application scenarios, such as search engines, intelligent question answering systems, etc., thus promoting the wide application and development of knowledge graph technology.
[0142] Embodiment 2
[0143] Figure 3 is a schematic structural diagram of a knowledge graph completion system based on multi-view contrast learning. As Figure 3 shown, this embodiment provides a knowledge graph completion system based on multi-view contrast learning, and the system includes:
[0144] A processing module 301, configured to receive the original knowledge graph data, and perform data cleaning and standardization processing, graph augmentation processing, multi-view mask encoding and masking processing, and low-dimensional vector mapping and feature regularization processing on the original knowledge graph data;
[0145] A training module 302, configured to construct the total loss function of the model, perform model training and optimization according to the total loss function, and obtain a knowledge graph completion model;
[0146] The completion module 303 is used to call the knowledge graph completion model and utilize the decoder to complete the knowledge graph.
[0147] The specific implementation processes of the functions implemented by each module in this Embodiment 2 are the same as those in Embodiment 1 and will not be elaborated here.
[0148] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural transformation made under the concept of the present invention by using the content of the specification and drawings of the present invention, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. A knowledge graph completion method based on multi-view contrastive learning, characterized in that, The method includes the following steps: S1: Receive the original knowledge graph data, and perform data cleaning and standardization processing, graph augmentation processing, multi-view masked encoding and masking processing, and low-dimensional vector mapping and feature regularization processing on the original knowledge graph data; S2: Construct the total loss function of the model, and perform model training and optimization according to the total loss function to obtain a knowledge graph completion model; S3: Invoke the knowledge graph completion model and use the decoder to complete the knowledge graph; Among them, the graph augmentation processing includes: Perform image augmentation processing through three graph augmentation strategies to obtain three enhanced graphs with different perspectives; the three graph augmentation strategies are: node dropping processing, feature noise injection processing, and subgraph sampling processing; among them, in the node dropping processing, randomly delete some nodes and their connecting edges, and automatically adjust the deletion ratio according to the number of relationships in the graph; in the feature noise addition processing, add different degrees of noise to the relationship features and entity features of the nodes; in the subgraph sampling processing, select some connected subgraphs from the original graph, and select neighbor nodes to join the subgraph according to a preset preferential sampling strategy until the predetermined subgraph size is reached; The construction of the total loss function of the model includes: Define independent contrast loss functions for the three enhanced views respectively to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs at the same time; combine the three contrast loss functions by weighted summation to obtain the total adversarial loss function; based on the difference between the actual label and the predicted probability, construct the cross-entropy loss function of the relationship prediction task, and balance according to the adversarial loss function and the cross-entropy loss function of the prediction task using hyperparameters to obtain the total loss function; The obtaining of the knowledge graph completion model by performing model training and optimization according to the total loss function includes: Use automatic differentiation to calculate the gradient of the total loss function with respect to the model parameters, and use the Adam optimization algorithm to adjust the model parameters according to the gradient and the current learning rate to achieve adaptive parameter update; use the validation set to evaluate the model performance, and stop training the model when the model performance meets the preset conditions to obtain the knowledge graph completion model.
2. The method according to claim 1, wherein The data cleaning and standardization processing includes: Perform deduplication, error detection and correction on the knowledge graph data in the dataset containing entities, relationships, and the connection information between the two to ensure the accuracy of entities and relationships, and standardize the naming of entities and the types of relationships.
3. The method according to claim 1, wherein The multi-view masked encoding and masking processing includes: Apply the masked graph autoencoder MaskedGAE to each independent view for encoding, extract high-dimensional feature vector representations, and update the nodes using the propagation rules of the path-by-bit random masking strategy.
4. The method according to claim 1, wherein The low-dimensional vector mapping and feature regularization processing includes: Convert the output of the final layer of the encoder into a low-dimensional vector through a projection layer, implement feature compression and regularization, and perform non-linear transformation using an activation function and a clipping function to convert the output of the final layer into a low-dimensional vector space.
5. The method according to claim 4, characterized in that, The invoking of the knowledge graph completion model and using the decoder to complete the knowledge graph includes: Restore the topology of the graph through the structure decoder in the MaskedGAE decoder, and use the interaction between node embeddings to predict the existence probability of edges; use the degree decoder in the MaskedGAE decoder to predict the degree of nodes by regression method to supplement the graph structure information, and integrate all the predicted new triple data back into the original knowledge graph to complete the original knowledge graph.
6. The method according to claim 5, characterized in that The knowledge graph completion using the decoder includes: Construct a MaskedGAE decoder from the structure decoder and the degree decoder; among them, the structure decoder directly calculates the inner product of two node embedding vectors and transforms it into an edge connection probability through the Sigmoid function: zi and zj represent the embedding vectors of two nodes, and T represents the transpose; The degree decoder uses the Softmax function to output a discrete distribution, representing the probability distribution of different values of the degree of node v: Among them, zv is the encoding vector of node v, and MLP represents a multi-layer perceptron.
7. A knowledge graph completion system based on multi-view contrastive learning, which is used to execute the knowledge graph completion method based on multi-view contrastive learning as described in any one of claims 1-6, characterized in that, The system includes: A processing module for receiving the original knowledge graph data and performing data cleaning and standardization processing, graph augmentation processing, multi-view masking encoding and masking processing, and low-dimensional vector mapping and feature regularization processing on the original knowledge graph data; A training module for constructing the total loss function of the model, training and optimizing the model according to the total loss function to obtain a knowledge graph completion model; a completion module for calling the knowledge graph completion model and using the decoder to complete the knowledge graph.
Citation Information
Patent Citations
Graph comparison recommendation method based on knowledge graph
CN117892815A
Knowledge graph structure optimization method based on graph contrast learning in self-supervised scene
CN118210929A