Uncertain knowledge graph reasoning method based on semi-supervised confidence distribution learning
By adopting semi-supervised confidence distribution learning and meta-self-training methods in the uncertainty knowledge graph, the problem of imbalance in the confidence distribution of triple tuples in the knowledge graph is solved, and the accuracy of embedding quality and confidence prediction is improved.
Patent Information
- Application Number
- CN202510270084.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing uncertainty knowledge graph inference methods deal with the problem of imbalance in confidence distribution of triplets, it is difficult to fit the model to samples with lower confidence, resulting in a decrease in embedding quality.
A method based on semi-supervised confidence distribution learning is proposed. By converting the triple confidence into a confidence distribution, and using the two modules CDL-RL and PCDG for meta-self-training, high-quality pseudo-labeled data, expanding the training data, and solving the problem of unbalanced confidence distribution.
Through semi-supervised learning and meta-self-training, the problem of unbalanced confidence distribution in the uncertain knowledge graph can be effectively dealt with, improving the quality of knowledge graph embedding and the accuracy of confidence prediction.
Smart Images

Figure CN120104810A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of uncertainty knowledge graph reasoning, and specifically relates to an uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning. Background Art
[0002] The Knowledge Graph (KG) is a multi-relational graph that uses triples to describe knowledge. Each triple is represented in the form of (subject, predicate, object), such as (Michael Jordan, Nationality, US). The Uncertain Knowledge Graph (UKG) associates each triple in the knowledge graph with a confidence score, which indicates the possibility that the triple is true. This setting is conducive to more accurate knowledge representation and reasoning in the real world. The Uncertain Knowledge Graph can be formally represented as a set of four-tuples: G = {(h, r, t, s)|h, t∈ε, s∈
[0003] [0,1]}, where ε and They represent the set of entities and relations respectively, and s is the confidence score used to describe the uncertainty of the triple. Typical representatives of uncertainty knowledge graphs include NELL, ConceptNet, etc.
[0004] The knowledge graph in the real world is incomplete because new knowledge is always generated over time, and the same is true for the uncertainty knowledge graph. Therefore, many uncertainty knowledge graph reasoning methods based on uncertainty knowledge graph embedding learning have emerged to perform confidence prediction and link prediction to achieve uncertainty knowledge graph completion, including UKGE, PASSLEAF, UKGsE, BEUrRE and other methods. Uncertainty knowledge graph embedding learning aims to learn the representation of entities and relationships in low-dimensional space and retain the information of graph structure and confidence in the low-dimensional representation. However, the above methods ignore the fact that the confidence distribution of triples in most uncertainty knowledge graphs is extremely unbalanced during the learning process, that is, only high-confidence triples are retained. For example, NELL only contains triplets with confidence greater than 0.9. Learning this unbalanced data will result in the inability of the embedding-based uncertainty knowledge graph reasoning method to fit the model to samples with relatively low confidence, thereby reducing the quality of the generated uncertainty knowledge graph embedding.
[0005] In order to solve this problem, the present invention proposes an uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning. In the uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning of the present invention, each triple confidence is converted into a confidence distribution. This method can make full use of the supervised information of a few confidences or unseen confidences in the data, and generate reliable confidences for unlabeled data (triples without confidence scores) to expand the training data, so it can handle the problem of unbalanced triple confidence distribution in the learning process. Summary of the invention
[0006] Technical problem: The present invention provides an uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning, which can solve the problem of uncertainty knowledge graph completion in the scenario of unbalanced confidence distribution of triples in uncertainty knowledge graphs in the real world.
[0007] Technical solution: The uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning of the present invention includes two key modules: a confidence distribution learning-based relational learner (CDL-RL) and a pseudo labeled data generator (PCDG). This method first converts all the confidences of the labeled data (i.e., triples with confidence) in the training set into confidence distributions. Then CDL-RL uses both the labeled data and the pseudo labeled data generated by PCDG to learn the embedding of the uncertainty knowledge graph. PCDG generates high-quality pseudo confidence labels for unlabeled data to obtain pseudo labeled data, and CDL-RL uses these pseudo labeled data in an iterative manner and further improves its performance. CDL-RL has the same structure as PCDG, but the training process is different. CDL-RL optimizes by minimizing the loss of labeled data and pseudo labeled data in confidence prediction and link prediction, while PCDG uses the performance of CDL-RL after using the pseudo labeled data generated by PCDG as its own meta-learning goal. The uncertain knowledge graph reasoning method based on semi-supervised confidence distribution learning performs meta-self-training by iteratively training CDL-RL and PCDG.
[0008] The uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning proposed in the present invention comprises the following steps:
[0009] S1, convert all triple confidences in the uncertainty knowledge graph training data into confidence distribution.
[0010] S2, using CDL-RL to simultaneously learn the embedding of the uncertain knowledge graph on labeled data and pseudo-labeled data generated by PCDG.
[0011] S3, uses PCDG to generate high-quality pseudo-confidence labels for unlabeled data.
[0012] S4, iteratively train CDL-RL and PCDG using meta-self-training until convergence.
[0013] S5: The data to be completed is passed into the trained CDL-RL for reasoning to achieve confidence prediction and link prediction.
[0014] In the uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning of the present invention, in step S1, the present invention defines the confidence distribution as a discrete distribution. Since the confidence interval is [0,1], the present invention directly sets the granularity of the confidence label set to The ordered confidence label set is For a given quadruple (h, r, t, s) in an uncertainty knowledge graph, the confidence distribution of (h, r, t) is defined as where s i ∈[0,1] is the confidence label The confidence distribution of the present invention is a Gaussian distribution. Generated, where σ is the standard deviation and s is the mean, so s has the highest descriptiveness. Therefore, in this step, each piece of knowledge in the uncertainty knowledge graph can be represented as a four-tuple (h, r, t, s), where is the confidence distribution.
[0015] In the uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning of the present invention, in step S2, CDL-RL is used to learn the embedding of entities and relationships in the uncertainty knowledge graph. The input of CDL-RL is the four-tuple l = (h, r, t, s) obtained in step S1. CDL-RL has three learning objectives: 1) minimize the difference between the predicted confidence distribution and the true confidence distribution; 2) minimize the difference between the predicted confidence and the true confidence; 3) accurately predict the tail entity given the head entity and the tail entity. The first two objectives are used to complete confidence prediction, and the last objective is used for link prediction. Specifically, given a triple (h, r, t), the embeddings of h, r and t are concatenated and passed into a two-layer fully connected neural network, which outputs an n+1 dimensional vector. The Softmax function is then applied as the activation function. Prediction confidence distribution of the triple (h, r, t) It can be calculated by the following formula:
[0016]
[0017] where || represents the connection between embeddings, FCN 1 is a function that converts the connection between embeddings into an n+1-dimensional vector using a fully connected network, and Softmax(·) is an activation function used to map an n+1-dimensional vector to an n+1-dimensional probability distribution.
[0018] The present invention uses Kullback-Leibler divergence (KL divergence) to measure the similarity between the predicted distribution and the true confidence distribution, and takes it as one of the optimization goals of the uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning. Assume that l i is the i-th quadruple in the training data, then this part of the loss function is defined as:
[0019]
[0020] where s i and Represents l i The true confidence distribution and predicted confidence distribution of , s i a and represent the ath element in the true confidence distribution and the predicted confidence distribution, respectively. Represents the training set.
[0021] In order to further improve the accuracy of confidence prediction, the present invention calculates the mean square error (MSE) loss function between the expected and true confidence of the predicted confidence distribution for each quadruple, and the calculation method is as follows:
[0022]
[0023] where s i Yes i The true confidence level, is the prediction confidence distribution The present invention takes this expectation as l i prediction confidence.
[0024] Will and Combined, we get the loss function of the confidence prediction task
[0025]
[0026] where w 1 and w 2 is a trade-off factor used to adjust and proportion.
[0027] In order to enable the uncertain knowledge graph reasoning method based on semi-supervised confidence distribution learning to solve the link prediction problem, it is necessary to accurately evaluate the ranking of possible tail entities under a given query (h, r, ?).
[0028] The present invention designs an edge-based ranking loss function to optimize the performance of the method on the link prediction task. Here, another two-layer fully connected neural network is used to calculate the ranking score of a triple. The difference from the fully connected neural network used for confidence prediction is that the output of this fully connected neural network is a scalar rather than a vector, and the scalar is normalized using the Sigmoid activation function. The ranking score of the triple is calculated as follows:
[0029]
[0030] in is the predicted triple ranking score, FCN 2 is a function that converts the connection between embeddings into a scalar using a fully connected network, and Sigmoid(·) is an activation function used to map a scalar to a probability in the range of 0 to 1. The edge-based ranking loss function for optimizing link prediction is to increase the ranking score difference between positive and negative samples so that the ranking score of the positive sample is higher than that of the negative sample. Given a positive sample, the present invention generates multiple negative samples by replacing the head or tail entity of the positive sample with a randomly selected entity. This edge-based ranking loss function is defined as:
[0031]
[0032] in and are the ranking scores of positive samples and negative samples respectively, is the set of negative samples, [x] + =max[0,x] is the standard hinge loss and γ is the margin size.
[0033] Use confidence prediction loss at the same time and link prediction loss When training, it is necessary to balance the ratio of the two tasks in the training process to avoid one task dominating the training. The present invention uses uncertainty weights to dynamically adjust the ratio of each task in the training process. The loss function of the kth task is adjusted to:
[0034]
[0035] where λ k is the noise parameter of the kth task, The weight of the kth task can be adjusted dynamically so that tasks with higher uncertainty have lower confidence, logλ k is a regularization term that prevents the loss from increasing due to λ k The increase of is the loss function of the kth task. The loss function of CDL-RL is calculated as:
[0036]
[0037] where λ 1 and λ 2 are the noise parameters of confidence prediction and link prediction, respectively, and θ represents the parameter of CDL-RL. When both labeled data and pseudo-labeled data of PCDG are used (which will be explained in detail in steps S3 and S4), the loss function of confidence prediction is redefined as:
[0038]
[0039] in is a labeled dataset, is a pseudo-labeled dataset, s j and denote the pseudo-label confidence distribution and prediction confidence distribution of the j-th pseudo-label quadruple of the pseudo-label dataset, respectively, and w p is the weight of pseudo-labeled data, s j b and represent the bth element in the pseudo confidence distribution and the corresponding predicted confidence distribution, respectively.
[0040] In the uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning of the present invention, in step S3, PCDG is used to generate high-quality pseudo-confidence labels for unlabeled data, thereby obtaining pseudo-labeled data. PCDG is based on the idea of meta-learning and solves the error accumulation problem in traditional self-training methods. The structure of PCDG is the same as that of CDL-RL, and its parameters are represented as η. During the training process, PCDG first generates a pseudo-confidence distribution for unlabeled data. These pseudo-labeled data contain the gradient information of PCDG, so that PCDG can be optimized. These pseudo-labeled data are represented as We then evaluate CDL-RL on labeled data and The performance after performing a gradient descent step on the CDL-RL is used to check whether CDL-RL is better trained. The smaller the loss of CDL-RL, the better its performance. Therefore, the training goal of PCDG is to minimize the loss of CDL-RL on the labeled data after one update. The training goal of PCDG can be expressed as:
[0041]
[0042] in When the parameter of CDL-RL is θ + At The loss on θ + Is CDL-RL in and The parameters after a gradient descent process is performed on , and its calculation method can be expressed as:
[0043]
[0044] Where α is the learning rate, When the parameter of CDL-RL is θ, and Finally, the loss function of PCDG can be explicitly expressed as:
[0045]
[0046] In the uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning of the present invention, in step S4, the specific process of meta-self-training is as follows: first, the labeled data The input is sent to CDL-RL, which uses CDL-RL to learn the embedding of the uncertainty knowledge graph. CDL-RL uses the loss function of step S2 to optimize itself by minimizing the loss on the confidence prediction and link prediction tasks. When the embedding of entities and relations gradually stabilizes, PCDG is trained to generate high-quality pseudo-confidence labels for unlabeled data. It uses the CDL-RL used in step S3 to train the labeled data and pseudo-labeled data. After training PCDG for a while, a screening strategy is applied to the pseudo-labeled data generated by PCDG, that is, in the labeled data, the original confidence of each triple should be in the leading position in the transformed confidence distribution. Therefore, in the pseudo-labeled data, if the highest descriptive degree of a pseudo-confidence label is greater than a fixed threshold, the corresponding pseudo-confidence distribution is regarded as high quality, and the pseudo-labeled data is selected for training CDL-RL, while other pseudo-labeled data are removed. These selected pseudo-labeled data are then Input into CDL-RL to enhance its training process. After that, CDL-RL and PCDG continuously repeat the above process and optimize each other, that is, PCDG provides pseudo-labeled data for CDL-RL, and CDL-RL provides meta-learning objectives for PCDG. This iterative training process is complete meta-self-training. The present invention then continuously optimizes CDL-RL and PCDG until convergence.
[0047] In the uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning of the present invention, in step S5, after the model training is completed through steps S1-S4, the data to be completed is passed into the trained CDL-RL for reasoning to obtain the prediction confidence and ranking score. The confidence prediction task and the link prediction task can be achieved according to the prediction confidence and ranking score. For the confidence prediction task, the given triple is passed into CDL-RL, and the prediction confidence output by CDL-RL is directly used as the result of the confidence prediction; and for the link prediction task, for the given head entity and relationship, all entities in the uncertainty knowledge graph are used as candidate tail entities, and constitute candidate triples with the given head entity and relationship. For each candidate triple, use CDL-RL to calculate its ranking score, and the candidate triple with a high ranking score is the possible correct completion item.
[0048] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, an uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning is implemented.
[0049] A computer-readable storage medium stores computer instructions, which, when executed by a processor, implement an uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning.
[0050] Compared with the prior art, the advantages of the present invention are as follows.
[0051] 1. The uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning proposed in the present invention applies meta-self-training in uncertainty knowledge graph embedding learning to generate reliable confidence for unlabeled data, and makes full use of unlabeled data to expand the training data, thereby solving the problem of unbalanced confidence distribution.
[0052] 2. This paper proposes a new confidence distribution learning strategy to convert triple confidence into confidence distribution, which is conducive to capturing the supervisory information of few confidences or unseen confidences in labeled data.
[0053] 3. The method of the present invention has been successfully applied on the uncertainty knowledge graph datasets CN15k and NL27k. In the confidence prediction scenario, the method is trained on the training set, and the mean square error MSE is used as the evaluation index to measure the method effect on the test set. The results show that the mean square error reaches 0.037 on the CN15k dataset and 0.012 on the NL27k dataset, indicating that the method has significant effects in the confidence prediction scenario. At the same time, in the link prediction scenario, the method is trained on the training set, and the weighted average reciprocal ranking WMRR is used as the evaluation index to measure the method effect on the test set. The results show that the weighted average reciprocal ranking reaches 0.207 on the CN15k dataset and 0.731 on the NL27k dataset, indicating that the method has significant effects in the link prediction scenario. The indicators of the method of the present invention on different tasks on the two datasets are superior to other uncertainty knowledge graph reasoning methods. In summary, the method of the present invention can effectively learn the embedding of the uncertainty knowledge graph and perform uncertainty knowledge graph reasoning, thereby achieving uncertainty knowledge graph completion. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a schematic diagram of the overall framework of the present invention.
[0055] Figure 2 It is a specific structural diagram of CDL-RL and PCDG of the present invention. DETAILED DESCRIPTION
[0056] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0057] Example 1
[0058] A method for uncertainty knowledge graph reasoning based on semi-supervised confidence distribution learning, such as Figure 1 As shown, the following steps are included:
[0059] Step S1: Convert all triple confidences in the given uncertainty knowledge graph training data into confidence distribution. This embodiment defines the confidence distribution as a discrete distribution. This embodiment sets the granularity of the confidence label set to The ordered confidence label set is For a given quadruple (h, r, t, s) in the uncertainty knowledge graph, the confidence distribution of (h, r, t) is defined as where s i ∈[0,1] is the confidence label The description degree of (h, r, t). This embodiment uses Gaussian distribution Generate a confidence distribution, where σ is the standard deviation and s is the mean, so s has the highest descriptiveness. In this step, each piece of knowledge in the uncertainty knowledge graph z is represented as a four-tuple (h, r, t, s), where is the confidence distribution. In this embodiment, n is set to 100, so there are 101 confidence labels in total.
[0060] Step S2: Use CDL-RL to learn the embedding of entities and relations in the uncertain knowledge graph. CDL-RL, such as Figure 2 As shown, its input is the four-tuple l = (h, r, t, s) obtained in step S1. CDL-RL has three learning objectives: 1) minimize the difference between the predicted confidence distribution and the true confidence distribution; 2) minimize the difference between the predicted confidence and the true confidence; 3) accurately predict the tail entity given the head entity and the tail entity. The first two objectives are used to complete the confidence prediction, and the last objective is used for link prediction. Specifically, given a triple (h, r, t), the embeddings of h, r and t are concatenated and passed into a two-layer fully connected neural network, which outputs an n+1 dimensional vector. The Softmax function is then applied as the activation function. Prediction confidence distribution of the triple (h, r, t) It can be calculated by the following formula:
[0061]
[0062] where || represents the connection between embeddings, FCN 1 is a function that converts the connection between embeddings into an n+1-dimensional vector using a fully connected network. Softmax(·) is an activation function that maps an n+1-dimensional vector to an n+1-dimensional probability distribution. This embodiment uses the Kullback-Leibler divergence (KL divergence) to measure the similarity between the predicted distribution and the true confidence distribution, and uses it as one of the optimization objectives. Assume that l i is the i-th quadruple of the training data, then this part of the loss function is defined as:
[0063]
[0064] where s i and Represents l i The true confidence distribution and predicted confidence distribution of , s i a and represent the ath element in the true confidence distribution and the predicted confidence distribution, respectively. Represents the training set.
[0065] In order to further improve the accuracy of confidence prediction, this embodiment calculates the mean square error (MSE) loss function between the expected and true confidence of the predicted confidence distribution for each quadruple, and the calculation method is as follows:
[0066]
[0067] where s i Yes i The true confidence level, is the prediction confidence distribution This embodiment takes this expectation as l i prediction confidence.
[0068] Will and Combined, we get the loss function of the confidence prediction task
[0069]
[0070] where w 1 and w 2 is a trade-off factor used to adjust and proportion.
[0071] To solve the link prediction problem, it is necessary to accurately evaluate the ranking of possible tail entities under a given query (h, r, ?). This embodiment uses an edge-based ranking loss function to optimize the performance of the method on the link prediction task. First, a new two-layer fully connected neural network is used to calculate the ranking score of a triple. The difference between this fully connected neural network and the fully connected neural network used for confidence prediction is that the output of this fully connected neural network is a scalar rather than a vector, and the scalar is normalized using the Sigmoid activation function. The ranking score of the triple is calculated as follows:
[0072]
[0073] in is the predicted triple ranking score, FCN 2 is a function that converts the connection between embeddings into a scalar using a fully connected network. Sigmoid(·) is an activation function used to map a scalar to a probability in the range of 0 to 1. The edge-based ranking loss function used to optimize link prediction is to increase the ranking score difference between positive and negative samples so that the ranking score of the positive sample is higher than the ranking score of the negative sample. Given a positive sample, this embodiment generates multiple negative samples by replacing the head or tail entity of the positive sample with a randomly selected entity. In this embodiment, the number of negative samples corresponding to each positive sample is set to 50. The edge-based ranking loss function is defined as:
[0074]
[0075] in and are the ranking scores of positive samples and negative samples respectively, is the set of negative samples, [x] + =max[0,x] is the standard hinge loss and γ is the margin size.
[0076] When using confidence prediction loss and link prediction loss When training CDL-RL simultaneously, it is necessary to balance the ratio of the two tasks during training to avoid one task dominating the training. This embodiment uses uncertainty weights to dynamically adjust the ratio of each task during training. The loss function of the kth task is adjusted to:
[0077]
[0078] where λ k is the noise parameter of the kth task, The weight of the kth task can be adjusted dynamically so that tasks with higher uncertainty have lower confidence, logλ k is a regularization term that prevents the loss from increasing due to λ k The increase of is the loss function of the kth task. The loss function of CDL-RL is calculated as:
[0079]
[0080] where λ 1 and λ 2 are the noise parameters of confidence prediction and link prediction, respectively, and θ represents the parameter of CDL-RL. When both labeled data and pseudo-labeled data of PCDG are used (as detailed in steps S3 and S4), the loss function of confidence prediction is redefined as:
[0081]
[0082] in is a labeled dataset, is a pseudo-labeled dataset, s j and denote the pseudo-label confidence distribution and prediction confidence distribution of the j-th pseudo-label quadruple of the pseudo-label dataset, respectively, and w p is the weight of pseudo-labeled data, s j b and represent the bth element in the pseudo confidence distribution and the corresponding predicted confidence distribution, respectively.
[0083] Step S3: Use PCDG to generate high-quality pseudo-confidence labels for unlabeled data, thereby obtaining pseudo-labeled data. The structure of PCDG is the same as that of CDL-RL, and its parameters are represented as η. During the training process of PCDG, PCDG first generates pseudo-confidence distributions for unlabeled data. These pseudo-labeled data contain the gradient information of PCDG, so that PCDG can be optimized. These pseudo-labeled data are represented as Note though and They are all generated by PCDG, but their roles are different. It is generated during the PCDG training phase. This example then evaluates CDL-RL on labeled data and The performance after performing a gradient descent step on the CDL-RL is used to check whether CDL-RL is better trained. The smaller the loss of CDL-RL, the better its performance. Therefore, the training goal of PCDG is to minimize the loss of CDL-RL on the labeled data after one update. The training goal of PCDG can be expressed as:
[0084]
[0085] in When the parameter of CDL-RL is θ + At The loss on θ + Is CDL-RL in and The parameters after a gradient descent process is performed on , and its calculation method can be expressed as:
[0086]
[0087] Where α is the learning rate, When the parameter of CDL-RL is θ, and Finally, the loss function of PCDG can be explicitly expressed as:
[0088]
[0089] Step S4: Iteratively train CDL-RL and PCDG using meta-self-training until convergence. First, label the data The input is sent to CDL-RL, which uses CDL-RL to learn the embedding of the uncertainty knowledge graph. CDL-RL uses the loss function of step S2 to optimize itself by minimizing the loss on the confidence prediction and link prediction tasks. When the embedding of entities and relations gradually stabilizes, PCDG is trained to generate high-quality pseudo-confidence labels for unlabeled data. It uses the CDL-RL used in step S3 to train the labeled data and pseudo-labeled data. After training PCDG for a while, a screening strategy is applied to the pseudo-labeled data generated by PCDG, that is, in the labeled data, the original confidence of each triple should be in the leading position in the transformed confidence distribution. Therefore, in the pseudo-labeled data, if the highest descriptive degree of a pseudo-confidence label is greater than a fixed threshold, the corresponding pseudo-confidence distribution is regarded as high quality, and the pseudo-labeled data is selected for training CDL-RL, while other pseudo-labeled data are removed. These selected pseudo-labeled data are then The input is then fed into CDL-RL to enhance its training process. CDL-RL and PCDG then repeat the above process and optimize each other, i.e., PCDG provides pseudo-labeled data for CDL-RL, and CDL-RL provides meta-learning objectives for PCDG. This iterative training process is complete meta-self-training. This embodiment then continuously optimizes CDL-RL and PCDG until convergence.
[0090] Algorithm 1 describes the process of meta-self-training in detail. This embodiment defines two important time points, namely, the round T at which the PCDG training starts. PCDG At the same time as the first round T of training CDL-RL with unlabeled data semiCDL In this embodiment, CDL-RL (θ) and PCDG (η) are first initialized, and the labeled data and sampled unlabeled data As the input of the uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning. PCDG When the current round exceeds T, this embodiment only uses labeled data to optimize CDL-RL. This setting allows the embedding of the uncertainty knowledge graph to be stable during this period, which will help this embodiment to train PCDG. PCDG But not reaching T semiCDL When we start training PCDG, PCDG first generates and will and And then input them into the latest updated CDL-RL, and further update the parameters of CDL-RL to get θ +, then PCDG will use the tag data to update itself. During this period, this embodiment does not update the PCDG generated Directly pass it to CDL-RL. Because the quality of the generated samples is not stable enough in the early stage of PCDG training, the process of training CDL-RL in this stage is the same as before. semiCDL In this embodiment, pseudo-labeled data is used To train CDL-RL. This embodiment applies a pseudo-label selection strategy and uses the filtered pseudo-labeled data and CDL-RL is optimized together, while PCDG is optimized the same as before. Afterwards, CDL-RL and PCDG will continue training until the maximum number of rounds is reached. In this embodiment, during the training process, stochastic gradient descent is used to update all parameters in each batch.
[0091]
[0092]
[0093] Step S5: After completing the model training through steps S1-S4, the data to be completed is passed into the trained CDL-RL to obtain the prediction confidence and ranking score. The confidence prediction task and link prediction task can be completed according to the prediction confidence and ranking score. For the confidence prediction task, the given triple is passed into CDL-RL, and the prediction confidence output by CDL-RL is directly used as the result of the confidence prediction; for the link prediction task, for the given head entity and relationship, all entities in the uncertainty knowledge graph are used as candidate tail entities, and they form candidate triples with the given head entity and relationship. For each candidate triple, CDL-RL is used to calculate its ranking score. The candidate triple with a high ranking score is the possible correct completion item.
[0094] The above embodiments are only preferred implementation modes of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and equivalent substitutions without departing from the principles of the present invention. These technical solutions after improvements and equivalent substitutions to the claims of the present invention all fall within the protection scope of the present invention.
Claims
1. An uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning, characterized in that: The steps include: S1, convert all triple confidences in the uncertainty knowledge graph training data into confidence distributions, S2, using the Confidence Distribution Learning-based Relational Learner (CDL-RL) to learn the embedding of the uncertainty knowledge graph on both the labeled data and the pseudo-labeled data generated by the pseudo-labeled data generator. S3, using the Pseudo Labeled Data Generator (PCDG) to generate high-quality pseudo confidence distribution labels for unlabeled data, S4, use meta-self-training to iteratively train CDL-RL and PCDG until they converge. S5: The data to be completed is passed into the trained CDL-RL for reasoning to achieve confidence prediction and link prediction.
2. The uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning as claimed in claim 1, characterized in that: In step S1, the confidence distribution is defined as a discrete distribution, and the granularity of the confidence label set is set to The ordered confidence label set is For a given quadruple (h, r, t, s) in an uncertainty knowledge graph, the confidence distribution of (h, r, t) is defined as where s i ∈[0,1] is the confidence label The description of (h, r, t), the confidence distribution is Gaussian distribution Generated, where σ is the standard deviation and s is the mean, so s has the highest descriptiveness. Therefore, in this step, each piece of knowledge in the uncertainty knowledge graph is re-expressed as a four-tuple (h, r, t, s), where is the confidence distribution.
3. The uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning as claimed in claim 2, characterized in that: In step S2, CDL-RL is used to learn the embedding of entities and relations in the uncertain knowledge graph. The input of CDL-RL is the four-tuple l = (h, r, t, s) obtained in step S1. CDL-RL has three learning objectives: 1) minimize the difference between the predicted confidence distribution and the true confidence distribution; 2) minimize the difference between the predicted confidence and the true confidence; 3) accurately predict the tail entity given the head entity and the tail entity, where the first two objectives are used to complete the confidence prediction, and the last objective is used for link prediction, as follows: given a triple (h, r, t), the embeddings of h, r and t are concatenated and passed into a two-layer fully connected neural network, which outputs an n+1 dimensional vector, and then the Softmax function is applied as the activation function. The predicted confidence distribution of the triple (h, r, t) Calculated by the following formula: where || represents the connection between embeddings, FCN1 is a function that converts the connection between embeddings into an n+1-dimensional vector using a fully connected network, and Softmax(·) is an activation function used to map an n+1-dimensional vector to an n+1-dimensional probability distribution. The Kullback-Leibler divergence (KL divergence) is used to measure the similarity between the predicted distribution and the true confidence distribution, and is used as one of the optimization objectives of the uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning. Assume that l i is the i-th quadruple in the training data, and this part of the loss function is defined as: where s i and Represents l i The true confidence distribution and predicted confidence distribution of , s i a and represent the ath element in the true confidence distribution and the predicted confidence distribution, respectively. represents the training set, The mean square error (MSE) loss function between the expected and true confidence of the predicted confidence distribution is calculated for each quadruple to further improve the accuracy of the confidence prediction. The calculation method is as follows: where s i Yes i The true confidence level, is the prediction confidence distribution expectations, and take this expectation as l i The prediction confidence of Will and Combined, we get the loss function of the confidence prediction task Where w1 and w2 are trade-off factors used to adjust and proportion; The ranking score of a triple is calculated as follows: in is the predicted triple ranking score, FCN2 is a function that converts the connection between embeddings into a scalar using a fully connected network, Sigmoid(·) is an activation function used to map a scalar to a probability in the range of 0 to 1. Given a positive sample, multiple negative samples are generated by replacing the head or tail entity of the positive sample with a randomly selected entity. The loss function is defined as: in and are the ranking scores of positive samples and negative samples respectively, is the set of negative samples, [x] + =max[0,x] is the standard hinge loss, γ is the edge size; Use confidence prediction loss at the same time and link prediction loss During training, it is necessary to balance the ratio of the two tasks in the training process to avoid one task dominating the training. The uncertainty weight is used to dynamically adjust the ratio of each task in the training process. After using the uncertainty weight, the loss function of the kth task is adjusted to: where λ k is the noise parameter of the kth task, The weight of the kth task can be adjusted dynamically so that tasks with higher uncertainty have lower confidence, logλ k is a regularization term that prevents the loss from increasing due to λ k The increase of is the loss function of the kth task. The loss function of CDL-RL is calculated as: Where λ1 and λ2 are the noise parameters for confidence prediction and link prediction, respectively, and θ represents the parameters of CDL-RL. when When both labeled data and pseudo-labeled data of PCDG are used, the loss function of confidence prediction is redefined as: in is a labeled dataset, is a pseudo-labeled dataset, s j and denote the pseudo-label confidence distribution and prediction confidence distribution of the j-th pseudo-label quadruple of the pseudo-label dataset, respectively, and w p is the weight of pseudo-labeled data, s j b and represent the bth element in the pseudo confidence distribution and the corresponding predicted confidence distribution, respectively.
4. The uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning as claimed in claim 3, characterized in that: In step S3, PCDG is used to generate high-quality pseudo-confidence labels for unlabeled data, thereby obtaining pseudo-labeled data. The structure of PCDG is the same as that of CDL-RL, and its parameters are represented as η. During the training process, PCDG first generates pseudo-confidence distributions for unlabeled data. These pseudo-labeled data contain the gradient information of PCDG, so that PCDG can be optimized. These pseudo-labeled data are represented as We then evaluate CDL-RL on labeled data and The performance after performing a gradient descent step on the CDL-RL is used to check whether CDL-RL is better trained. The smaller the loss of CDL-RL, the better its performance. The training goal of PCDG is to minimize the loss of CDL-RL on the labeled data after one update. The training goal of PCDG is expressed as: in When the parameter of CDL-RL is θ + At The loss on θ + Is CDL-RL in and The parameters after a gradient descent process is performed on , and the calculation method is expressed as: Where α is the learning rate, When the parameter of CDL-RL is θ, and Finally, the loss function of PCDG is explicitly expressed as:
5. The uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning as claimed in claim 4, characterized in that: In step S4, meta-self-training is applied to iteratively train CDL-RL and PCDG. The specific process of meta-self-training is as follows: First, the labeled data Input into CDL-RL, use CDL-RL to learn the embedding of the uncertain knowledge graph, and CDL-RL uses the loss function of step S2 to optimize itself by minimizing the loss on the confidence prediction and link prediction tasks. When the embedding of entities and relations gradually stabilizes, PCDG is trained to generate high-quality pseudo-confidence labels for unlabeled data. It uses the CDL-RL used in step S3 to train the labeled data and pseudo-labeled data. After training PCDG for a period of time, the screening strategy is applied to the pseudo-labeled data generated by PCDG, and then these selected pseudo-labeled data are The data is input into CDL-RL to enhance its training process. After that, CDL-RL and PCDG continuously repeat the above process and optimize each other. That is, PCDG provides pseudo-labeled data for CDL-RL, and CDL-RL provides meta-learning objectives for PCDG. This iterative training process is the complete meta-self-training, and then CDL-RL and PCDG are continuously optimized until convergence.
6. The uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning as claimed in claim 5, characterized in that: In step S5, the triple data to be completed is passed into the CDL-RL trained through steps S1-S4 for reasoning to obtain the prediction confidence and ranking score, thereby realizing the confidence prediction and link prediction tasks in the uncertain knowledge graph completion task.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements an uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning as described in any one of claims 1 to 6 above.
8. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, an uncertainty knowledge graph reasoning method based on semi-supervised confidence distribution learning as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Knowledge representation reasoning method based on encoder and decoder architecture
CN113836312A
Knowledge graph inference model, system and inference method for Bayesian small sample learning
CN114861917A
Cited By
Unbalanced mark distribution learning method suitable for movie recommendation
CN120952112A
Imbalanced Label Distribution Learning Method for Movie Recommendation
CN120952112B