Non-sampling knowledge graph embedding training method based on dynamic weighting
By adopting dynamically weighted non-sampling training method and improved loss function in the knowledge graph embedding model, the training accuracy and stability problems caused by negative sampling are solved, and the training performance and stability of the model are improved.
Patent Information
- Application Number
- CN202510173868.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The existing knowledge graph embedding models rely on negative sampling during training, resulting in a decrease in training accuracy and stability, and a square-based loss function weakens the model performance.
A non-sampled knowledge graph embedding framework based on dynamic weighting is proposed. By using non-positive triplets composed of all entities and relationships in the knowledge graph as negative triplets for training, and improving the least squares non-sampled loss function, the Focal-Loss weight factor and self-adversarial negative sampling weight factor are used to dynamically weight the loss function of negative triplets.
The training performance of the model is improved, the impact of easily divided negative triplets on model training is reduced, and the model's attention to difficultly divided negative triplets is enhanced, thereby improving the training accuracy and stability of the model.
Smart Images

Figure CN120012894A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital equipment, and in particular relates to a non-sampling knowledge graph embedding training method based on dynamic weighting. Background Art
[0002] At present, knowledge graph (KG) has attracted wide attention in academia and industry as an important carrier for storing and processing structured data and as an advanced knowledge interconnection solution that provides a wide range of intelligent system power for many applications. Knowledge graph embedding (KGE) embeds entities and relationships in KG into a continuous vector space, which can retain the inherent structure of the knowledge graph while facilitating operations such as entity search and matching, thus making the large-scale application of knowledge graph a reality. Therefore, in recent years, knowledge graph embedding has dominated the vast field of structured knowledge interconnection, and its effectiveness has been reflected in equipment digitalization scenarios. For example, in equipment fault prediction, a fault prediction model is established based on historical maintenance records, sensor data, and expert knowledge to detect potential problems in advance and avoid unexpected downtime.
[0003] In recent years, researchers have proposed various KGE models, among which the classic models include DistMult, SimplE, ComplEx, TransE and RotatE. Figure 1 As shown in the figure, in order to better mine useful information in the data set to obtain better knowledge embedding, the common training method of KGE is to provide the model with positive and negative samples at the same time, that is, positive triplets and negative triplets, and calculate the loss function of positive and negative triplets based on the constructed objective function for optimization training, so as to obtain the embedding vector representation of each element of the knowledge graph. However, the data set generally only has positive samples (positive samples, the triplets in the data set can be regarded as positive samples of the model), and the negative samples are constructed based on the positive samples by a certain method. Therefore, the training of most models depends on negative sampling, which is the process of constructing and selecting negative samples relative to positive samples based on a certain strategy. However, while the training method based on negative sampling improves the training efficiency, it also has some challenges. First, sampling only part of the negative triplets for training will reduce the accuracy of model training. Second, the different negative triplets sampled in different training rounds will affect the stability of model training.
[0004] Inspired by the latest research on non-sampling recommendation systems, some researchers have tried to apply non-sampling methods to KGE training, and proposed a non-sampling training method that uses all non-positive triplets composed of entities and relationships in the knowledge graph as negative triplets for training, and solves the problems of training accuracy and stability based on the negative sampling training method by skipping the negative sampling process. However, this training method uses a square-based loss function to weaken the training performance of the original knowledge graph embedding model, and a large number of easy-to-separate negative triplets far away from the model decision boundary are treated equally with other negative triplets such as difficult negative triplets in training (this application refers to negative triplets that are difficult to distinguish as positive samples or negative samples as difficult negative triplets, and easy-to-separate negative triplets are a concept relative to difficult negative samples), making it difficult for the model to learn useful gradient information from the training samples and unable to perform good parameter updates, thereby affecting the convergence speed and training accuracy of the model training.
[0005] ComplEx explored the impact of the number of negative samples corresponding to each positive sample on model performance on the FB15K dataset. Figure 2 As shown in the figure, as the number of negative samples increases, the performance of the model improves. When the number of negative samples reaches a certain number, the performance of the model decreases. This experimental result shows that sampling more negative samples for training usually achieves better results. However, when the number of negative samples reaches a certain level, too many negative samples for training will reduce the effect of model training due to the large number of easy-to-classify negative samples.
[0006] Existing KGE models can be roughly divided into three categories:
[0007] Geometric transformation-based methods: This type of method uses geometric transformation to model the semantic relationship between entities and relations. Among them, the most classic TransE model regards the relationship as a translation transformation between the head entity vector and the tail entity vector in the same space. Models such as TransH, TransR, TransD, and RotatE are expanded on the basis of TransE and can handle more complex relationships.
[0008] Tensor decomposition-based methods: This type of model constructs a score function through tensor decomposition to measure the probability of a triple being true. Among them, the widely used DistMult models the score function as the inner product of the head and tail entity vectors and the relationship diagonal matrix. The ComplEx, HolE, ANALOGY, and SimplE models proposed later are improvements and extensions of the DistMult model to obtain better embedding performance.
[0009] Neural network-based methods: This type of model uses the characteristics of deep neural networks such as CNN, RNN, and GNN to model complex relationships between entities and relations. Among them, models such as ConvE and ConvKB use CNN to model the latent semantics between entities and relations; Gradner et al. use RNN to model relationship paths to capture longer relationship dependencies in KG; R-GCN, RGHAT, and CompGCN use GNN to capture structured information in KG.
[0010] Negative sampling methods in knowledge embedding models can be roughly divided into three categories: random negative sampling methods, adversarial negative sampling, and negative sampling with additional information (Additional Data Enhanced Sampling), which will be introduced below. The random negative sampling method is the most basic and widely used negative sampling method. TransE randomly extracts an entity from the entity set and replaces it with the positive triple to generate a negative sample. In order to obtain better training results, TransH proposed an improved strategy, that is, for 1-N relationship triples, a higher replacement probability is given to the head entity, and correspondingly, for N-1 relationship triples, a higher replacement probability is given to the tail entity. The subsequent TransR, TransD and other models have adopted this sampling strategy. In summary, the random negative sampling method often samples overly simple negative samples, that is, easy-to-divided negative samples, which cannot provide useful information for the optimization training of the model, causing the model to fall into a local optimal solution.
[0011] Adversarial Sampling is based on the principle of GAN, and uses the semantic information of samples in the embedding space to select negative samples in order to sample better hard-to-distinguish negative samples. However, it has disadvantages such as complex framework, unstable training results, and long training time. To this end, RotatE proposes a self-adversarial negative sampling strategy, which dynamically adjusts the sampling probability of negative samples during the model training process to sample hard-to-distinguish negative samples for training.
[0012] The negative sampling method that introduces additional information introduces external information in the negative sampling process to increase the sampling probability of difficult negative samples.
[0013] Regardless of the negative sampling method, the training method based on negative sampling only uses part of the negative sample information in the data set for training, which will inevitably waste the training information of the negative samples that have not been sampled. In addition, since the negative samples sampled in different rounds are different, it will also have a certain impact on the stability of the model training.
[0014] Knowledge graph embedding model based on non-sampling: In order to overcome the problems of high training complexity and prediction bias caused by negative sampling's over-reliance on negative sample selection, a regularization term is added to the loss function to prevent the model from generating high scores for all triplets, thereby eliminating the need for negative sampling. Although this method does not explicitly use negative sampling, the calculation of the regularization term implicitly uses all possible triplets in the knowledge graph, including positive and negative triplets. KG-NSF proposes that negative samples are not required for training, and KGE learning is performed directly based on the cross-correlation matrix of the embedding vectors of entities and relations in the knowledge graph. However, due to the need to calculate multiplications between large matrices, the space complexity of this method is high. The non-sampling knowledge graph embedding framework NS-KGE treats all non-positive triplets composed of entities and relations as negative triplets for training. This method solves the shortcomings of the negative sampling training method by skipping negative sampling. However, this method uses a square-based loss function, which weakens the training performance of the original KGE model, and treating all negative triplets equally in training will result in a decrease in training effect. Summary of the invention
[0015] To this end, the present application proposes a scalable non-sampling knowledge graph embedding framework based on dynamic weighting. First, a non-sampling training method is adopted, and all non-positive triples composed of entities and relationships in the knowledge graph are used as negative triples for training, which alleviates the problem of reduced training accuracy and stability caused by the training method based on negative sampling; secondly, the non-sampling loss function based on least squares is improved, and the prediction value based on the score function is understood as the probability value of the triple being predicted as a positive triple. On this basis, a new non-sampling loss function is simplified and defined to improve the training performance of the model. Finally, the loss function of the negative triple is dynamically weighted by using the Focal-Loss weight factor and the weight factor based on self-adversarial negative sampling to reduce the impact of easy-to-divided negative triples on model training. In order to evaluate the performance of the non-sampling KGE model training framework based on dynamic weighting proposed in this application, the training framework is applied to five KGE models DistMult, SimplE, ComplEx, TransE and RotatE for experiments, and the effectiveness of the training framework proposed in this application is verified.
[0016] To achieve the above purpose, the non-sampling knowledge graph embedding training method based on dynamic weighting disclosed in this application includes the following steps:
[0017] Construct a knowledge graph embedding model based on the translation distance model based on the historical maintenance records of equipment, sensor data and expert knowledge;
[0018] In the knowledge graph embedding model based on the translation distance model, the performance weakening problem of the model caused by the square-based loss function is alleviated by using the non-sampling loss function based on the probability prediction function;
[0019] The weight factor is used to dynamically weight the loss function of negative triples to reduce the adverse effects of easy-to-classify negative triples on model training;
[0020] Divide the training set and test set, train the knowledge graph embedding model, optimize the model parameters, and discover potential equipment problems.
[0021] Preferably, the knowledge graph embedding model based on the translation distance model includes at least one of the TransE and RotatE models.
[0022] Preferably, for the knowledge graph embedding model G based on the translation distance model, for the triple (h, r, t), where h, t∈E, E is an entity set, h and t are the head entity and the tail entity in the triple, r∈R, R is a relationship set, r is a relationship, and f r (h, t) represents the true label value of the triple (h, r, t), where for a positive triple (h, r, t), f r (h, t) = 1, indicating that the head entity h and the tail entity t can be connected by the relationship r, while for the negative triple (h, r, t), f r (h, t) = 0, indicating that the head entity h and the tail entity t cannot be connected by the relationship r; It represents the predicted function value based on the score function of the knowledge graph embedding model, which aims to determine whether the head entity h and the tail entity t can be connected through the relationship r; the training framework of the model aims to learn the vector representations h, r, t of entities h, t and relationship r in the knowledge graph embedding model by minimizing the difference between the true label values of positive and negative triples and the predicted values based on the score function.
[0023] Preferably, As the probability prediction function P(r|h,t) that the triple (h,r,t) is predicted to be a positive triple; for the triple (h,r,t), its probability prediction function P(r|h,t) is defined as:
[0024] P(r|h,t)=sigmoid(δ r -F r (h,t))
[0025] Among them, δ r >0, indicating the threshold hyperparameter related to the relation r, F r (h, t) is the score function of the triple (h, r, t); if the score function F of (h, r, t) rThe value of (h,t) is greater than the threshold δ r , the value of P will be close to 0, that is, the probability of (h, r, t) being predicted as a positive triple is close to 0, then it is predicted as a negative triple; conversely, the value of the predicted probability function is close to 1, then it is predicted as a positive triple;
[0026] The training framework uses a least squares-based loss function to define its optimization training loss function, namely
[0027]
[0028] C hrt Represents the weight of the corresponding triple (h, r, t), indicating the importance of the triple in model training;
[0029] Expand the above formula according to the positive and negative sample sets, and we get:
[0030]
[0031] Among them, C + , C - represents the weight of positive and negative triples, S + is a set of positive triples, S - is the set of negative triples;
[0032] For a positive triple, its true value f r (h, t) = 1, for negative triples, its true value f r (h,t)=0,
[0033] Substituting into the above formula, we get:
[0034]
[0035] Preferably, since the value range of the probability prediction function P(r|h,t) of the triple is [0,1], therefore, P(r|h,t)≥0, 1-P(r|h,t)≥0;
[0036] The loss function is simplified to:
[0037]
[0038] Preferably, in the loss function, a model-based weight factor is added to all negative triplets, and the influence of different negative triplets on model training is dynamically adjusted during the training process; the negative triplet weight factor includes a Focal-Loss weight factor and a self-adv weight factor.
[0039] Preferably, the negative triplet Focal-Loss weight factor for:
[0040]
[0041] Among them, β1,γ>0 are hyperparameters. β1 is to balance the weights between positive and negative samples. F represents the probability of a negative triplet being misclassified, that is, the probability that a negative triplet is misclassified as a positive triplet; for a negative triplet (h, r, t), P F (r|h,t) is defined as:
[0042] P F (r|h,t)=sigmoid(δ r -F r (h,t))
[0043] Among them, δ r >0 is the threshold hyperparameter, F r (h,t)>0 is the score function of the triple (h,r,t);
[0044] The role of the Focal-Loss weight factor is: in model training, for easy-to-classify negative triples, its P F The closer it is to 0, the greater its weight factor C - The closer it is to 0, and for difficult negative triples, its P F The greater the value is than 0, the - The value is also greater than 0; therefore, the model dynamically adjusts the weights of easy-to-separate and hard-to-separate negative triples based on the predicted probability. By scaling the weight value to a greater degree, the weight of the easy-to-separate negative triples is reduced, so that the model pays less attention to this part of the samples, thereby reducing the impact of the easy-to-separate negative triples on the model training; for the hard-to-separate negative triples, their weight is less than the scaling degree of the easy-to-separate negative triples, so that the model training focuses on the hard-to-separate negative triples, thereby improving the effect of model training;
[0045] Combined with the negative triplet Focal-Loss weight factor, the final loss function is as follows:
[0046]
[0047] Among them, β1,γ>0 are hyperparameters.
[0048] Preferably, the self-adv weighting factor for:
[0049]
[0050] Among them, β2,α>0 are hyperparameters, β2 is to balance the weights between positive and negative samples as a whole; the self-adv weight factor C - The role is: the negative triplet (hj ',r,t j ') to the sum of the score function values of all negative triples with relation r after function transformation;
[0051] Combined with the self-adv weight factor, the final loss function is as follows:
[0052]
[0053] Among them, β2,α>0 are hyperparameters.
[0054] Preferably, for TransE, the score function F of the triplet r is defined as:
[0055] F r =||h+rt||2
[0056] The probability prediction function P(r|h,t) based on TransE is defined as:
[0057] P(r|h,t)=sigmoid(δ r -||h+rt||2)
[0058] For RotatE, the score function F of the triplet r is defined as:
[0059]
[0060] The probability prediction function P(r|h,t) based on RotatE is defined as:
[0061]
[0062] Calculate the model's score function F r After obtaining the probability prediction function P(r|h,t), we substitute it into the loss function to obtain the final loss function of the training framework, and then perform optimization training to learn the optimal embedding h, r, t for h, r, t.
[0063] The beneficial technical effects of this application include:
[0064] Firstly, a new non-sampling knowledge graph embedding framework is proposed, which alleviates the performance weakening problem of the model caused by the square-based loss function by simplifying the definition of the non-sampling loss function based on the probability prediction function.
[0065] Secondly, two weight factors are used to dynamically weight the loss function of negative triplets to reduce the adverse effects of easy-to-classify negative triplets on model training.
[0066] Thirdly, the proposed framework is scalable and can be applied to almost all negative sampling based KGE models. Taking multiple KGE models as examples, it demonstrates how the training framework can be applied to existing KGE models. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 Schematic diagram of the traditional training method based on negative sampling;
[0068] Figure 2 Experimental results of the effect of negative sampling number on model training;
[0069] Figure 3 Flowchart of the present invention. DETAILED DESCRIPTION
[0070] The present invention is further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention belong to the protection scope of the present invention.
[0071] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0072] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0073] The technical solutions provided in the embodiments of the present application involve technologies such as machine learning and natural language processing of artificial intelligence, which are specifically introduced and explained through the following embodiments.
[0074] A formal description is given to the dynamically weighted and scalable non-sampling knowledge graph embedding framework of this application.
[0075] Table 1 Basic symbols in this application
[0076]
[0077] Table 1 shows the basic symbols used. Given a knowledge graph G, for a triple (h, r, t), where h, t∈E, r∈R, f r (h, t) represents the true label value of the triple (h, r, t), where for a positive triple (h, r, t), f r (h, t) = 1, indicating that the head entity h and the tail entity t can be connected by the relationship r, while for the negative triple (h, r, t), f r (h, t) = 0, indicating that the head entity h and the tail entity t cannot be connected by the relationship r. represents the predicted function value based on the score function of the knowledge graph embedding model, which aims to determine whether the head entity h and the tail entity t can be connected through the relationship r. A general KGE training framework aims to learn the vector representations h, r, t of entities h, t and relations r in KG by minimizing the difference between the true label values of positive and negative triples and the predicted values based on the score function. Generally, the non-sampling-based training framework uses a least squares-based loss function to define its loss function for optimization training, that is,
[0078]
[0079] Among them, C hrt Indicates the weight of the corresponding triple (h, r, t), indicating the importance of the triple in model training. In the traditional KGE model based on negative sampling, such as the TransE model, since it only samples some non-positive triples in the knowledge graph as negative triples for training, it can be considered that the C of the positive triples and the sampled negative triples hrt The values are uniformly set to 1, and the C of the negative triples that are not sampled hrt The value is set to 0; in the NS-KGE model framework, the weight value C of all positive triples hrt The weights of all negative triplets are uniformly set to 1, and the weights of all negative triplets are uniformly set to a positive real number less than 1. The optimal weight value obtained through experiments is 0.001.
[0080] refer to Figure 3 The non-sampling knowledge graph embedding training method based on dynamic weighting disclosed in this application comprises the following steps:
[0081] Construct a knowledge graph embedding model based on the translation distance model based on the historical maintenance records of equipment, sensor data and expert knowledge;
[0082] In the knowledge graph embedding model based on the translation distance model, the performance weakening problem of the model caused by the square-based loss function is alleviated by using the non-sampling loss function based on the probability prediction function;
[0083] The weight factor is used to dynamically weight the loss function of negative triples to reduce the adverse effects of easy-to-classify negative triples on model training;
[0084] Divide the training set and test set, train the knowledge graph embedding model, optimize the model parameters, and discover potential equipment problems.
[0085] In one embodiment, the knowledge graph is composed of heterogeneous information of sensor data, where temperature, pressure, and rotation speed are regarded as nodes, and interaction relationships (such as the connection between temperature and rotation speed) are regarded as edges. Therefore, a knowledge graph containing a large amount of information can be obtained. Figure 3 Tuples, including topological structures and semantic relations.
[0086] The existing non-sampling-based training framework converts the optimization training of each original KGE model into a square-based loss function, which weakens the training performance of the original KGE model. The meaning of is to understand it as the probability prediction function P(r|h,t) that the triple (h,r,t) is predicted to be a positive triple. For the triple (h,r,t), its probability prediction function P(r|h,t) is defined as:
[0087] P(r|h,t)=sigmoid(δ r -F r (h,t)) (2)
[0088] Among them, δ r >0, indicating the threshold hyperparameter related to the relation r, F r (h, t) is the score function of the triple (h, r, t). If the score function F of (h, r, t) r The value of (h,t) is greater than the threshold δ r , the value of P will be close to 0, that is, the probability of (h,r,t) being predicted as a positive triple is close to 0, and it is predicted as a negative triple; conversely, the predicted probability function value is close to 1, and it is predicted as a positive triple. Therefore, the value range of P(r|h,t) is [0,1].
[0089] Substituting the probability prediction function P(r|h,t) into formula (1), we get:
[0090]
[0091] Expand the formula according to the positive and negative sample sets, and we get:
[0092]
[0093] Among them, C + , C- Represents the weight of positive and negative triples. For positive triples, its true value f r (h, t) = 1, for negative triples, its true value f r (h, t) = 0, and substituting it into formula (4), we get:
[0094]
[0095] Since the probability prediction function P(r|h,t) of the triplet ranges from [0,1], P(r|h,t)≥0, 1-P(r|h,t)≥0. In order to simplify the calculation and improve the training efficiency, this application simplifies the loss function of the non-sampled knowledge embedding model to:
[0096]
[0097] It can be seen from formula (6) that compared with the loss function based on least squares, the loss function of the simplified non-sampling knowledge graph embedding framework is defined based on the probability prediction function defined by the original KGE model, which can maintain the characteristics of the original KGE model and thus improve the training performance.
[0098] For the non-sampling training method, since there are only positive samples in the dataset, all negative samples are constructed later based on the closed world hypothesis, that is, all triples composed of entities and relations in the dataset that are not in the dataset are negative samples. Since positive samples are the main basis for training, they should be treated equally in training. Based on this, this application sets the weight factor (weight score) C of all positive triples + Uniformly set to 1. As for the negative triplets constructed later, since there are a huge number of obviously erroneous easy-to-separate negative triplets, these easy-to-separate negative triplets cannot provide effective information for the gradient descent in the model optimization process, making the model easily fall into the local optimal solution and unable to learn the optimal embedding. Therefore, if these huge number of easy-to-separate negative triplets are treated the same as other negative triplets such as difficult-to-separate negative triplets during the training process, it is easy to affect the direction of gradient descent, and then affect the accuracy and convergence speed of training. To this end, the present application considers adding a model-based weight factor to all negative triplets in the loss function, and dynamically adjusting the impact of different negative triplets on model training during the training process. Next, two different negative triplet weight factors will be introduced.
[0099] (1) Focal-Loss weight factor
[0100] The core idea of Focal Loss is to scale the loss function of the model as a whole. The easy-to-classify samples are scaled more than the difficult-to-classify samples, thereby reducing the weight of the easy-to-classify samples in the loss function and highlighting the weight of the difficult-to-classify samples. This allows the model to reduce the impact of easy-to-classify negative samples during training and focus more on difficult-to-classify samples.
[0101] Define the negative triplet Focal-Loss weight factor for:
[0102]
[0103] Among them, β1,γ>0 are hyperparameters. β1 is to balance the weights between positive and negative samples. F represents the probability of a negative triplet being misclassified, that is, the probability that a negative triplet is misclassified as a positive triplet. For a negative triplet (h, r, t), P F (r|h,t) is defined as:
[0104] P F (r|h,t)=sigmoid(δ r -F r (h,t)) (8)
[0105] Among them, δ r >0 is the threshold hyperparameter, F r (h, t)>0 is the score function of the triple (h, r, t). It can be seen that P F This is consistent with the probability function P defined in formula (2).
[0106] The Focal-Loss weight factor can be understood as: the score function F of the negative triple (h, r, t) r The value of (h,t) is greater than the threshold δ r , predicted probability function P F The closer the value of is to 0, the closer the probability of the negative triplet being predicted as a positive triplet is to 0, that is, the probability of the negative triplet being predicted as a positive triplet is close to 0. F is close to 0; accordingly, P of the difficult negative triples F Greater than 0. By β1,γ>0,1≥P F ≥0, get C - ≥0, it can be seen that C - With P F Positive correlation, that is, P F The closer it is to 0, the - The closer it is to 0, the F The larger the C - The larger the value, the larger the P value. F The closer it is to 0, the greater its weight factor C -The closer it is to 0, and for difficult negative triples, its P F The greater the value is than 0, the - The value is also greater than 0. Therefore, the model dynamically adjusts the weights of easy and difficult negative triples based on the predicted probability. Compared with the difficult negative triples, the weight of the easy negative triples is reduced by scaling the weight value to a greater extent, so that the model pays less attention to this part of the samples, thereby reducing the impact of the easy negative triples on the model training. For the difficult negative triples, the weight is lower than the scaling degree of the easy negative triples, which is equivalent to indirectly increasing the weight value of the difficult negative triples, so that the model training focuses on the difficult negative triples, thereby improving the effect of model training. This is similar to students studying for exams. For the knowledge points that have been mastered, similar to the easy triples, the effect of paying attention to such knowledge points on improving grades will not be too obvious, so you can spend less energy. For the knowledge points that have not been mastered well, similar to the difficult negative triples, you need to spend more energy. Strengthening the mastery of these knowledge points is the key to improving grades.
[0107] Combining formulas (6), (7) and (8), we can obtain:
[0108]
[0109] Among them, β1,γ>0 are hyperparameters.
[0110] (2) Self-adv weight factor
[0111] Self-adversarial negative sampling believes that in the negative sampling process, for those difficult-to-distinguish negative triplets, since the model’s classification ability for these triplets is not enough, its sampling probability should be increased and the training of difficult negative samples should be strengthened to improve the model’s classification ability for difficult-to-distinguish negative triplets; while for easy-to-distinguish negative triplets, they should be sampled with the smallest possible sampling probability, ultimately improving the training effect of the model.
[0112] For easy-to-classify negative triples, since they cannot provide much meaningful information for training, their weight in the loss function should be quite small compared to hard-to-classify negative triples; while for hard-to-classify negative triples that can provide more meaningful training information, a larger weight should be assigned in the loss function. Based on this, this application defines the self-adv weight factor for:
[0113]
[0114] Among them, β2, α>0 are hyperparameters, and β2 is also used to balance the weights between positive and negative samples. The weight factor C is defined as - In fact, it is: the negative triplet (h j ',r,t j') to the sum of the score function values of all negative triples with relation r after function transformation.
[0115] Combining formulas (6) and (10), we can obtain:
[0116]
[0117] Among them, β2,α>0 are hyperparameters.
[0118] The similarities between the two weight factors are that the basic concept is the same, both of which reduce the weight of easy-to-classify negative triples and increase the weight of hard-to-classify negative triples. The difference is that the Focal-Loss weight factor can directly obtain a weight by calculating the score function value of the triple, while the weight factor based on self-adversarial negative sampling needs to calculate the function value of all negative triplets with relationship r, and obtains a relative weight value. This application will test the training effect of the two weight factors in the experimental part.
[0119] In one embodiment, five typical knowledge graph embedding models, including DistMult, SimplE, ComplEx, TransE and RotatE, are taken as examples to demonstrate the application of the newly proposed non-sampling training framework on different KGE models.
[0120] For TransE, the score function F of the triplet r is defined as:
[0121] F r =||h+rt||2 (12)
[0122] Therefore, the probability prediction function P(r|h,t) based on TransE can be defined as:
[0123] P(r|h,t)=sigmoid(δ r -||h+rt||2) (13)
[0124] For RotatE, the score function F of the triplet r is defined as:
[0125]
[0126] Therefore, the probability prediction function P(r|h,t) based on RotatE can be defined as:
[0127]
[0128] Calculate the model's score function F rAfter obtaining the probability prediction function P(r|h,t), we substitute it into formulas (9) and (11) to obtain the final loss function of the training framework, and then perform optimization training to learn the optimal embedding h, r, t for h, r, t.
[0129] This application will evaluate the training performance of the NSDW-KGE framework through experiments.
[0130] This application will conduct experiments on the CMAPSS benchmark dataset used in the NS-KGE framework to evaluate the training performance of the NSDW-KGE framework. Both datasets are commonly used benchmark datasets in knowledge graph embedding. Among them, UCIMachine Learning Repository1.CMAPSSData Set1 (Commercial Modular Aero-Propulsion System Simulation), which is a dataset for predicting the remaining useful life (RUL) of aircraft engines. It contains sensor data for multiple flight cycles, such as temperature, pressure, speed, etc. UCI CMAPSSDataSet2.Pump Sensor Data, this dataset comes from an industrial pump system and contains time series data from multiple sensors, which can be used for fault detection and diagnosis.
[0131] According to the common method of setting up data sets in knowledge graph embedding experiments, the triple set is divided into training set, validation set and test set, accounting for approximately 85%, 5% and 10% respectively. The specific experimental steps for link prediction are as follows:
[0132] First, the optimal parameters of the model are trained through training and validation sets, including the optimal embedding representations of all entities and relations.
[0133] Then, a link prediction task is performed on the test set to evaluate the model's ability to predict the missing head entity or tail entity in the triple, that is, the ability to predict t given (h, r) or the ability to predict h given (r, t). Specifically, for each test triple (h, r, t), the tail entity t is replaced with each entity x in the knowledge graph to obtain a set of candidate triples (h, r, x), and the score of the score function F(h, r, x) of this set of candidate triples is calculated. The replaced triples are sorted in ascending order according to the score, and the ranking of the test triple (h, r, t) among all candidate triples can be obtained, denoted as rank(t). Similarly, the ranking of all candidate triples (x, r, t) obtained by replacing the head entity h of the test triple (h, r, t) with each entity x in the knowledge graph can be obtained, denoted as rank(h).
[0134] In addition, the present application adopts the "Filter" setting. Specifically, before the test set is evaluated to obtain the ranking of each test triple, the candidate triples existing in the training set, the validation set, and the test set are removed from the set of candidate triples. This is because the candidate triple is actually correct, and it is reasonable that its ranking based on the objective function score is higher than the ranking of the original test triple. Before obtaining the ranking score of each test triple, the candidate triple is removed to ensure that the candidate triple does not participate in the ranking to eliminate the influence of interference factors.
[0135] In the link prediction task, this application uses the following evaluation metrics to evaluate the quality of the learned embedding.
[0136] (1) MR value, which is the average ranking of all test triples, is defined as follows:
[0137]
[0138] (2) MRR value, which is the average of the reciprocal rankings of all test triples, is defined as follows:
[0139]
[0140] (3) Hits@N, that is, the ratio of the number of test triplets ranked not greater than N to the total number of test triplets. This application used three indicators, namely Hits@1, Hits@3 and Hits@10, in this experiment.
[0141] The larger the MRR value and Hits@N value, or the smaller the MR value, the better the model performance; otherwise, the worse the model performance.
[0142] The comparison methods are divided into three categories:
[0143] (1) Original knowledge graph embedding model, that is, the knowledge graph embedding model that does not use the non-sampling framework. This application selects two classic models, TransE and RotatE, as comparison methods.
[0144] TransE: Translation embedding model that minimizes the distance between the head entity vector and the tail entity vector after translation of the relation vector.
[0145] RotatE: Rotational embedding model, in complex space, the distance between the head entity vector and the tail entity vector after being rotated by the relation vector is minimized.
[0146] (2) KGE model using NS-KGE framework, that is, the original knowledge graph embedding model using NS-KGE framework, including NS-TransE and NS-ComplEx models (RotatE has NS-KGE version).
[0147] (3) KGE model using NSDW-KGE framework, that is, the original knowledge graph embedding model using NSDW-KGE framework, including two models: NSDW-ComplEx and NSDW-RotatE. Since two dynamic weighting methods are applied, it can be further divided into Focal Loss version and Self-adv version of NSDW-KGE, which are respectively recorded as: NSDW-TransE-FL, NSDW-RotatE-FL and NSDW-TransE-SA.
[0148] The learning rate λ used by stochastic gradient descent SGD is {0.0005, 0.0001, 0.01, 0.1}, and the value range of hyperparameters β1 and β2 is {0.01, 0.1, 1, 10, 100}. The value range of hyperparameter γ is {0.1, 0.2, 0.5, 1, 2, 5}. The embedding dimension d ranges from {100, 300, 500, 1000}, α∈{0.1, 0.3, 0.5, 0.8, 8, 1.0}, and batch=128. The model is iteratively trained for 1000 rounds on each dataset. The optimal parameters are determined by Hits@10 of the validation set. The optimal parameter configurations of the models for the two datasets are shown in Table 2.
[0149] Table 2 Optimal hyperparameter settings
[0150]
[0151] The model can also solve the shortcomings of data sparsity and data imbalance to a certain extent. Because all entities and relationships in the data set are either connected to form the positive samples of the model or connected to construct the negative samples required for model training, the samples involved in training are enhanced to a certain extent, and the frequency imbalance of some entities and relationships in the data set is alleviated to a certain extent. The reason for saying "to a certain extent" here is that the useful training information provided by easy-to-separate negative samples is limited.
[0152] The reasons why the model works are: first, sample enhancement (the number of negative samples for training is greatly increased), second, the data imbalance problem is solved to a certain extent, and third, the loss function is dynamically weighted to reduce the impact of easy-to-classify negative triplets on model training. The experimental results are shown in Table 3 and Figure 3 shown.
[0153] Table 3 Link prediction results
[0154]
[0155] This application proposes a scalable non-sampled knowledge graph embedding framework NSDW-KGE based on dynamic weighting. The model adopts a non-sampled training method, that is, all non-positive triples composed of entities and relationships in the knowledge graph are used as negative triples for training, and a new non-sampled loss function is simplified and defined. The loss function of negative triples is dynamically weighted using two weighting methods, Focal-Loss weight factor and self-adv weight factor. Taking two KGE models TransE and RotatE as examples, the scalability of the training framework of this application in various different KGE models is demonstrated.
[0156] This application proposes a scalable non-sampling knowledge graph embedding method based on dynamic weighting. First, the least squares-based training loss function is transformed, and the redefined loss function can maintain the performance of the original knowledge graph embedding model; secondly, by dynamically weighting the loss function, the adverse effects of easily separable negative triples on model training are reduced, thereby improving the accuracy of model training; finally, this application is used for training on different types of original knowledge graph embedding models, demonstrating the scalability of this application. Experiments using a variety of different knowledge graph embedding models on benchmark datasets show that this application can achieve good training performance and can be applied to a variety of different knowledge graph embedding models.
[0157] As used herein, the word "preferred" is intended to be used as an example, instance, or illustration. Any aspect or design described herein as "preferred" is not necessarily to be construed as being more advantageous than other aspects or designs. On the contrary, the use of the word "preferred" is intended to present concepts in a specific way. The term "or" as used in this application is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" means any one of the naturally included permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.
[0158] Moreover, although the present disclosure has been shown and described with respect to one or implementations, those skilled in the art will think of equivalent variations and modifications based on the reading and understanding of this specification and the accompanying drawings. The present disclosure includes all such modifications and variations, and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above-mentioned components (such as elements, etc.), the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (such as it is functionally equivalent), even if the structure is not equivalent to the disclosed structure of the function in the exemplary implementation of the present disclosure shown herein. In addition, although the specific features of the present disclosure have been disclosed with respect to only one of several implementations, such features can be combined with one or other features of other implementations that may be desired and advantageous for a given or specific application. Moreover, insofar as the terms "including", "having", "containing" or their variations are used in specific embodiments or claims, such terms are intended to be included in a manner similar to the term "comprising".
[0159] The functional units in the embodiments of the present invention may be integrated into a processing module, or each unit may exist physically separately, or multiple or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc. The above-mentioned devices or systems may execute the storage method in the corresponding method embodiment.
[0160] To sum up, the above embodiment is an implementation mode of the present invention, but the implementation mode of the present invention is not limited by the embodiment. Any other changes, modifications, substitutions, combinations, and simplifications that deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A non-sampling knowledge graph embedding training method based on dynamic weighting, characterized in that: The following steps are involved: Construct a knowledge graph embedding model based on the translation distance model based on the historical maintenance records of equipment, sensor data and expert knowledge; In the knowledge graph embedding model based on the translation distance model, the performance weakening problem of the model caused by the square-based loss function is alleviated by using the non-sampling loss function based on the probability prediction function; The weight factor is used to dynamically weight the loss function of negative triples to reduce the adverse effects of easy-to-classify negative triples on model training; Divide the training set and test set, train the knowledge graph embedding model, optimize the model parameters, and discover potential equipment problems.
2. The non-sampling knowledge graph embedding training method based on dynamic weighting according to claim 1 is characterized in that: The knowledge graph embedding model based on the translation distance model includes at least one of the TransE and RotatE models.
3. The non-sampling knowledge graph embedding training method based on dynamic weighting according to claim 2 is characterized in that: For the knowledge graph embedding model G based on the translation distance model, for the triple (h, r, t), where h, t∈E, E is an entity set, h and t are the head entity and tail entity in the triple, r∈R, R is a relationship set, r is a relationship, f r (h, t) represents the true label value of the triple (h, r, t), where for a positive triple (h, r, t), f r (h, t) = 1, indicating that the head entity h and the tail entity t can be connected by the relationship r, while for the negative triple (h, r, t), f r (h, t) = 0, indicating that the head entity h and the tail entity t cannot be connected by the relationship r; It represents the predicted function value based on the score function of the knowledge graph embedding model, which aims to determine whether the head entity h and the tail entity t can be connected through the relationship r; the training framework of the model aims to learn the vector representations h, r, t of entities h, t and relationship r in the knowledge graph embedding model by minimizing the difference between the true label values of positive and negative triples and the predicted values based on the score function.
4. The non-sampling knowledge graph embedding training method based on dynamic weighting according to claim 3 is characterized in that: Will As the probability prediction function P(r|h,t) that the triple (h,r,t) is predicted to be a positive triple; for the triple (h,r,t), its probability prediction function P(r|h,t) is defined as: P(r|h,t)=sigmoid(δ r -F r (h,t)) Among them, δ r >0, indicating the threshold hyperparameter related to the relation r, F r (h, t) is the score function of the triple (h, r, t); if the score function F of (h, r, t) r The value of (h,t) is greater than the threshold δ r , the value of P will be close to 0, that is, the probability of (h, r, t) being predicted as a positive triple is close to 0, then it is predicted as a negative triple; conversely, the value of the predicted probability function is close to 1, then it is predicted as a positive triple; The training framework uses a least squares-based loss function to define its optimization training loss function, namely C hrt Represents the weight of the corresponding triple (h, r, t), indicating the importance of the triple in model training; Expand the above formula according to the positive and negative sample sets, and we get: Among them, C + , C - represents the weight of positive and negative triples, S + is a set of positive triples, S - is the set of negative triples; For a positive triple, its true value f r (h, t) = 1, for negative triples, its true value f r (h, t) = 0, substituting into the above formula, we get:
5. The non-sampling knowledge graph embedding training method based on dynamic weighting according to claim 4 is characterized in that: Since the probability prediction function P(r|h,t) of the triplet ranges from [0,1], therefore, P(r|h,t)≥0, 1-P(r|h,t)≥0; The loss function is simplified to:
6. The non-sampling knowledge graph embedding training method based on dynamic weighting according to claim 5 is characterized in that: In the loss function, a model-based weight factor is added to all negative triplets, and the influence of different negative triplets on model training is dynamically adjusted during the training process; The negative triplet weight factors include a Focal-Loss weight factor and a self-adv weight factor.
7. The non-sampling knowledge graph embedding training method based on dynamic weighting according to claim 6 is characterized in that: The negative triplet Focal-Loss weight factor for: Among them, β1,γ>0 are hyperparameters. β1 is to balance the weights between positive and negative samples. F represents the probability of a negative triplet being misclassified, that is, the probability that a negative triplet is misclassified as a positive triplet; for a negative triplet (h, r, t), P F (r|h,t) is defined as: P F (r|h,t)=sigmoid(δ r -F r (h,t)) Among them, δ r >0 is the threshold hyperparameter, F r (h,t)>0 is the score function of the triple (h,r,t); The role of the Focal-Loss weight factor is: in model training, for easy-to-classify negative triples, its P F The closer it is to 0, the greater its weight factor C - The closer it is to 0, and for difficult negative triples, its P F The greater the value is than 0, the - The value is also greater than 0; therefore, the model dynamically adjusts the weights of easy-to-separate and hard-to-separate negative triples based on the predicted probability. By scaling the weight value to a greater degree, the weight of the easy-to-separate negative triples is reduced, so that the model pays less attention to this part of the samples, thereby reducing the impact of the easy-to-separate negative triples on the model training; for the hard-to-separate negative triples, their weight is less than the scaling degree of the easy-to-separate negative triples, so that the model training focuses on the hard-to-separate negative triples, thereby improving the effect of model training; Combined with the negative triplet Focal-Loss weight factor, the final loss function is as follows: Among them, β1,γ>0 are hyperparameters.
8. The non-sampling knowledge graph embedding training method based on dynamic weighting according to claim 6 is characterized in that: self-adv weight factor C s - for: Among them, β2,α>0 are hyperparameters, β2 is to balance the weights between positive and negative samples as a whole; the self-adv weight factor C - The role is: the negative triplet (h j ',r,t j ') to the sum of the score function values of all negative triples with relation r after function transformation; Combined with the self-adv weight factor, the final loss function is as follows: Among them, β2,α>0 are hyperparameters.
9. The non-sampling knowledge graph embedding training method based on dynamic weighting according to claim 8 is characterized in that: For TransE, the score function F of the triplet r is defined as: Fr=||h+rt||2 The probability prediction function P(r|h,t) based on TransE is defined as: P(r|h,t)=sigmoid(δ r -||h+r-t||2) For RotatE, the score function F of the triplet r is defined as: The probability prediction function P(r|h,t) based on RotatE is defined as: Calculate the model's score function F r After obtaining the probability prediction function P(r|h,t), we substitute it into the loss function to obtain the final loss function of the training framework, and then perform optimization training to learn the optimal embedding h, r, t for h, r, t.
Citation Information
Patent Citations
Knowledge graph embedding model training method based on comparative learning
CN114741530A
Knowledge graph completion method based on multi-granularity hierarchy and dynamic embedding
CN116842199A
Small sample knowledge graph completion method based on graph structure information
CN119227790A