Mode and feature fusion knowledge graph embedding method and model

Through the knowledge graph embedding method of modal and feature fusion, combined with BERT pre-trained language model and multiple attention mechanisms, the problem of difficulty in comprehensively capturing the knowledge graph embedding features in the existing technology is solved, and the accuracy and efficiency of the link prediction task is improved.

CN120106197APending Publication Date: 2025-06-06GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510211374.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When the existing knowledge graph embedding algorithms deal with knowledge bases with strong text correlation, it is difficult to fully capture embedded feature information, and due to the sparsity and noise of the data, further improvement of accuracy is limited.

Method used

Using a knowledge graph embedding method of modality and feature fusion, the BERT pre-trained language model is serialized from text and triplets, combined with local attention modules, cross attention modules and similarity fitting strategies, effective embedded features are generated, and feature summary and output through multi-layer perceptron modules.

Benefits of technology

The various evaluation indicators of the link prediction task algorithm embedded in the knowledge graph are improved, the prediction accuracy of strongly associated entities of text is enhanced, the long-tail distribution problem is effectively solved, and the sampling cost is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106197A_ABST
    Figure CN120106197A_ABST
Patent Text Reader

Abstract

The invention discloses a mode and feature fusion knowledge graph embedding method and model, and the method comprises the following steps: 1, a feature extraction stage: importing a BERT pre-training language model, carrying out the serialized coding of a text and a triple, carrying out the code input and word embedding of a BERT model, and carrying out the code embedding of a target node and a node relation group; step 2, a feature fusion stage: carrying out effective feature fitting on information of different modalities and different dimensions by adopting interactive transmission of various modules, and generating effective embedded features; 3, a feature output stage: performing binary classification on feature output, and enabling model training to enter an iterative generalization stage; and 4, reasoning the link prediction task. Compared with the prior art, the method has the advantages that by adopting the modal fusion embedding and multi-path feature combination algorithm, under the condition that the scheme quality is kept or the precision is better improved, various evaluation indexes of the link prediction task algorithm embedded by the knowledge graph are improved, and more effective modeling is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge representation and reasoning technology, and specifically refers to a knowledge graph embedding method and model for modality and feature fusion. Background Art

[0002] The knowledge graph is essentially a semantic network composed of nodes and edges, where each node represents an "entity" in the real world, and each edge is a "relationship" between entities. Knowledge graphs usually use triples (head entity, relationship, tail entity) as a formal expression, such as (Guangdong, provincial capital, Guangzhou), where "Guangdong" and "Guangzhou" are the head entity and tail entity respectively, and "provincial capital" is the relationship between the two entities. This triple describes the fact that "the capital of Guangdong Province is Guangzhou". The incompleteness of the knowledge graph makes it impossible for downstream applications to perform as they should. Researchers in related fields began to think about how to complete the knowledge graph, and knowledge graph completion technology came into being. The methods of knowledge graph completion vary, and the main ideas are rule reasoning, path finding, and knowledge representation learning.

[0003] Knowledge Representation Learning is a key link in the field of knowledge graph technology. Among them, the graph-based knowledge reasoning method mainly adopts a neural network model to represent the entities in the knowledge base in the form of word vectors, and then uses various neural tensor network models to perform reasoning, mine hidden associations in the knowledge graph, and make the knowledge graph more complete. The quality of knowledge reasoning will directly affect the effect of knowledge graph completion, thereby having a significant impact on the integrity of knowledge base update and construction. In addition, due to the directional constraints between entities and relationships, knowledge reasoning derives complex relationship modeling, long-tail distribution problems, dynamic knowledge update problems, etc., and it is necessary to use various logical reasoning algorithms or neural network model algorithms to complete the various indicators in the link prediction task in knowledge reasoning with higher accuracy to achieve the purpose of effective prediction.

[0004] In the past few decades, researchers have designed a variety of knowledge graph embedding algorithms to optimize the accuracy of link prediction task indicators and effectively represent entities and relationships in learning graphs. These algorithms include distance and semantic matching methods, convolutional neural network methods, graph convolution methods, pre-trained language model methods, and so on. These graph embedding algorithms represent entity and relationship triplets as vectors and then look for efficient feature engineering solutions. Among them, the application of pre-trained language models has been widely studied in the field of graph embedding in recent years because it can effectively introduce information from other modalities.

[0005] The resources for building a knowledge base are usually obtained from text, so entity description text can naturally interact effectively with the knowledge space. Knowledge representation models based only on graph structured information cannot handle entities that are not in the knowledge graph, and the joint text embedding method can be complementary, so that the model can learn entities that appear in the text but are not in the knowledge graph. When using the encoding vector of a pre-trained language model to implement text embedding, first use the pre-trained language model to encode text information, such as masked language modeling. Then, the text encoding output is combined with the entity and relationship embedding of the knowledge graph to perform token mask embedding on the current tensor. Next, the confidence of its output and the label entity is calculated based on the feature fusion model: in the link prediction task, if the confidence is higher than the threshold, it means that the prediction is valid; otherwise, it means that the prediction is invalid. With the iteration of the model, the model is fitted to a specific knowledge graph, and the algorithm gradually converges. The fitting accuracy of the knowledge graph embedding algorithm based on the pre-trained language model is affected by the adjustable parameters of the internal feature model and the composition of the model. These parameters are usually set to fixed values ​​before starting training, and the composition of the model is mostly composed of neural network models. Without loss of generality, this fixed parameter value and diversified network layer combination is difficult to adapt to all application scenarios.

[0006] In order to improve the model's effective extraction of text information and entity information and accelerate the convergence of the algorithm, predecessors have conducted extensive research on the key aspects of modeling strategies and model iteration strategies for entity and relationship embedding features that integrate text information in knowledge graph embedding algorithms. Regarding the entity and relationship embedding strategy that uses a pre-trained language model to fuse text information, the paper (Yao L, Mao C, Luo Y. KG-BERT: BERT for knowledge graph completion [J]. arXiv preprint arXiv: 1909.03193, 2019.) adopts a method based on a contrastive learning framework to allow zero-cost negative sampling, reducing the expensive negative sampling cost and significantly reducing the time complexity of training and evaluation. The paper (Wang X, He Q, Liang J, et al. Language models as knowledge embeddings [J]. arXiv preprint arXiv: 2206.12617, 2022.) is based on a pre-trained language model and proposes a staged encoding embedding. The algorithm embeds the text description, the knowledge graph fact pair embedding and the combined encoding training are passed layer by layer, and the BERT model is used to construct a text encoder. The paper (Zhang Z, Liu X, Zhang Y, et al. Pretrain-KGE: learning knowledge representation from pretrained language gemodels[C] / / FindingsoftheAssociationforComputationalLinguistics:EMNLP2020.2020:259-266.) Different from the past KGC model, which predicts the tail entity when the head entity and relationship are given in the tail batch, a masked entity model is proposed, and its extensions for different application tasks are proposed. The literature (ChoiB, JangD, KoY.MEM-KGC: MaskedEntityModelforKnowledgeGraphCompletionWithPre-trainedLanguageModel[J].IEEEAccess,2021,9:132025-132032.) However, since the above-mentioned knowledge graph embedding modeling work based on text description is only optimized for certain single-point problems and focuses on local areas, the final result index improvement is limited. Summary of the invention

[0007] The technical problem to be solved by the present invention is to provide a knowledge graph embedding method and model that fuses modalities and features in response to the deficiencies raised in the above-mentioned background technology.

[0008] In order to solve the above technical problems, the technical solution provided by the present invention is: a knowledge graph embedding method and model for modality and feature fusion, which includes the following processes:

[0009] Step 1: Feature extraction stage: import the BERT pre-trained language model, serialize the code from the text and triples, encode the BERT model input and word embedding, and encode the target node and node relationship group;

[0010] Step 2: In the feature fusion stage, the interactive transmission of multiple modules is used to fit the effective features of information of different dimensions in different modalities to generate effective embedding features, that is, the encoding embedding is effectively fitted between the modules of the neural network;

[0011] Step 3: Feature output stage: feature output is classified into two categories, and model training enters the iterative generalization stage, calculates loss, back propagates, updates gradients, and finally converges to the global optimal model;

[0012] Step 4: Reasoning for link prediction task.

[0013] Furthermore, the feature fusion stage includes the following modules:

[0014] Module 1: Local attention module;

[0015] Module 2: Cross-Attention Module;

[0016] Module 3: Similarity fitting strategy;

[0017] The feature fusion stage is to first perform a local attention capture module on the embedding of the batch of triples with MASK fused with text and the embedding of the target entity. After the local attention capture is completed, it is output to the cross attention capture module to calculate the effective information embedding of the two modal combinations. At the same time, the local attention capture output is used to calculate the similarity of the features to adapt to the overall feature fusion, and the embedding of the batch of triples with MASK fused with the text earlier and the embedding of the target entity and the node degree features of the calculated triples are interacted with the idea of ​​residual fusion. Finally, the information of the three different feature attentions is summed up and then fitted and updated, and finally converges to the global optimal model.

[0018] Furthermore, the local attention module mainly includes linear transformation and activation and attention weight calculation;

[0019] The attention weight is calculated as shown in the formula:

[0020] a_weights = softmax(W 2 *Dropout(tanh(W 1 X*b 1 )+b2 ),

[0021] The weighted input and weighted sum are shown in the formula:

[0022] weights_sum = Σ i X T *a_weights.

[0023] Furthermore, the cross attention module is mainly used for weight calculation of attention scores;

[0024] The weight calculation of the attention score is calculated by the following formula:

[0025] query=W q *tensor text&triple +b q ,

[0026] key=W k *tensor entity +b k ,

[0027] value=W v *tensor entity +b v ,

[0028]

[0029] The weighted input and weighted sum are shown in the formula:

[0030] attention_output=attention_weights*value,

[0031] The output feature embedding is shown in the formula:

[0032] attention_output = W o *attention_output+b o .

[0033] Furthermore, the similarity fitting strategy quantifies the similarity between two vectors by calculating the angle between them, as shown in the following formula:

[0034]

[0035] Furthermore, the feature output stage splices the embedding of the previous text-fused triple batch with MASK and the embedding of the target entity in the row direction, and then continues to splice the value difference and Hadamard product of the two and the degree features of the nodes of the triple in the row direction, and finally passes them through a multi-layer perceptron module as a feature embedding operation, and the similarity fitting strategy feature fusion output and the cross attention module feature fusion output are both used as one output, and the three outputs are finally summarized into a batch of one iteration to perform global feature fusion;

[0036] Objective function of feature summary output:

[0037] M(q,f q-dra ,k,f k-dra )=σ(MLP(cat(q,k,qk,H(q,k),D(h,t))))+S(f q-dra ,f k-dra )+Cro_att(f q-dra ,f k-dra ).

[0038] After adopting the above structure, the present invention has the following advantages: by adopting the algorithm of modal fusion embedding and multi-way feature combination, the evaluation indicators of the link prediction task algorithm of knowledge graph embedding are improved while maintaining the quality of the solution or better improving the accuracy, thereby achieving more effective modeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a knowledge graph embedding method that integrates modality and features and a schematic diagram of the overall model CAFF-BERT framework;

[0040] Figure 2 It is a schematic diagram of the local attention module structure of a knowledge graph embedding method and model that integrates modality and features;

[0041] Figure 3 It is a schematic diagram of the cross-attention module structure of a knowledge graph embedding method and model that integrates modality and features;

[0042] Figure 4 It is a schematic diagram of the structure of a knowledge graph embedding method that integrates modality and features and a similarity fitting strategy module of the model;

[0043] Figure 5 It is a schematic diagram of the multi-layer perceptron module structure of a knowledge graph embedding method and model that integrates modality and features;

[0044] Figure 6It is a schematic diagram of embedding a batch of triples with MASK that is a knowledge graph embedding method that fuses modality and features and a text fusion of a model;

[0045] Figure 7 It is a schematic diagram of the embedding of the knowledge graph embedding method of modality and feature fusion and the target entity of the model;

[0046] Figure 8 This is a knowledge graph embedding method that integrates modality and features, and the model's model training loss results in WN18RR;

[0047] Fig. 9 It is a knowledge graph embedding method and model that integrates modality and features. The model is tested in WN18RR result diagram;

[0048] Fig.10 This is a knowledge graph embedding method that integrates modality and features, and the model's training loss results on FB15K-237;

[0049] Fig.11 It is a knowledge graph embedding method and model that integrates modality and features. The model is tested in FB15K-237.

[0050] Fig.12 It is a schematic diagram of a comparative learning framework between a knowledge graph embedding method that fuses modalities and features and a knowledge embedding of a model. DETAILED DESCRIPTION

[0051] In the methods described in the background technology section, in past studies, researchers have made a lot of improvements and innovations in improving the accuracy of link prediction indicators based on graph structure methods. However, these improvements usually only focus on the improvement of a single modality vector, such as the two-dimensional convolution model that uses convolution and fully connected layers to organize the propagation network and the graph convolution model that aggregates information through neural networks and adjacency matrices, ignoring the semantic information between entities and relationships in fact triples. This single-modality optimization may result in incomplete capture of embedded feature information when the entire algorithm is processing a knowledge base with strong text associations, or limit further improvements in accuracy due to the sparsity and noise of the data.

[0052] In order to overcome this limitation, it is necessary to consider changing multiple components in the graph embedding algorithm at the same time. For example, the combination of embedding vectors of multiple modes can be used to improve the limitations of a single mode. The feature fusion of multi-source embedding vectors optimizes the graph embedding algorithm. Secondly, the idea of ​​extracting matching similarity calculation is integrated into the embedding vectors of multiple sources. First, this method can obtain effective information at different levels, thereby improving the prediction accuracy of entity pairs with strong associations with texts and effectively solving the problem of long-tail distribution of knowledge bases. Secondly, the fusion of multi-way features to extract information forms a synergistic effect, and the effective information of different links promotes each other, improving the algorithm's ability to capture effective information globally. In addition, for the problem of reducing sampling costs, the strategy of using contrastive learning architecture can better handle the problem. At the same time, the transfer learning method of text embedding with a pre-trained language model better endows the algorithm with scalability and enhances the adaptability of the algorithm, making it more flexible for various downstream tasks.

[0053] The present invention is further described in detail below in conjunction with the accompanying drawings.

[0054] Combined with Figure 1-12 , the algorithm flow of the present invention includes the following steps:

[0055] Step 1: Import the BERT pre-trained language model, serialize the code from the text and triples, encode the BERT model input and word embedding, and encode the target node and node relationship group embedding;

[0056] Step 2: Encoding embedding is effectively fitted between the modules of the neural network;

[0057] Step 3: The feature output is classified into two categories, and the model training enters the iterative generalization stage, calculates the loss, back-propagates, updates the gradient, and finally converges to the global optimal model;

[0058] The loss function of back-propagation gradient update is:

[0059]

[0060] In the formula, K+ and K- are positive and negative keys, i.e., label keys, respectively, and K = K+∪K-. We use self-adversarial negative sampling for effective KE learning, and its calculation weight is

[0061] Step 4: Reasoning for link prediction task.

[0062] In the present invention, the name of the inventive method is CAFF-BERT, and its overall model architecture flow is as follows: Figure 1As shown. The process is divided into three stages. In the feature extraction stage, it starts with the serialization encoding of text and triples, the encoding input of the BERT model, the word embedding, and the encoding embedding of the target node and node relationship group. Then in the feature fusion stage, the interactive transmission of multiple modules is used to fit the effective features of information of different dimensions of different modalities to generate effective embedding features. Finally, binary classification is performed in the feature output stage, and the algorithm enters the iterative generalization stage. The key to modeling lies in the feature fusion stage. First, the embedding of the batch of triples with MASK fused with text and the embedding of the target entity are subjected to the local attention capture module. After the local attention capture is completed, it is output to the cross attention capture module to calculate the effective information embedding of the combination of the two modalities. At the same time, the local attention capture output is used to calculate the similarity of the features to adapt to the overall feature fusion. The embedding of the batch of triples with MASK fused with text earlier and the embedding of the target entity and the node degree features of the calculated triples are interacted with the idea of ​​residual fusion. Finally, the information of the three different feature attentions is summed up and then fitted and updated, and finally converges to the global optimal model. Next, each module and its innovations will be introduced in turn:

[0063] Local Attention Module

[0064] The local attention calculation is to calculate the attention weight through the tanh activation function after a linear transformation of the input tensor, which reduces the complex inner product steps in the traditional QK calculation and uses only two layers of linear transformation to directly calculate the weighted sum on the input, avoiding the complex Q, K, V branch structure, and making the operation process more efficient. By implicitly taking the input as Q, K, V, the complexity of the module is reduced, and it can also be well integrated into the scene that does not need to explicitly distinguish Q, K, V. It still retains the core idea of ​​the attention mechanism, and weights the input of different parts according to the correlation of the input tensor, captures the key local information of the embedded features, and refines the embedded information of different input modalities to enhance the expression ability of the model. Compared with the traditional attention mechanism, this module is not only simpler in implementation, but also has certain advantages for resource-constrained environments (such as embedded or mobile devices).

[0065] In the local attention module, the main tasks are linear transformation, activation, and attention weight calculation, weighted input, and weighted sum. Figure 2 As shown;

[0066] The attention weight calculation method is shown in formula (1):

[0067] a_weights = softmax(W 2 *Dropout(tanh(W 1 X*b 1 )+b 2 )(1)

[0068] The weighted input and weighted sum are shown in formula (2):

[0069] weights_sum = Σ i X T *a_weights(2)

[0070] Cross-Attention Module

[0071] In Transformer, the traditional attention mechanism is usually used, which is to calculate the similarity between different positions in the sequence and weight the input sequence through the self-attention mechanism. However, in the feature fusion task involving information extraction from different modalities (such as text descriptions and knowledge graph triples), the cross-attention proposed in this model algorithm provides a more flexible and powerful way to combine information from different sources. This cross-attention module, especially when it can process and fuse information streams from two different modalities, can effectively capture the relationship between them. This design enables the model to learn the interaction between text data and knowledge graph data, thereby enhancing the multimodal understanding ability of the model. By using the embedding representation of the batch of triples with mask fused with text as the query (Query), and the embedding representation of the target entity of the knowledge graph as the key (Key) and value (Value), the model can establish a dynamic relationship between the two different modal fusion embeddings and adaptively adjust the attention to each input modality. This mechanism provides effective contextual information for nodes and edges with complex relationships in the knowledge graph, thereby improving the understanding of complex information. In many past applications, the relationship between text description information and knowledge graphs is not easy to capture, especially in terms of the diversity of entities and relationships. The cross-attention module can provide a more effective fusion method by explicitly modeling their interactions, which is particularly suitable for downstream tasks involving knowledge graphs. When dealing with complex modal fusion, the cross-attention module can provide more accurate and efficient cross-modal information fusion compared to traditional self-attention or single modal models.

[0072] In the cross-attention module, the main task is to calculate the weight of the attention score, the weighted input and weighted sum, and the output feature embedding as follows Figure 3 As shown:

[0073] The weight calculation of the attention score is calculated by the following formula (3):

[0074] query=W q *tensor text&triple +b q

[0075] key=Wk *tensor entity +b k

[0076] value=W v *tensor entity +b v

[0077]

[0078] The weighted input and weighted sum are shown in formula (4):

[0079] attention_output=attention_weights*value (4);

[0080] The output feature embedding is shown in formula (5):

[0081] attention_output = W o *attention_output+b o (5);

[0082] By fusing two different modality embeddings (text description and knowledge graph entity representation) into a unified space and modeling them through a cross-attention mechanism, it can better capture the correlation between them. In this way, it provides a more efficient and flexible solution for the work in the feature fusion stage.

[0083] Similarity Fitting Strategy

[0084] In multimodal learning, especially in the process of feature calculation of text description and knowledge graph triples, the similarity measurement between different modalities is crucial. In order to effectively capture the connection between text description and knowledge graph entities and relations, the semantic matching model is combined with the idea of ​​similarity scoring function, that is, the rationality of facts is measured by semantic matching. Semantic matching usually uses the multiplication formula. This paper proposes to add L2 norm calculation on the basis of the multiplication formula, and finally adopts the similarity fitting strategy. Cosine similarity is a measure of the angle between two non-zero vector directions. It is often used in text analysis because it can effectively measure the extent to which two vectors point to the same direction regardless of their length. By calculating the angle between two vectors, their similarity is quantified. This measurement method is particularly suitable for processing high-dimensional sparse data such as text and knowledge graphs because it emphasizes directionality rather than amplitude, so that it can effectively capture the relationship between different modalities when the vector dimension is large. As shown in the following formula (6):

[0085]

[0086] It can more naturally measure the similarity between text and knowledge graph entities, especially at the semantic level. This similarity measurement can help the model better understand the correspondence between entities in the knowledge graph and text descriptions, thereby improving the accuracy of modal fusion, especially when it is necessary to understand and reason about the entity relationships in the knowledge graph to strengthen cross-modal alignment, such as Figure 4 .

[0087] In the formula, the denominator is calculated by multiplying the L2 norm and then performing the max operation with a small constant ∈ to ensure that it will not be divided by zero.

[0088] Feature summary output

[0089] First, the embedding of the previous text-fused triple batch with MASK and the embedding of the target entity are concatenated along the row direction, and then the value difference and Hadamard product of the two and the degree features of the nodes of the triple are also concatenated along the row direction. Finally, they are passed through a multi-layer perceptron module and used as a feature embedding as follows Figure 5 The embedding of the triple batch with MASK of the text fusion and the sequence representation and embedding representation of the target entity are as follows: Figure 6 and Figure 7 The feature fusion output of the similarity fitting strategy and the feature fusion output of the cross attention module are both taken as one output. The three outputs in total are finally aggregated into one iteration of a batch to perform global feature fusion.

[0090] Objective function of feature summary output:

[0091] M(q,f q-dra ,k,f k-dra )=σ(MLP(cat(q,k,qk,H(q,k),D(h,t))))+S(f q-dra ,f k-dra )+Cro_att(f q-dra ,f k-dra ).

[0092] This algorithm has been verified on open source and industry-authoritative benchmarks, namely FB15K-237 and WN18RR. These two datasets differ in the scale of the knowledge base. The benchmark datasets used for link prediction tasks are usually obtained by sampling a relatively complete knowledge graph. The data in the dataset is generally divided into a training set, a validation set, and a test set. The specific information is as follows:

[0093] Table 1 FB15K-237 and WN18RR datasets

[0094]

[0095] FB15k-237 is a sub-dataset of FB15k proposed by Toutanova et al. The authors extracted 237 most mentioned relations from the FB15k dataset and collected all knowledge related to these relations, while removing all knowledge in the training set that is equivalent to or has an inverse relationship with the test set data. In this way, the authors ensure that entities with connection relationships in the training set are not directly connected in both the validation set and the test set, which greatly solves the problem of data leakage. WN18RR is a sub-dataset of WN18 constructed by Dettmers et al. In order to solve the data leakage problem of WN18, the authors constructed a more challenging WN18RR dataset through a filtering method similar to FB15k-237. However, WN18RR has obvious defects, that is, there are 212 entities in the test set that do not appear in the training set, resulting in 6.7% of the test knowledge can never be correctly predicted.

[0096] All experiments were conducted on a Linux platform. The algorithm proposed in this invention was implemented using Python language and run on a workstation with Intel(R) Xeon(R) Platinum 8368 CPU @ 2.40 GHz and 80G A100-nVidia GPU.

[0097] The hyperparameters of the algorithm of the present invention are set as follows: the learning rate of the pre-trained model BERT is set to 1e-5, the model algorithm learning rate is 5e-4, the weight decay coefficient is 1e-7, the training batch size is dynamically adjusted according to the size of the data set, FB15K-237 is 64, WN18RR is 256, the hidden transfer layer dimension of the local attention module is 768, the neuron random dropout rate is 0.15, the hidden transfer layer dimension of the cross attention module is 768, and the neuron random dropout rate is 0.35. The method of the literature (Wang X, He Q, Liang J, et al. Language models as knowledge embeddings [J]. arXivpreprint arXiv: 2206.12617, 2022.) is implemented in the same environment as a benchmark algorithm for comparison. For the convenience of the following description, this algorithm is called C-LMKE. In addition, a dataset that is recognized for performance measurement in publicly published literature is selected as a benchmark, and a unimodal algorithm based on structure embedding and a previous bimodal algorithm based on text description embedding are used for comparison. The performance of the solutions is compared with that of C-LMKE and the algorithm of the present invention to measure the strength of the solution quality preservation.

[0098] Using the link prediction task as an experimental evaluation application, the link prediction model needs to fill all entities or relationships in the data set into the knowledge gaps that need to be predicted, and score each entity or relationship (depending on whether the gap is an entity or a relationship), and then sort according to the scoring results. Based on this, the mean reciprocal rank (MRR) and hit ratio Hits@ metrics are often used in link prediction experiments to evaluate the prediction effect of the model. The hit ratio includes Hits@1, Hits@3, and Hits@10. The mean rank (MR) is the average of the rankings of the true answers corresponding to each link prediction task in the prediction results, and its definition is as follows:

[0099]

[0100] Among them, there exists MR = [1, |E|], and the smaller the result value of MR, the better the prediction effect. Since the result is directly averaged, this indicator is extremely sensitive to outliers. Therefore, to solve this problem, researchers proposed the MRR metric, which is to take the average of the reciprocal ranking of the real answer corresponding to each link prediction task in the prediction result. Its specific definition is as follows:

[0101]

[0102] Among them, MRR exists MRR=[0,1], and the larger its value is, the better the prediction effect is.

[0103] H@K is the ratio of the ranking results of the correct answers in the prediction to be equal to or less than the threshold K, which is defined as follows:

[0104]

[0105] The higher the value of H@K, the better the prediction effect of the model. In the link prediction task, the commonly used values ​​of K are K = [1, 3, 10]. The lower the K value, the more obvious the difference between different models. Therefore, in the two metrics of Hits@1 and Hits@10, in general, more attention is paid to the measurement results of H@1. Due to different focuses, MRR and H@K cannot replace each other. MRR takes the average value and focuses on the overall prediction effect of the evaluation model; H@K focuses on targeted testing and evaluation based on the specific applications of the upper layer and the different requirements for the granularity of prediction accuracy. At the same time, the MR metric is highly sensitive to outliers and its stability is obviously less than that of MRR.

[0106] The specific experimental results are as follows:

[0107] WN18RR dataset

[0108] The published literature involves models based on structure embedding, including TransE, DistMult, ComplEx, ConvE, ComplexGCN, and models based on text description embedding, including Pretrain-KGE, KG-BRET, StAR, MEM-KGC, and C-LMKE as third-party benchmark methods.

[0109] Table 2 shows the comparison of the results of the proposed algorithm and other algorithms on WN18RR. Compared with the model based on structural embedding, the proposed algorithm has an average improvement of 60% in the MRR evaluation index, of which the average improvement ratios on Hits@1, Hits@3 and Hits@10 are 64%, 54% and 61% respectively. Compared with the C-LMKE model based on text description embedding, the proposed algorithm has an improvement of 2.49% in the MRR evaluation index, of which the improvement ratios on Hits@1, Hits@3 and Hits@10 are 2.54%, 1.05% and 0.50% respectively, that is, the proposed algorithm has no loss of solution performance.

[0110] Experimental results show that the algorithm of the present invention can achieve the highest efficiency while maintaining the same or even better solution quality. Figure 8-Figure 9 It is a visualization of the model’s training iteration loss and time changes under this dataset, as well as the verification results of various indicators.

[0111] Table 2 Experimental results on WN18RR

[0112]

[0113]

[0114] FB15K-237 dataset

[0115] The published literature involves models based on structure embedding, including TransE, DistMult, ComplEx, ConvE, ComplexGCN, and models based on text description embedding, including Pretrain-KGE, KG-BRET, StAR, MEM-KGC, and C-LMKE as third-party benchmark methods.

[0116] Table 3 shows the comparison of the results of the proposed algorithm and other algorithms on FB15K-237. Compared with the C-LMKE model based on text description embedding, the proposed algorithm improves the MRR evaluation index by 0.67%, and the improvement ratios of Hits@3 and Hits@10 are 1.86% and 1.47% respectively.

[0117] From the experimental results, it can be concluded that although the two methods are not much different in evaluation indicators, the algorithm of the present invention still has improvements, which illustrates the effectiveness of the algorithm of the present invention in mining text and knowledge graph information. Figure 10-11 It is a visualization of the model’s training iteration loss and time changes under this dataset, as well as the verification results of various indicators.

[0118] Table 3 Experimental results on FB15K-237

[0119]

[0120]

[0121] Analysis of structural characteristics and functional relationships

[0122] In the training architecture, a contrastive learning framework is adopted to better utilize the description-based knowledge graph embedding for link prediction. Given an entity-relationship pair and a target entity as a query q and a key k to be matched, compared to MEM-KGC, H is the hidden layer dimension, K is the number of entities input to the softmax, and the output embedding q of the mask entity is the encoded query. The weight of the linear layer We∈R K×H The i-th row in is the key corresponding to the i-th entity for each 1≤i≤|ε|. Inputting q into the linear layer is considered as matching the query q with the key. And the key is a directly learned embedding representation, rather than encoding text information like the query.

[0123] The learned entity representation (row of We) is replaced by a logits vector with the encoding of the target entity and text description. Comparative matching calculations are performed in small batches, which makes full use of information for effective link prediction while avoiding the additional cost of encoding negative samples. The query key in the batch is defined as the candidate key of the query, such as Fig.12 As shown in Figure 2. Encoding negative samples is crucial for knowledge graph embedding learning, but due to the cost of language models, most existing text description-based methods only assign a few negative samples to each positive sample. Tying the negative sampling size to the batch size solves the expensive negative sampling problem and allows text description-based knowledge graph embedding to learn features from more negative samples.

[0124] In summary, the algorithm of the present invention has the advantages of higher prediction accuracy and improved model extraction of effective information while maintaining the same or better solution quality. The specific advantages are as follows:

[0125] Aiming at the field of knowledge representation learning in knowledge graph construction technology, a knowledge graph embedding method and model based on modal feature fusion of pre-trained language model is proposed, including a newly proposed local attention mechanism module, cross attention module and similarity fitting strategy.

[0126] The main innovation of this method is to take into account the feature fusion strategy and similarity calculation strategy of embedding from the text description modality and the target entity graph modality at the same time, which is different from the embedding algorithm of the knowledge graph structure or the embedding algorithm based on the text description modality and the graph modality in the existing methods. Specifically, in the feature fusion stage, the attention mechanism is used as the core to perform progressive feature calculation and residual feature fusion, and the similarity fitting value is calculated. Then, the fused and matched feature embedding is used to update and fit the model under the contrastive learning training framework of early stopping supervision setting so that the model converges to the optimal. Through this piecewise comprehensive optimization, the global optimal solution or model parameters close to the optimal solution can be searched more effectively in the feature space. In addition, the transfer learning strategy of the pre-trained language model can adjust the algorithm at the encoding level of the pre-trained language model, and adjust different mask models for different application backgrounds or platforms, thereby improving the portability of the model. The introduction of weight decay in the model can effectively avoid falling into the local optimal solution. Secondly, the modules in different links form a synergistic effect, promote each other, and improve the ability of the model algorithm to extract effective information.

[0127] Experimental results show that the algorithm of the present invention has the advantages of higher prediction accuracy and improved model extraction of effective information while maintaining or improving the solution quality.

[0128] The present invention and its implementation methods are described above, and such description is not restrictive, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by it, and does not deviate from the purpose of the invention, and does not creatively design a structure and implementation method similar to the technical solution, they should all fall within the protection scope of the present invention.

Claims

1. A knowledge graph embedding method and model for modality and feature fusion, characterized by: It includes the following processes: Step 1: Feature extraction stage: import the BERT pre-trained language model, serialize the code from the text and triples, encode the BERT model input and word embedding, and encode the target node and node relationship group; Step 2: In the feature fusion stage, the interactive transmission of multiple modules is used to fit the effective features of information of different dimensions in different modalities to generate effective embedding features, that is, the encoding embedding is effectively fitted between the modules of the neural network; Step 3: Feature output stage: feature output is classified into two categories, and model training enters the iterative generalization stage, calculates loss, back propagates, updates gradients, and finally converges to the global optimal model; Step 4: Reasoning for link prediction task.

2. According to the knowledge graph embedding method and model of modality and feature fusion according to claim 1, it is characterized by: The feature fusion stage includes the following modules: Module 1: Local attention module; Module 2: Cross-Attention Module; Module 3: Similarity fitting strategy; The feature fusion stage is to first perform a local attention capture module on the embedding of the batch of triples with MASK fused with text and the embedding of the target entity. After the local attention capture is completed, it is output to the cross attention capture module to calculate the effective information embedding of the two modal combinations. At the same time, the local attention capture output is used to calculate the similarity of the features to adapt to the overall feature fusion. The idea of ​​residual fusion is used to interact with the embedding of the batch of triples with MASK fused with the earlier text, the embedding of the target entity and the node degree features of the calculated triples. Finally, the information of the three different feature attention levels is summed up and then fitted and updated, and finally converges to the global optimal model.

3. According to the knowledge graph embedding method and model of modality and feature fusion according to claim 2, it is characterized by: The local attention module is mainly linear transformation and activation and attention weight calculation; The attention weight is calculated as shown in the formula: a_weights=softmax(W2*Dropout(tanh(W1X*b1)+b2), The weighted input and weighted sum are shown in the formula: weights_sum=Σ i X T *a_weights。 4. According to the knowledge graph embedding method and model of modality and feature fusion according to claim 2, it is characterized by: The cross attention module is mainly used for weight calculation of attention scores; The weight calculation of the attention score is calculated by the following formula: query=W q *tensor text&triple +b q , key=W k *tensor entity +b k , value=W v *tensor entity +b v , The weighted input and weighted sum are shown in the formula: attention_output=attention_weights*value, The output feature embedding is shown in the formula: attention_output=W o *attention_output+b o 。 5. According to claim 2, a knowledge graph embedding method and model for modality and feature fusion, characterized in that: The similarity fitting strategy quantifies the similarity between two vectors by calculating the angle between them, as shown in the following formula:

6. A method and model for embedding a knowledge graph for fusing modality and features according to claims 1-5, characterized in that: In the feature output stage, the embedding of the previous text-fused triple batch with MASK and the embedding of the target entity are spliced ​​together along the row direction, and the value difference and Hadamard product of the two and the degree features of the nodes of the triple are also spliced ​​along the row direction. Finally, they are passed through a multi-layer perceptron module as a feature embedding operation, and the feature fusion output of the similarity fitting strategy and the feature fusion output of the cross attention module are both used as one output. The three outputs are finally summarized into a certain iteration of a batch to perform global feature fusion. Objective function of feature summary output: M(q,f q-dra ,k,f k-dra )=σ(MLP(cat(q,k,q-k,H(q,k),D(h,t))))+S(f q-dra ,f k-dra )+Cro_att(f q-dra ,f k-dra )。