A few-shot knowledge graph completion method based on feature fusion
Patent Information
- Application Number
- CN202611011019.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
AI Technical Summary
当直接邻居信息稀疏时,模型难以获得充分的关系推理证据
[0108]本发明的补全方法,首先,显式地引入头实体和尾实体之间的多跳路径,解决少样本补全中多跳证据利用不足的问题;并通过路径筛选机制,保留与目标关系高度相关的有效路径,解决路径噪声干扰较强的问题。其次,将有效路径转换为自然语言描述,提取其路径语义特征,为路径表示注入外部常识语义信息;通过多粒度路径编码,获取有效路径的结构特征,其中,通过细粒度编码捕捉相邻实体与关系之间的局部交互,通过粗粒度编码捕捉整条路径的整体语义模式,解决单一编码方式对路径信息刻画不充分的问题;最终,通过融合机制,学习目标关系的表示。
Smart Images

Figure CN122819409A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to knowledge graph reasoning technology, specifically to a few-shot knowledge graph completion method based on feature fusion. Background Technology
[0002] Knowledge graphs typically represent knowledge in the form of triples: "head entity, relation, and tail entity." For example, "Washington Capitals, Coach, Bruce Boudreau" can be represented as a triple. Knowledge graphs can organize entities and their relations into structured networks and are widely used in tasks such as intelligent question answering, recommendation systems, search engines, medical knowledge reasoning, and enterprise knowledge management.
[0003] However, in reality, due to the vast number of entities, the complexity of relationship types, and the dispersed data sources, many real-world facts are not recorded in the knowledge graph, resulting in incomplete knowledge graphs. Therefore, knowledge graph reasoning tasks are needed to complete the knowledge graph. The goal of a knowledge graph reasoning task is to reason and complete the missing facts based on the facts already existing in the graph. For example, given a head entity and its relation, predict possible tail entities; or given a tail entity and its relation, predict possible head entities.
[0004] Existing knowledge graph completion methods mainly include embedding representation-based methods, graph neural network-based methods, and meta-learning-based methods. Embedding representation-based methods first use a knowledge graph embedding model to map the head entity, target relation, and tail entity contained in each input triplet into a continuous vector space, obtaining embedding vectors for the entities and relations. Then, based on the embedding vectors of the head entity, target relation, and tail entity, a matching score for each triplet is calculated using vector distance or similarity. Finally, based on the matching scores of each input triplet, the validity of the match is determined, ultimately obtaining the knowledge graph completion information. Graph neural network-based methods further aggregate entity neighbor information to enhance entity representation. Meta-learning-based methods treat different relations as different tasks, quickly learning the representation of new relations using a small number of support samples.
[0005] However, in real-world knowledge graphs, relation distribution typically exhibits a pronounced long-tail characteristic. A few high-frequency relations possess a large number of triples, while a large number of low-frequency relations have only a few samples. For these long-tailed relations, traditional knowledge graph completion methods struggle to learn sufficiently stable relation representations. Therefore, few-sample knowledge graph completion has gradually become an important research direction in knowledge graph reasoning.
[0006] Existing few-shot knowledge graph completion methods typically learn target relation patterns using a small number of entity pairs in the support set, and then rank and predict candidate entities in the query set. While these methods alleviate the problem of insufficient samples, they still mainly rely on one-hop neighbors or the overall representation of support sample entity pairs, lacking explicit mining and modeling of multi-hop paths between entities. When direct neighbor information is sparse, the model struggles to obtain sufficient evidence for relational reasoning. Summary of the Invention
[0007] The technical problem to be solved by this invention is to propose a few-sample knowledge graph completion method based on feature fusion, which solves the technical problem in the prior art that "when the information of direct neighbors is sparse, the model has difficulty obtaining sufficient evidence for relational reasoning" by introducing multi-hop path information that supports entity pairs.
[0008] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0009] A few-shot knowledge graph completion method based on feature fusion, which is based on the head entities contained in each input triple. Target Relationship Tail-end entity The embedding vectors are used to calculate the matching score for each input triplet; based on the matching scores of each input triplet, the knowledge graph completion information is obtained; where each input triplet contains the head entity. Tail-end entity The embedding vectors are obtained using a knowledge graph embedding model; for each input triple, a feature fusion model is used to obtain the target relation for calculating its matching score. Embedding vector:
[0010] The target relationship is calculated using a feature fusion model. The embedding vector includes:
[0011] A1. Based on the topological information of the knowledge graph, for each input triple, separately mine the head entities it contains. Tail-end entity The length between them does not exceed The multi-hop path, the The preset maximum path length;
[0012] The multi-hop path is a de novo entity. Reaching the tail entity via a path of more than one hop The path is represented as:
[0013] ;
[0014] in, This is the sequence number of the multi-hop path. Indicates a multi-hop path The included first A relationship, Indicates a multi-hop path The included first One entity, For multi-hop paths Number of contained relationships; multi-hop paths The included entities All are unique entities; multi-hop paths All the included triples belong to the triple set of the knowledge graph;
[0015] A2. For each input triple, respectively, calculate the multi-hop paths based on that triple. Relationship with the target Based on the degree of relevance, multi-hop paths that meet the preset conditions are selected as valid paths for the triplet;
[0016] A3. Utilize a fine-grained feature extraction network to capture each effective path. Local interaction information to obtain each effective path Fine-grained features are extracted; a coarse-grained feature extraction network is used to capture each effective path. Global structural information to obtain each valid path The coarse-grained characteristics; the effective paths Convert them into natural language descriptions respectively. Using the coding model, each effective path is... Natural Language Description Obtained by Transformation Encode the valid paths. semantic features;
[0017] For each valid path Separately, integrate this effective path The fine-grained and coarse-grained features are used to obtain the effective path. The structural characteristics; utilizing a fusion network to integrate this effective path. Based on the structural and semantic features, obtain the effective path. Cross-modal fusion features;
[0018] A4. For each input triple, aggregate the valid paths for that triple separately. Cross-modal fusion features are used to obtain the target relation for calculating the matching score of the triple. The embedding vector.
[0019] Specifically, in step A1, the Bi-BFS algorithm is used. Based on the topological information of the knowledge graph, for each input triple, the head entity it contains is extracted. Tail-end entity The length between them does not exceed Multi-hop paths.
[0020] Furthermore, in step A2, for each multi-hop path of each input triple, the relationship between the multi-hop path and the target is calculated separately according to the following steps. Relevance:
[0021] Based on multi-hop path The relationships included Relationship with the target Similarity to obtain multi-hop paths Relationship with the target Similarity between them; based on multi-hop paths Relationship with the target The similarity between them is used to obtain multi-hop paths. Relationship with the target Correlation weights between them; based on multi-hop paths Each triplet contained The connection strength is used to calculate multi-hop paths. Confidence weights; fusion of multi-hop paths The relevance weights and confidence weights are used to obtain the characteristics of multi-hop paths. Relationship with the target The overall weight of the degree of relevance.
[0022] Specifically, in step A2, based on the following formula, the multi-hop path... The relationships included Relationship with the target Similarity to obtain multi-hop paths Relationship with the target similarity between :
[0023] ;
[0024] in, This represents a multi-hop path obtained using a knowledge graph embedding model. The included number Embedding vectors of each relation, This represents the target relation obtained using a knowledge graph embedding model. The embedding vector; For multi-hop paths The number of relationships contained; Represents the vector dot product. Represents the L2 norm;
[0025] In step A2, based on the following formula, the multi-hop path... Relationship with the target similarity between Obtain multi-hop paths Relationship with the target Correlation weights between :
[0026] ;
[0027] in, For head entity Tail-end entity A set of multi-hop paths between; Represents the natural exponential function;
[0028] In step A2, based on the following formula, the multi-hop path... Each triplet contained The connection strength is used to calculate multi-hop paths. Confidence weight :
[0029] ;
[0030] ;
[0031] in, Represents a triplet The connection strength, Represents a triplet Frequency of occurrence in knowledge graphs This represents the total number of entities contained in the knowledge graph;
[0032] In step A2, multi-hop paths are merged according to the following formula. Relevance weights and confidence weight Obtain the representation of multi-hop paths Relationship with the target The overall weight of the degree of relevance :
[0033] ;
[0034] in, To integrate the relevance weight and confidence weight, a balance coefficient is needed. This is a smoothing term.
[0035] Specifically, the fine-grained feature extraction network includes a bidirectional gated recurrent network and a graph attention network;
[0036] In step A3, a fine-grained feature extraction network is used to capture each effective path. Local interaction information to obtain each effective path Fine-grained features include:
[0037] AA1. Based on the effective paths of each triple obtained in step A2, construct the target relation. Path subgraph;
[0038] AA2. Using a bidirectional gated recurrent network, each valid path is encoded to obtain the encoded features of the entities and relationships contained in each valid path;
[0039] AA3. Using the encoding features of entities and relations contained in each effective path as input, and based on the correspondence between entities and relations, the fine-grained features of each effective path of each triple are calculated using a graph attention network based on the path subgraph.
[0040] In step AA2, a bidirectional gated cyclic network is used to analyze the effective paths. Encode to obtain a valid path The encoded features of the entities and relations included include:
[0041] AA21. Using the embedding vectors of entities and relations contained in the knowledge graph obtained by the knowledge graph embedding model, obtain effective paths. path sequence Embedded sequences;
[0042] AA22, using an effective path Using the embedded sequence as input, and employing a bidirectional gated recurrent network, the effective path is obtained according to the following formula. Encoded features of the entities and relations included:
[0043] ;
[0044] ;
[0045] ;
[0046] in, This represents the first path in the input path sequence. An embedding vector of elements; This represents the first path in the input path sequence. The positive hidden state of an element. This represents the first path in the input path sequence. The reverse hidden state of each element. Indicates the first position in the input path sequence The encoding features of each element, This indicates vector concatenation;
[0047] In step AA3, effective paths are calculated using a graph attention network based on the path subgraph. Fine-grained features include:
[0048] AA31. Based on the target relationship Obtain valid paths from the path subgraph. The set of neighbors of each entity contained therein;
[0049] AA32. Based on the correspondence between entities and relations, and according to the encoding characteristics of the entities and relations contained in each valid path, obtain the valid paths. The initial characteristics of each entity and its neighbors;
[0050] AA33, Based on Valid Paths The initial features of each entity and its neighbors are used to calculate the attention weights using the following formula:
[0051] ;
[0052] AA34. Based on attention weights, aggregate the data according to the following formula to obtain the effective path. Fine-grained features :
[0053] ;
[0054] in, For an effective path The The entity and its first Attention weights of each neighbor; For an effective path The The coding characteristics of an entity, and Valid paths The The first entity The neighbor and the first The encoded features of each neighbor; Let be the learnable linear transformation matrix of the graph attention network. For the learnable attention vectors of a graph attention network, This is the matrix transpose. For activation functions; Representation layer normalization.
[0055] Specifically, the encoding model is an encoder of a pre-trained large language model; the coarse-grained feature extraction network adopts a Transformer encoder.
[0056] In step A3, the coding model is used to analyze each valid path. Natural Language Description Obtained by Transformation Encode the model and extract the hidden states of the last layer of the encoding model that contain special classification labels, as the valid paths. semantic features;
[0057] In step A3, a coarse-grained feature extraction network is used to capture each effective path. Global structural information to obtain each valid path The coarse-grained characteristics include:
[0058] AB1, Valid Path path sequence By filling the mark Expand to the preset uniform length;
[0059] AB2, Valid Path The expanded path sequence is input into the Transformer encoder to extract... The output of the marker is used as the valid path. The coarse-grained characteristics;
[0060] In step A3, for each valid path Specifically, the effective paths are merged according to the following formula. Fine-grained features and coarse-grained characteristics Obtain the valid path Structural features :
[0061] ;
[0062] in, To integrate fine-grained features and coarse-grained characteristics The balance coefficient.
[0063] Specifically, the fusion network is a gated network; in step A3, the effective path is fused using the fusion network. Based on the structural and semantic features, obtain the effective path. The cross-modal fusion features include:
[0064] AC1. According to the following formula, the structural features are... and semantic features Mapped to the same dimension:
[0065] ;
[0066] ;
[0067] in, and These are the weight matrices for the learnable linear transformation. and These are the bias matrices for the learnable linear transformations;
[0068] AC2, mapping structural features to the same dimension and semantic features The gating vector is obtained by calculating using the following formula. :
[0069] ;
[0070] in, This represents the Sigmoid activation function. This represents the gating weight matrix of the gating network. Represents the bias vector of the gated network;
[0071] AC3, based on the gate vector According to the following formula, the structural features are integrated. and semantic features Obtain a valid path Cross-modal fusion features :
[0072] ;
[0073] in, This represents element-wise multiplication.
[0074] Specifically, in step A4, for each input triple, an attention mechanism is used to aggregate the effective paths of that triple according to the following formula. Cross-modal fusion features Obtain the target relation used to calculate the matching score of the triple. Embedded vector :
[0075] ;
[0076] ;
[0077] ;
[0078] in, Here is the transformation matrix for the attention mechanism. The set of valid paths corresponding to the triples. Transpose of the representation matrix.
[0079] Specifically, the training of the feature fusion model includes the following steps:
[0080] B1. Based on the target relationship of the input, construct triples for this round of training according to the known support samples;
[0081] B2. For each input triple, mine the multi-hop path of the triple using the method in step A1.
[0082] B3. For each input triplet, filter and obtain the valid path of the triplet according to the method in step A2.
[0083] B4. For each input triplet, obtain the cross-modal fusion features of each effective path of the triplet according to the method in step A3.
[0084] B5. For each input triplet, obtain the embedding vector of the target relation of the triplet according to the method in step A4.
[0085] B6. Calculate the total loss using the following formula. :
[0086] ;
[0087] in, The loss is for sorting triples. For path relationship consistency loss, and As weight;
[0088] B7. Determine whether the training is complete. If yes, end the training; otherwise, return to step B1.
[0089] The triplet sorting loss Calculate using the following steps:
[0090] ;
[0091] in, For the first A triplet, To replace triples The negative sample triples consisting of the tail entities, For triples and negative sample triples; For triples The score based on the embedding vector of the target relation obtained in step B5, negative sample triples The score based on the embedding vector of the target relation obtained in step B5, These are marginal parameters;
[0092] The path relationship consistency loss Calculate using the following steps:
[0093] First, for each valid path of each triple, cross-modal fusion features are performed through a linear layer. Mapping this to the embedding space of the knowledge graph embedding model yields its path relationship representation. :
[0094] ;
[0095] in, The weight matrix is a learnable linear transformation. The bias matrix is the learnable linear transformation. Represents the first of the corresponding triples One valid path;
[0096] Then, calculate the path relationship consistency loss of the effective path using the following formula:
[0097] ;
[0098] in, This refers to the embedding vector of the target relation obtained using a knowledge graph embedding model. Represents the L2 norm;
[0099] Then, the average of the path relationship consistency loss for each effective path in each triplet is used as the path relationship consistency loss. .
[0100] Furthermore, the encoding model is an encoder of a pre-trained large language model;
[0101] In step B4, each valid path is... Convert to natural language description Then, using coding models to describe natural language Before encoding, it also includes:
[0102] Determine if this is the first round of training; if so, perform the fine-tuning step for the encoding model; otherwise, skip the fine-tuning step for the encoding model.
[0103] The fine-tuning steps of the encoding model include:
[0104] Based on each effective path Described in its natural language As input, the labels of the target relations contained in the triples to which they belong are used as the output targets to construct training samples; using the constructed training samples, the contrastive loss is calculated according to the following formula through a contrastive learning algorithm to fine-tune the encoding model;
[0105] ;
[0106] in, Indicates the number of valid paths. Indicates the first The semantic features obtained by encoding the natural language description of each effective path through the encoding model. The semantic features obtained by encoding the labels representing the target relationship through an encoding model; Indicates the temperature coefficient. This represents the similarity function.
[0107] The beneficial effects of this invention are:
[0108] The completion method of this invention firstly introduces a multi-hop path between the head and tail entities to address the problem of insufficient utilization of multi-hop evidence in few-sample completion; and retains effective paths highly relevant to the target relationship through a path filtering mechanism to address the problem of strong path noise interference. Secondly, the effective paths are converted into natural language descriptions, and their path semantic features are extracted to inject external common-sense semantic information into the path representation; the structural features of the effective paths are obtained through multi-granularity path encoding, wherein fine-granular encoding captures the local interactions between adjacent entities and relationships, and coarse-granular encoding captures the overall semantic pattern of the entire path, addressing the problem that a single encoding method is insufficient to characterize path information; finally, the representation of the target relationship is learned through a fusion mechanism.
[0109] In summary, the method of this invention effectively solves the problems of weak path evidence, insufficient semantic support, and inadequate feature fusion in existing few-sample knowledge graph completion through path mining, path filtering, semantic enhancement, multi-granularity coding, and cross-modal fusion, and can improve the prediction effect of missing entities in long-tail and complex relationship scenarios. Attached Figure Description
[0110] Figure 1 This is a schematic diagram of the training process of the feature fusion model in the embodiment. Detailed Implementation
[0111] This invention aims to provide a few-shot knowledge graph completion method based on feature fusion. Its reasoning process is the same as existing embedding representation-based methods. First, based on the head entities contained in each input triple... Target Relationship Tail-end entity The embedding vectors are used to calculate the matching score for each input triple; then, based on the matching scores of each input triple, the knowledge graph completion information is obtained. This includes the head entities contained in each input triple. Tail-end entity The embedding vectors are obtained using knowledge graph embedding models; however, unlike existing embedding representation-based methods, in this invention, for each input triple, a feature fusion model is used to obtain the target relation used to calculate its matching score. The embedding vector.
[0112] Furthermore, by employing a feature fusion model, the target relationship can be calculated. The embedding vector includes:
[0113] A1. Based on the topological information of the knowledge graph, for each input triple, separately mine the head entities it contains. Tail-end entity The length between them does not exceed The multi-hop path, the The preset maximum path length;
[0114] The multi-hop path is a de novo entity. Reaching the tail entity via a path of more than one hop The path is represented as:
[0115] ;
[0116] in, This is the sequence number of the multi-hop path. Indicates a multi-hop path The included first A relationship, Indicates a multi-hop path The included first One entity, For multi-hop paths Number of contained relationships; multi-hop paths The entities included All are unique entities; multi-hop paths All the included triples belong to the triple set of the knowledge graph;
[0117] A2. For each input triple, respectively, calculate the multi-hop paths based on that triple. Relationship with the target Based on the degree of relevance, multi-hop paths that meet the preset conditions are selected as valid paths for the triplet;
[0118] A3. All valid paths Convert them into natural language descriptions respectively. ; Utilize a fine-grained feature extraction network to capture each effective path Local interaction information to obtain each effective path Fine-grained features are extracted; a coarse-grained feature extraction network is used to capture each effective path. Global structural information to obtain each valid path The coarse-grained features; using the coding model, each effective path Natural Language Description Obtained by Transformation Encode the valid paths. semantic features;
[0119] For each valid path Separately, integrate this effective path The fine-grained and coarse-grained features are used to obtain the effective path. The structural characteristics; utilizing a fusion network to integrate this effective path. Based on the structural and semantic features, obtain the effective path. Cross-modal fusion features;
[0120] A4. For each input triple, aggregate the valid paths for that triple separately. Cross-modal fusion features are used to obtain the target relation for calculating the matching score of the triple. The embedding vector.
[0121] Step A1 mines multi-hop paths by explicitly introducing them between head and tail entities, addressing the issue of insufficient utilization of multi-hop evidence in few-sample completion. Step A2 filters multi-hop paths, retaining effective paths highly relevant to the target relationship to address strong path noise interference. In Step A3, effective paths are converted into natural language descriptions. An encoding model extracts their semantic features, injecting external common-sense semantic information into the path representation. A fine-grained feature extraction network captures local interaction information between adjacent entities and relationships. A coarse-grained feature extraction network captures the global structural pattern of the entire path. Then, fine-grained and coarse-grained features are fused to obtain structural features. A fusion network then fuses the semantic and structural features of the effective paths to obtain cross-modal fusion features. Finally, by aggregating the cross-modal fusion features, a relationship embedding that fully integrates multi-hop path information is obtained.
[0122] Therefore, the method of this invention effectively solves the problems of weak path evidence, insufficient semantic support, and inadequate feature fusion in existing few-sample knowledge graph completion, and can improve the prediction effect of missing entities in long-tail and complex relationship scenarios. It should be noted that the core of this invention lies in the introduction of multi-hop paths and their cross-modal fusion features. Therefore, the mining of the above-mentioned multi-hop paths, as well as the encoding model, fine-grained feature extraction network, coarse-grained feature extraction network, fusion network, and scoring function, can all adopt any existing algorithm.
[0123] To ensure efficient searching while recalling as many multi-hop paths as possible, step A1 preferably employs the Bi-BFS algorithm. Based on the topological information of the knowledge graph, for each input triple, it separately mines the head entity it contains. Tail-end entity The length between them does not exceed Multi-hop paths. The Bi-Breadth-First Search (BFS) algorithm searches outwards simultaneously from both the head and tail entities. When the two ends meet at an intermediate entity, a complete path is formed.
[0124] Multi-hop path Relationship with the target The correlation can be calculated using any existing method. Preferably, in step A2, for each multi-hop path of each input triple, the relationship between the multi-hop path and the target is calculated separately according to the following steps. Relevance:
[0125] Based on multi-hop path The relationships included Relationship with the target Similarity to obtain multi-hop paths Relationship with the target Similarity between them; based on multi-hop paths Relationship with the target The similarity between them is used to obtain multi-hop paths. Relationship with the target Correlation weights between them. Based on multi-hop paths. Each triplet contained The connection strength is used to calculate multi-hop paths. Confidence weights; fusion of multi-hop paths The relevance weights and confidence weights are used to obtain the characteristics of multi-hop paths. Relationship with the target The overall weight of the degree of relevance.
[0126] The similarity weights mentioned above measure whether the relation sequence in a multi-hop path can reflect the semantic pattern of the target relation. The confidence weights mentioned above measure the reliability of the connections between each triple in a multi-hop path.
[0127] Pre-trained large language models accumulate rich semantic and common-sense knowledge through large-scale text pre-training. Therefore, using the encoder of a pre-trained large language model to construct an encoding model can enhance the semantic expressive power of path representation. Large language models such as QianWen3-8B, ChatGLM, LLaMA, and Baichuan, which possess text encoding capabilities, can be employed. The hidden states of the last layer of the encoding model, containing special classification labels, can be extracted as effective paths. The semantic features. In mainstream large language models, the last layer is a special classification label, typically the [CLS] label. The encoding model gathers the global semantic information of the input into the hidden state of the [CLS] label.
[0128] Fine-grained feature extraction networks are networks that focus on capturing local, individual, and subtle differences in features. They focus on the precise semantics of the "single element" itself and aim to perform refined feature extraction and representation learning of input information from multiple dimensions and levels. They are not limited to the combination of Bi-GRU and GAT used in the following embodiments. They can also adaptively select different structures such as Long Short-Term Memory Network (LSTM), Convolutional Neural Network (CNN), Graph Convolutional Network (GCN), or various improved graph attention mechanisms, depending on task requirements and data characteristics.
[0129] Coarse-grained feature extraction networks are feature networks focused on capturing overall, global, and general patterns. They ignore individual differences and focus on the statistical regularities of "groups" or "overall structures," aiming to capture global dependencies and long-range associations. In terms of structural selection, they are not limited to the Transformer used in subsequent embodiments, but can also adopt any network structure with global context modeling and long sequence dependency capture capabilities to achieve efficient encoding of global semantics and overall distribution information.
[0130] To dynamically adjust the contribution ratio of semantic features and structural features, a gated network is preferably used in the fusion network. For paths with more complete semantic information, the gated network can increase the contribution of path semantic features; for paths with more reliable graph structure connections, the gated network can increase the contribution of path structural features.
[0131] Similarly, the scoring function can be a distance scoring function, a semantic matching scoring function, or a neural network scoring function, as long as it can rank candidate entities according to the target relation representation.
[0132] As mentioned above, the neural networks included in the feature fusion model are all existing algorithms. Therefore, the training of the feature fusion model can be performed using a matching loss function based on the specific architecture of the feature fusion model. Preferably, the training of the feature fusion model includes the following steps:
[0133] B1. Based on the target relationship of the input, construct triples for this round of training according to the known support samples;
[0134] B2. For each input triple, mine the multi-hop path of the triple using the method in step A1.
[0135] B3. For each input triplet, filter and obtain the valid path of the triplet according to the method in step A2.
[0136] B4. For each input triplet, obtain the cross-modal fusion features of each effective path of the triplet according to the method in step A3.
[0137] B5. For each input triplet, obtain the embedding vector of the target relation of the triplet according to the method in step A4.
[0138] B6. Calculate the total loss using the following formula. :
[0139] ;
[0140] in, The loss is for sorting triples. For path relationship consistency loss, and As weight;
[0141] B7. Determine whether the training is complete. If yes, end the training; otherwise, return to step B1.
[0142] The total loss for the above training, in addition to the conventional triplet ranking loss, also introduces a path relation consistency loss. The path relation consistency loss maps the cross-modal fused features of the effective paths to the relation embedding space and calculates the similarity or distance between them and the initial embedding of the target relation. The consistency loss constrains the cross-modal fused features to maintain semantic alignment with the target relation.
[0143] Specifically, the triplet sorting loss Calculate using the following steps:
[0144] ;
[0145] in, For the first A triplet, To replace triples The negative sample triples consisting of the tail entities, For triples and negative sample triples; For triples The score is based on the embedding vector of the target relation obtained in step A24. negative sample triples The score is based on the embedding vector of the target relation obtained in step A24. This is a marginal parameter.
[0146] The path relationship consistency loss Calculate using the following steps:
[0147] First, for each valid path of each triple, cross-modal fusion features are performed through a linear layer. Mapping this to the embedding space of the knowledge graph embedding model yields its path relationship representation. :
[0148] ;
[0149] in, The weight matrix is a learnable linear transformation. The bias matrix is the learnable linear transformation. Represents the first of the corresponding triples One valid path;
[0150] Then, calculate the path relationship consistency loss of the effective path using the following formula:
[0151] ;
[0152] in, This refers to the embedding vector of the target relation obtained using a knowledge graph embedding model. Represents the L2 norm;
[0153] Then, the average of the path relationship consistency loss for each effective path in each triplet is used as the path relationship consistency loss. .
[0154] Furthermore, for the case where the encoding model is a pre-trained large language model encoder, in order to conveniently utilize the obtained effective paths without having to separately select effective paths for fine-tuning, thus saving training computing power, in step B4, each effective path... Convert to natural language description Then, using coding models to describe natural language Before encoding, it also includes:
[0155] Determine whether this is the first round of training. If so, perform a fine-tuning step for the encoding model; otherwise, skip the fine-tuning step. The fine-tuning step for the encoding model includes:
[0156] Based on each effective path Described in its natural language As input, the labels of the target relations contained in the triples to which they belong are used as the output targets to construct training samples; using the constructed training samples, the contrastive loss is calculated according to the following formula through a contrastive learning algorithm to fine-tune the encoding model;
[0157] ;
[0158] in, Indicates the number of valid paths. Indicates the first The semantic features obtained by encoding the natural language description of each effective path through the encoding model. The semantic features obtained by encoding the labels representing the target relationship through an encoding model; Indicates the temperature coefficient. This represents the similarity function.
[0159] The following description, in conjunction with specific examples, provides further details.
[0160] Example:
[0161] A few-shot knowledge graph completion method based on feature fusion, which is based on the head entities contained in each input triple. Target Relationship Tail-end entity The embedding vectors are used to calculate the matching score for each input triplet; based on the matching scores of each input triplet, the knowledge graph completion information is obtained; where each input triplet contains the head entity. Tail-end entity The embedding vectors are obtained using a knowledge graph embedding model; for each input triple, a feature fusion model is used to obtain the target relation for calculating its matching score. The embedding vector.
[0162] In this embodiment, the matching score of the triple is calculated using the following distance scoring function:
[0163] ;
[0164] in, Represents head entity The embedding vectors obtained using the knowledge graph embedding model Represents tail entity The embedding vectors obtained using the knowledge graph embedding model Indicates target relationship The embedding vector obtained using the feature fusion model.
[0165] In this embodiment, a feature fusion model is used to calculate the target relationship. The embedding vector includes:
[0166] A1. Path Mining
[0167] In this step, the Bi-BFS algorithm is used. Based on the topological information of the knowledge graph, for each input triple, the head entity it contains is extracted. Tail-end entity The length between them does not exceed Multi-hop paths.
[0168] The multi-hop path is a de novo entity. Reaching the tail entity via a path of more than one hop The path is represented as:
[0169] ;
[0170] in, This is the sequence number of the multi-hop path. Indicates a multi-hop path The included first A relationship, Indicates a multi-hop path The included first One entity, For multi-hop paths Number of contained relationships; multi-hop paths The entities included All are unique entities; multi-hop paths All the included triples belong to the triple set of the knowledge graph, that is:
[0171] ;
[0172] in, This represents the set of triples in a knowledge graph.
[0173] The This is the preset maximum path length. To control path semantic relevance and computational complexity, in this embodiment, the maximum path length is... Set it to 5.
[0174] A2. Path Filtering
[0175] In this step, for each input triple, the multi-hop paths based on that triple are calculated separately. Relationship with the target The degree of relevance is used to filter multi-hop paths that meet preset conditions, and these are considered valid paths for the triple.
[0176] In this embodiment, for each multi-hop path of each input triple, the relationship between the multi-hop path and the target is calculated separately according to the following steps. Relevance:
[0177] Based on multi-hop path The relationships included Relationship with the target Similarity to obtain multi-hop paths Relationship with the target Similarity between them; based on multi-hop paths Relationship with the target The similarity between them is used to obtain multi-hop paths. Relationship with the target Correlation weights between them; based on multi-hop paths Each triplet contained The connection strength is used to calculate multi-hop paths. Confidence weights; fusion of multi-hop paths The relevance weights and confidence weights are used to obtain the characteristics of multi-hop paths. Relationship with the target The overall weight of the degree of relevance.
[0178] Specifically, based on the following formula, using multi-hop paths... The relationships included Relationship with the target Similarity to obtain multi-hop paths Relationship with the target similarity between :
[0179] ;
[0180] in, This represents a multi-hop path obtained using a knowledge graph embedding model. The included number Embedding vectors of each relation, This represents the target relation obtained using a knowledge graph embedding model. The embedding vector; For multi-hop paths The number of relationships contained; Represents the vector dot product. This represents the L2 norm.
[0181] Based on the following formula, multi-hop paths Relationship with the target similarity between Obtain multi-hop paths Relationship with the target Correlation weights between :
[0182] ;
[0183] in, For head entity Tail-end entity A set of multi-hop paths between; This represents the natural exponential function.
[0184] Based on the following formula, multi-hop paths Each triplet contained The connection strength is used to calculate multi-hop paths. Confidence weight :
[0185] ;
[0186] ;
[0187] in, Represents a triplet The connection strength, Represents a triplet Frequency of occurrence in knowledge graphs This represents the total number of entities contained in the knowledge graph.
[0188] Merge multi-hop paths according to the following formula. Relevance weights and confidence weight Obtain the representation of multi-hop paths Relationship with the target The overall weight of the degree of relevance :
[0189] ;
[0190] in, To integrate the relevance weight and confidence weight, a balance coefficient is needed. This is a smoothing term.
[0191] In this embodiment, the overall screening weight meets a preset threshold. Multi-hop paths are considered valid paths. Preset threshold. Set it to 0.3, and set the maximum number of valid paths. The value is 50. During the screening process, the overall weight is considered. Initial path set for triples Perform filtering and retain overall weight. Greater than the preset threshold Multi-hop paths; when the number of retained multi-hop paths exceeds the maximum number of valid paths. At that time, according to the comprehensive weight Sort from highest to lowest, keeping the paths that appear earlier in the list.
[0192] A3. Extracting cross-modal fusion features
[0193] In this step, each valid path Convert them into natural language descriptions respectively. ; Utilize a fine-grained feature extraction network to capture each effective path Local interaction information to obtain each effective path Fine-grained features are extracted; a coarse-grained feature extraction network is used to capture each effective path. Global structural information to obtain each valid path The coarse-grained features; using the coding model, each effective path Natural Language Description Obtained by Transformation Encode the valid paths. The semantic features. For each valid path Separately, integrate this effective path The fine-grained and coarse-grained features are used to obtain the effective path. The structural characteristics; utilizing a fusion network to integrate this effective path. Based on the structural and semantic features, obtain the effective path. Cross-modal fusion features.
[0194] In this embodiment, the encoding model is an encoder of a pre-trained large language model; the fine-grained feature extraction network includes a bidirectional gated recurrent network and a graph attention network; the coarse-grained feature extraction network uses a Transformer encoder; and the fusion network is a gated network.
[0195] Each effective path Convert them into natural language descriptions respectively This involves representing the entity and relation sequences in a knowledge graph into a text format that can be processed by a large language model. Specifically, in this embodiment, the effective path is represented using the following formula. Convert to natural language description :
[0196] ;
[0197] For example, the path "h→r1→e1→r2→t" can be transformed into "entity h is connected to entity e1 through relation r1, and then connected to entity t through relation r2".
[0198] In this embodiment, the coding model is used to assign each valid path according to the following formula. Natural Language Description Obtained by Transformation Encode the data and extract the special classification markers from the last layer of the encoding model. The hidden state, as a valid path semantic features :
[0199] ;
[0200] in, This refers to the encoder of a pre-trained large language model that has undergone complete fine-tuning as an encoding model; Indicate extraction The hidden state of the marker.
[0201] This semantic feature, containing common-sense semantic information from large language models, can compensate for the lack of semantics in structured graph data in scenarios with few samples.
[0202] In this embodiment, a fine-grained feature extraction network is used to capture each effective path. Local interaction information to obtain each effective path Fine-grained features include:
[0203] AA1. Based on the effective paths of each triple obtained in step A2, construct the target relation. Path subgraph;
[0204] AA2. Using a bidirectional gated recurrent network, each valid path is encoded to obtain the encoded features of the entities and relationships contained in each valid path;
[0205] AA3. Using the encoded features of entities and relations contained in each effective path as input, and based on the correspondence between entities and relations, the fine-grained features of each effective path of each triple are calculated using a graph attention network based on the path subgraph.
[0206] In this embodiment, in step AA2, a bidirectional gated loop network is used to analyze the effective paths. Encode to obtain a valid path The encoded features of the entities and relations included include:
[0207] AA21. Using the embedding vectors of entities and relations contained in the knowledge graph obtained by the knowledge graph embedding model, obtain effective paths. path sequence Embedded sequences;
[0208] AA22, using an effective path Using the embedded sequence as input, and employing a bidirectional gated recurrent network, the effective path is obtained according to the following formula. Encoded features of the entities and relations included:
[0209] ;
[0210] ;
[0211] ;
[0212] in, This represents the first path in the input path sequence. An embedding vector of elements; This represents the first path in the input path sequence. The positive hidden state of an element. This represents the first path in the input path sequence. The reverse hidden state of each element. Indicates the first position in the input path sequence The encoding features of each element, This indicates vector concatenation.
[0213] In this embodiment, in step AA3, an effective path is calculated using a graph attention network based on the path subgraph. Fine-grained features include:
[0214] AA31. Based on the target relationship Obtain valid paths from the path subgraph. The set of neighbors of each entity contained therein;
[0215] AA32. Based on the correspondence between entities and relations, and according to the encoding characteristics of the entities and relations contained in each valid path, obtain the valid paths. The initial characteristics of each entity and its neighbors;
[0216] AA33, Based on Valid Paths The initial features of each entity and its neighbors are used to calculate the attention weights using the following formula:
[0217] ;
[0218] AA34. Based on attention weights, aggregate the data according to the following formula to obtain the effective path. Fine-grained features :
[0219] ;
[0220] in, For an effective path The The entity and its first Attention weights of each neighbor; For an effective path The The coding characteristics of an entity, and Valid paths The The first entity The neighbor and the first The encoded features of each neighbor; Let be the learnable linear transformation matrix of the graph attention network. For the learnable attention vectors of a graph attention network, This is the matrix transpose. For activation functions; Representation layer normalization.
[0221] In this embodiment, a coarse-grained feature extraction network is used to capture each effective path. Global structural information to obtain each valid path The coarse-grained characteristics include:
[0222] AB1, Valid Path path sequence By filling the mark Expand to the preset uniform length ;
[0223] AB2, Valid Path The expanded path sequence Input the Transformer encoder and extract using the following formula. The output of the marker is used as the valid path. coarse-grained characteristics :
[0224] ;
[0225] In this embodiment, for each valid path Specifically, the effective paths are merged according to the following formula. Fine-grained features and coarse-grained characteristics Obtain the valid path Structural features :
[0226] ;
[0227] in, To integrate fine-grained features and coarse-grained characteristics The balance coefficient.
[0228] In this embodiment, the effective path is fused using a fusion network. Based on the structural and semantic features, obtain the effective path. The cross-modal fusion features include:
[0229] AC1. Semantic features and structural features come from different sources and differ in representation space, dimension, and information emphasis. To avoid modal conflicts caused by direct concatenation, firstly, the structural features are processed according to the following formula. and semantic features Mapped to the same dimension:
[0230] ;
[0231] ;
[0232] in, and These are the weight matrices for the learnable linear transformation. and These are the bias matrices for the learnable linear transformations;
[0233] AC2, mapping structural features to the same dimension and semantic features The gating vector is obtained by calculating using the following formula. :
[0234] ;
[0235] in, This represents the Sigmoid activation function. This represents the gating weight matrix of the gating network. Represents the bias vector of the gated network;
[0236] AC3, based on the gate vector According to the following formula, the structural features are integrated. and semantic features Obtain a valid path Cross-modal fusion features :
[0237] ;
[0238] in, This represents element-wise multiplication.
[0239] ;
[0240] in, To integrate fine-grained features and coarse-grained characteristics The balance coefficient.
[0241] A4. Obtain the target relationship Embedded vector
[0242] In this step, for each input triple, the effective paths of that triple are aggregated separately. Cross-modal fusion features are used to obtain the target relation for calculating the matching score of the triple. The embedding vector.
[0243] In this embodiment, for each input triple, an attention mechanism is used to aggregate the effective paths of the triple according to the following formula. Cross-modal fusion features Obtain the target relation used to calculate the matching score of the triple. Embedded vector :
[0244] ;
[0245] ;
[0246] ;
[0247] in, Here is the transformation matrix for the attention mechanism. The set of valid paths corresponding to the triples. Transpose of the representation matrix.
[0248] The training of the feature fusion model, such as Figure 1 As shown, it includes the following steps:
[0249] B1. Constructing Training Samples
[0250] In this step, based on the input target relation and the known support samples, triples are constructed for this round of training.
[0251] The constructed triples constitute the target relation. The support set is represented as:
[0252] ;
[0253] in, Indicates the first The header entities of the known supporting samples, Indicates the first Tail entities of known supporting samples, This indicates the number of known supporting samples.
[0254] B2. Path Mining
[0255] In this step, for each input triple, the multi-hop path of the triple is mined according to the method in step A1.
[0256] B3. Path Filtering
[0257] In this step, for each input triplet, the valid paths for that triplet are obtained by filtering according to the method in step A2.
[0258] B4. Extracting cross-modal fusion features
[0259] In this step, for each input triple, the cross-modal fusion characteristics of each effective path of the triple are obtained according to the method in step A3.
[0260] Because the feature fusion model in this embodiment uses an encoder for a pre-trained large language model, it facilitates the use of already obtained effective paths without the need for separate selection of effective paths for fine-tuning of a large model, thus saving training computational power. In this embodiment, each effective path... Convert to natural language description Then, using coding models to describe natural language Before encoding, it also includes:
[0261] Determine if this is the first round of training. If so, perform the fine-tuning step of the encoding model; otherwise, skip the fine-tuning step of the encoding model.
[0262] B5. Obtaining the target relationship Embedded vector
[0263] In this step, for each input triple, the embedding vector of the target relation of the triple is obtained according to the method in step A4.
[0264] B6. Loss Calculation
[0265] In this step, the total loss is calculated using the following formula. :
[0266] ;
[0267] in, The loss is for sorting triples. For path relationship consistency loss, and As weight.
[0268] B7. Model Update
[0269] In this step, it is determined whether the training is complete. If yes, the training ends; otherwise, return to step B1.
[0270] In this embodiment, the triplet sorting loss Calculate using the following steps:
[0271] ;
[0272] in, For the first A triplet, To replace triples The negative sample triples consisting of the tail entities, For triples and negative sample triples; For triples The score is based on the embedding vector of the target relation obtained in step A24. negative sample triples The score is based on the embedding vector of the target relation obtained in step A24. These are marginal parameters;
[0273] In this embodiment, the path relationship consistency loss Calculate using the following steps:
[0274] First, for each valid path of each triple, cross-modal fusion features are performed through a linear layer. Mapping this to the embedding space of the knowledge graph embedding model yields its path relationship representation. :
[0275] ;
[0276] in, The weight matrix is a learnable linear transformation. The bias matrix is the learnable linear transformation. Represents the first of the corresponding triples One valid path;
[0277] Then, calculate the path relationship consistency loss of the effective path using the following formula:
[0278] ;
[0279] in, This refers to the embedding vector of the target relation obtained using a knowledge graph embedding model. Represents the L2 norm;
[0280] Then, the average of the path relationship consistency loss for each effective path in each triplet is used as the path relationship consistency loss. .
[0281] In this embodiment, the fine-tuning step of the encoding model includes:
[0282] Based on each effective path Described in its natural language As input, the labels of the target relations contained in the corresponding triples are used as the output targets to construct training samples. These training samples enable the encoding model to learn the mapping relationship between the natural language description of the effective path and the target relations.
[0283] Using the constructed training samples, the contrastive loss is calculated according to the following formula through the contrastive learning algorithm, and the coding model is fine-tuned using the low-rank adaptation method;
[0284]
[0285] in, Indicates the number of valid paths. Indicates the first The semantic features obtained by encoding the natural language description of each effective path through the encoding model. The semantic features obtained by encoding the labels representing the target relationship through an encoding model; Indicates the temperature coefficient. This represents the similarity function.
[0286] Low-Rank Adaptation (LoRA) reduces training costs by inserting a low-rank matrix into the Transformer layer of a large language model, training only the parameters of the newly added low-rank matrix, and freezing the main parameters of the base model.
[0287] In some embodiments, natural language descriptions can be expanded through synonym substitution, sentence transformation, and other methods. This is to enhance the coding model's adaptability to different natural language expressions.
[0288] Test 1:
[0289] The method described in this embodiment was applied to the NELL-One few-shot knowledge graph completion dataset for testing. During testing, the hidden layer dimension of the bidirectional gated recurrent network was set to 256, the output dimension of the graph attention network was set to 256, the number of layers in the Transformer encoder (used as a coarse-grained feature extraction network) was set to 2, and the number of attention heads was set to 8. The learning rate for fine-tuning the large model was set to 5e-5, and the overall model training learning rate was set to 1e-4.
[0290] The results are shown in Table 1, where TSKFC is the implementation method. Test results show that in the 1-shot scenario, the implementation method has the highest MRR value compared to other baseline models, and its Hits@1, 5, and 10 are also among the first and second highest. Therefore, the implementation method can achieve more stable inference performance through multi-hop path evidence and large model semantic enhancement.
[0291] Table 1. Experimental results using the NELL-One dataset
[0292]
[0293] Test 2:
[0294] The method described in the examples was applied to the Wiki-One dataset for testing. The Wiki-One dataset contains a large number of encyclopedic entities and relationships, with complex relationship types and a small number of samples for some relationships. This scenario can test the model's ability to complete complex semantic relationships and long-tail relationships.
[0295] Table 2. Experimental results on the Wilki-One dataset.
[0296]
[0297] The results are shown in Table 2, where TSKFC is the method used in the embodiment. The results show that the method in this embodiment has good completion performance in the 1-shot scenario where samples are extremely scarce, and the changes are not significant in the 3-shot and 5-shot scenarios, maintaining good stability.
[0298] Finally, it should be noted that the above embodiments are merely preferred embodiments and are not intended to limit the present invention. It should be pointed out that those skilled in the art can make various modifications, equivalent substitutions, and improvements without departing from the spirit and scope of the claims, and all such modifications, substitutions, and improvements should be included within the scope of protection of the present invention.
Claims
1. A few-shot knowledge graph completion method based on feature fusion, which is based on the head entities contained in each input triple. Target Relationship Tail-end entity The embedding vectors are used to calculate the matching score for each input triplet; based on the matching scores of each input triplet, the knowledge graph completion information is obtained; where, The head entities contained in each input triple Tail-end entity The embedding vectors are obtained using knowledge graph embedding models; the key feature is that, for each input triple, a feature fusion model is used to obtain the target relation for calculating its matching score. Embedding vector: The target relationship is calculated using a feature fusion model. The embedding vector includes: A1. Based on the topological information of the knowledge graph, for each input triple, separately mine the head entities it contains. Tail-end entity The length between them does not exceed The multi-hop path, the The preset maximum path length; The multi-hop path is a de novo entity. Reaching the tail entity via a path of more than one hop The path is represented as: ; in, This is the sequence number of the multi-hop path. Indicates a multi-hop path The included first A relationship, Indicates a multi-hop path The included first One entity, For multi-hop paths Number of contained relationships; multi-hop paths The included entities All are unique entities; multi-hop paths All the included triples belong to the triple set of the knowledge graph; A2. For each input triple, respectively, calculate the multi-hop paths based on that triple. Relationship with the target Based on the degree of relevance, multi-hop paths that meet the preset conditions are selected as valid paths for the triplet; A3. Utilize a fine-grained feature extraction network to capture each effective path. Local interaction information to obtain each effective path Fine-grained features are extracted; a coarse-grained feature extraction network is used to capture each effective path. Global structural information to obtain each valid path The coarse-grained characteristics; the effective paths Convert them into natural language descriptions respectively. Using the coding model, each valid path Natural Language Description Obtained by Transformation Encode the valid paths. semantic features; For each valid path Separately, integrate this effective path The fine-grained and coarse-grained features are used to obtain the effective path. The structural characteristics; utilizing a fusion network to fuse this effective path. Based on the structural and semantic features, obtain the effective path. Cross-modal fusion features; A4. For each input triple, aggregate the valid paths for that triple separately. Cross-modal fusion features are used to obtain the target relation for calculating the matching score of the triple. The embedding vector.
2. The few-shot knowledge graph completion method based on feature fusion as described in claim 1, characterized in that: Step A1 employs the Bi-BFS algorithm, based on the topological information of the knowledge graph, to extract the head entities contained in each input triple. Tail-end entity The length between them does not exceed Multi-hop paths.
3. The few-shot knowledge graph completion method based on feature fusion as described in claim 1, characterized in that: In step A2, for each multi-hop path of each input triple, the relationship between the multi-hop path and the target is calculated separately according to the following steps. Relevance: Based on multi-hop path The relationships included Relationship with the target Similarity to obtain multi-hop paths Relationship with the target Similarity between them; based on multi-hop paths Relationship with the target The similarity between them is used to obtain multi-hop paths. Relationship with the target Correlation weights between them; based on multi-hop paths Each triplet contained The connection strength is used to calculate multi-hop paths. Confidence weights; fusion of multi-hop paths The relevance weights and confidence weights are used to obtain the characteristics of multi-hop paths. Relationship with the target The overall weight of the degree of relevance.
4. The few-shot knowledge graph completion method based on feature fusion as described in claim 3, characterized in that: In step A2, based on the following formula, the multi-hop path... The relationships included Relationship with the target Similarity to obtain multi-hop paths Relationship with the target Similarity between : ; in, This represents a multi-hop path obtained using a knowledge graph embedding model. The included number Embedding vectors of relations, This represents the target relation obtained using a knowledge graph embedding model. The embedding vector; For multi-hop paths The number of relationships contained; Represents the vector dot product. Represents the L2 norm; In step A2, based on the following formula, the multi-hop path... Relationship with the target Similarity between Obtain multi-hop paths Relationship with the target Correlation weights between : ; in, For head entity Tail-end entity A set of multi-hop paths between; Represents the natural exponential function; In step A2, based on the following formula, the multi-hop path... Each triplet contained The connection strength is used to calculate multi-hop paths. Confidence weight : ; ; in, Represents a triplet The connection strength, Represents a triplet Frequency of occurrence in knowledge graphs This represents the total number of entities contained in the knowledge graph; In step A2, multi-hop paths are merged according to the following formula. Relevance weights and confidence weight Obtain the representation of multi-hop paths Relationship with the target The overall weight of the degree of relevance : ; in, To integrate the relevance weight and confidence weight, a balance coefficient is needed. This is a smoothing term.
5. A few-shot knowledge graph completion method based on feature fusion as described in any one of claims 1 to 4, characterized in that: The fine-grained feature extraction network includes a bidirectional gated recurrent network and a graph attention network; In step A3, a fine-grained feature extraction network is used to capture each effective path. Local interaction information to obtain each effective path Fine-grained features include: AA1. Based on the effective paths of each triple obtained in step A2, construct the target relation. Path subgraph; AA2. Using a bidirectional gated recurrent network, each valid path is encoded to obtain the encoded features of the entities and relationships contained in each valid path; AA3. Using the encoding features of entities and relations contained in each effective path as input, and based on the correspondence between entities and relations, the fine-grained features of each effective path of each triple are calculated using a graph attention network based on the path subgraph. In step AA2, a bidirectional gated cyclic network is used to analyze the effective paths. Encode to obtain a valid path The encoded features of the entities and relations included include: AA21. Using the embedding vectors of entities and relations contained in the knowledge graph obtained by the knowledge graph embedding model, obtain effective paths. path sequence Embedded sequences; AA22, using an effective path Using the embedded sequence as input, and employing a bidirectional gated recurrent network, the effective path is obtained according to the following formula. Encoded features of the entities and relations included: ; ; ; in, This represents the first path in the input path sequence. An embedding vector of elements; This represents the first path in the input path sequence. The positive hidden state of an element. This represents the first path in the input path sequence. The reverse hidden state of each element. Indicates the first position in the input path sequence The encoding features of each element, This indicates vector concatenation; In step AA3, effective paths are calculated using a graph attention network based on the path subgraph. Fine-grained features include: AA31. Based on the target relationship Obtain valid paths from the path subgraph. The set of neighbors of each entity contained therein; AA32. Based on the correspondence between entities and relations, and according to the encoding characteristics of the entities and relations contained in each valid path, obtain the valid paths. The initial characteristics of each entity and its neighbors; AA33, Based on Valid Paths The initial features of each entity and its neighbors are used to calculate the attention weights using the following formula: ; AA34. Based on attention weights, aggregate the data according to the following formula to obtain the effective path. Fine-grained features : ; in, For an effective path The The entity and its first Attention weights of each neighbor; For an effective path The The coding characteristics of an entity, and Valid paths The The first entity The neighbor and the first The encoded features of each neighbor; Let be the learnable linear transformation matrix of the graph attention network. For the learnable attention vectors of a graph attention network, This is the matrix transpose. For activation functions; Representation layer normalization.
6. A few-shot knowledge graph completion method based on feature fusion as described in any one of claims 1 to 4, characterized in that: The encoding model is an encoder for a pre-trained large language model; the coarse-grained feature extraction network uses a Transformer encoder. In step A3, the coding model is used to analyze each valid path. Natural Language Description Obtained by Transformation Encode the model and extract the hidden states of the last layer of the encoding model that contain special classification labels, as the valid paths. semantic features; In step A3, a coarse-grained feature extraction network is used to capture each effective path. Global structural information to obtain each valid path The coarse-grained characteristics include: AB1, Valid Path path sequence By filling the mark Expand to the preset uniform length; AB2, Valid Path The expanded path sequence is input into the Transformer encoder to extract... The output of the marker is used as the valid path. The coarse-grained characteristics; In step A3, for each valid path Specifically, the effective paths are merged according to the following formula. Fine-grained features and coarse-grained characteristics Obtain the valid path Structural features : ; in, To integrate fine-grained features and coarse-grained characteristics The balance coefficient.
7. A few-shot knowledge graph completion method based on feature fusion as described in any one of claims 1 to 4, characterized in that: The fusion network is a gated network; in step A3, the effective path is fused using the fusion network. Based on the structural and semantic features, obtain the effective path. The cross-modal fusion features include: AC1. According to the following formula, the structural features are... and semantic features Mapped to the same dimension: ; ; in, and These are the weight matrices for the learnable linear transformation. and These are the bias matrices for the learnable linear transformations; AC2, mapping structural features to the same dimension and semantic features The gating vector is obtained by calculating using the following formula. : ; in, This represents the Sigmoid activation function. This represents the gating weight matrix of the gating network. Represents the bias vector of the gated network; AC3, based on the gate vector According to the following formula, the structural features are integrated. and semantic features Obtain a valid path Cross-modal fusion features : ; in, This represents element-wise multiplication.
8. A few-shot knowledge graph completion method based on feature fusion as described in any one of claims 1 to 4, characterized in that: In step A4, for each input triple, an attention mechanism is used to aggregate the effective paths of that triple according to the following formula. Cross-modal fusion features Obtain the target relation used to calculate the matching score of the triple. Embedded vector : ; ; ; in, Here is the transformation matrix for the attention mechanism. The set of valid paths corresponding to the triples. Transpose of the representation matrix.
9. The few-shot knowledge graph completion method based on feature fusion as described in claim 1, characterized in that: The training of the feature fusion model includes the following steps: B1. Based on the target relationship of the input, construct triples for this round of training according to the known support samples; B2. For each input triple, mine the multi-hop path of the triple using the method in step A1. B3. For each input triplet, filter and obtain the valid path of the triplet according to the method in step A2. B4. For each input triplet, obtain the cross-modal fusion features of each effective path of the triplet according to the method in step A3. B5. For each input triplet, obtain the embedding vector of the target relation of the triplet according to the method in step A4. B6. Calculate the total loss using the following formula. : ; in, The loss is for sorting triples. For path relationship consistency loss, and As weight; B7. Determine whether the training is complete. If yes, end the training; otherwise, return to step B1. The triplet sorting loss Calculate using the following steps: ; in, For the first A triplet, To replace triples The negative sample triples consisting of the tail entities, For triples and negative sample triples; For triples The score based on the embedding vector of the target relation obtained in step B5, negative sample triples The score based on the embedding vector of the target relation obtained in step B5, These are marginal parameters; The path relationship consistency loss Calculate using the following steps: First, for each valid path of each triple, cross-modal fusion features are performed through a linear layer. Mapping this to the embedding space of the knowledge graph embedding model yields its path relationship representation. : ; in, The weight matrix is a learnable linear transformation. is the bias matrix of the learnable linear transformation; Represents the first of the corresponding triples One valid path; Then, calculate the path relationship consistency loss of the effective path using the following formula: ; in, This refers to the embedding vector of the target relation obtained using a knowledge graph embedding model. Represents the L2 norm; Then, the average of the path relationship consistency loss for each effective path in each triplet is used as the path relationship consistency loss. .
10. The few-shot knowledge graph completion method based on feature fusion as described in claim 9, characterized in that: The encoding model is an encoder for a pre-trained large language model; In step B4, each valid path is... Convert to natural language description Then, using coding models to describe natural language Before encoding, it also includes: Determine if this is the first round of training; if so, perform the fine-tuning step for the encoding model; otherwise, skip the fine-tuning step for the encoding model. The fine-tuning steps of the encoding model include: Based on each effective path Described in its natural language As input, the labels of the target relations contained in the triples to which they belong are used as the output targets to construct training samples; using the constructed training samples, the contrastive loss is calculated according to the following formula through a contrastive learning algorithm to fine-tune the encoding model; ; in, Indicates the number of valid paths. Indicates the first The semantic features obtained by encoding the natural language description of each effective path through the encoding model. The semantic features obtained by encoding the labels representing the target relationship through an encoding model; Indicates the temperature coefficient. This represents the similarity function.