Multi-hop knowledge graph reasoning method based on query structure coding

By encoding and fusion of context information in the knowledge graph, the problem of difficulty in generalizing to unknown entities and relationships in the prior art is solved, and efficient multi-hop reasoning and complex problem-solving capabilities of knowledge graphs are achieved.

CN119988645AInactive Publication Date: 2025-05-13DALIAN NATIONALITIES UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510158233.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to generalize to unseen entities and relationship types in the knowledge graph, and it is difficult to effectively perform multi-hop query inference.

Method used

By encoding the query structure, the pre-trained language model is used to capture the logical structure information of the query, and multi-hop reasoning is performed based on the context information of the query structure. Specific methods include designing structured prompts, using geometric operation modeling, graph neural network coding and similarity calculation, etc.

Benefits of technology

The ability to generalize entities and relationships without seeing is realized, the performance and efficiency of inductive logical reasoning is improved, and the performance of knowledge graphs in complex problem-solving and knowledge discovery is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988645A_ABST
    Figure CN119988645A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-hop knowledge graph reasoning method based on query structure coding, and belongs to the field of artificial intelligence. According to the technical scheme, a linearization template is selected according to a query type, and a query structure is converted into a text sequence; an instruction for indicating a geometric operation execution sequence is introduced, and implicit structure information is used for prompting a pre-training encoder. The context of the query structure is captured by identifying the location, the function role and the query type of the node. A BERT is used to pre-train an encoder, and geometric operations are modeled. In the graph neural network, the message is transmitted and updated, the feature vector of the node is updated to be embedded representation, then the similarity between the query node and the candidate entity is calculated, and the prediction task of the entity and the relationship is processed through the cross entropy loss. The method has the beneficial effects that the query statement is processed, multi-hop reasoning is effectively performed through structured prompt and geometric operation modeling, and meanwhile, a more comprehensive knowledge graph is constructed according to the context information of the query structure, so that the reasoning accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of knowledge graph reasoning based on artificial intelligence, and in particular to an inductive logic reasoning method based on a knowledge graph and a query structure encoding method. Background Art

[0002] The knowledge graph is a structured representation of facts or knowledge. It is a network structure composed of entities and relationships between entities. An entity is an independent thing with clear characteristics that can be distinguished from other things. In the knowledge graph, the information used to describe these things is the entity. Entities are represented by vertices in the attribute graph. The type of entity association is the entity type, which is represented by vertex labels in the attribute graph. Relationships express a certain semantic relationship between two entities, usually represented by semantic labels and represented as directed edges in the attribute graph. In other words, the knowledge graph G consists of a series of triples:<h,r,t> where h and t represent the head entity and the tail entity respectively, and r represents the directed relationship from h to t.

[0003] Knowledge graph reasoning is knowledge reasoning for knowledge graphs. There are many definitions of knowledge graphs, which are knowledge bases that represent the relationships between entities in the real world in the form of graphs. Knowledge reasoning in knowledge graphs aims to identify errors and infer new conclusions from existing data. Through knowledge reasoning, new relationships between entities can be derived and feedback can be provided to enrich the knowledge graph and provide support for advanced applications.

[0004] At present, there are two main methods for inductive logic reasoning. One is to inherit the embedding-based method and combine types as additional information to improve the inductive ability, but this method cannot be generalized to invisible entity and relationship types. The second is to use a pre-trained language model to encode the text information of entities and relationships to generalize to unseen elements. In order to solve the above problems, the query structure is modeled and integrated into the inductive logic reasoning of the knowledge graph. The query structure encoding and the contextual information of the query structure are used to enhance the performance and efficiency of multi-hop queries. It is of great research value to construct a multi-hop knowledge graph reasoning method based on query structure encoding. Summary of the invention

[0005] In order to meet the needs of the above-mentioned prior art, the present invention provides a multi-hop knowledge graph reasoning method based on query structure encoding, which uses a pre-trained language model to encode the query structure and performs inductive logic reasoning, and explicitly captures the logical structure information of the query by designing structured prompts and geometric operation modeling. Combined with the structural context inherent in the query structure and the relationship-induced context unique to each node in the query graph, a refined internal representation is obtained in the entire multi-hop reasoning step. This method can be generalized to unseen entities and relationships, and even different knowledge graphs, rather than being limited to generalization within the same knowledge graph, thereby improving the ability of inductive logic reasoning. Through context-aware query representation learning, the intent of the question can be understood more accurately, and accurate answers can be generated through multi-hop reasoning, significantly improving the performance of knowledge graphs in complex problem solving and knowledge discovery.

[0006] The technical solution is as follows:

[0007] A multi-hop knowledge graph reasoning method based on query structure encoding, the steps are as follows:

[0008] Step 1: According to the type of query statement, select a suitable linearization template to convert the query structure into a text sequence, thereby retaining the structural information of the query statement;

[0009] Step 2: Design structured hints for each query type to indicate the order in which geometric operations are performed in the query, thereby helping the pre-trained language model capture the structural information of the query;

[0010] Step 3: Use position embedding, role embedding, and type embedding to learn structural context and fuse it with the model;

[0011] Step 4: Use the pre-trained encoder, attention layer, and maximum layer to model the three geometric operations of projection, intersection, and union in the representation space respectively;

[0012] Step 5: Use graph neural network to encode query and candidate entities respectively and calculate the similarity between them;

[0013] Step 6: Optimize the contrast loss and classification loss for training. During inference, use the query structure that incorporates context information for query reasoning.

[0014] For steps 1 and 2, for each triple in the query, denote it as [anchor]t(e i )[relation]t(r i), introduces step-by-step instructions that indicate the order in which geometric operations are performed. The instructions include two parts: the total number of execution steps and the order of operations in the query. The structure hint s and the linearized query structure t(q) are connected to form the input format [CLS][qtype]s[SEP]t(q) and input to the training encoder.

[0015] For step 3, an adjacency matrix is ​​constructed, where the rows and columns correspond to the nodes in the query graph. If there is an edge between node i and node j, the element in the i-th row and j-th column of the matrix is ​​1, otherwise it is 0, indicating that there is no edge between node i and node j. The adjacency matrix can effectively represent the connection relationship between nodes, thereby obtaining the set of neighbor nodes of a node. The calculation process is:

[0016] N(i)={j|A(i,j)}

[0017] Among them, A(i, j) represents the element in the i-th row and j-th column in the adjacency matrix.

[0018] Using the head-to-tail mapping method, a mapping table is established. The key of the table is the relationship number, and the value is a set of two tuples. Each node is assigned a position number and a role number. The first number of the tuple of each node in the query graph represents the position number, and the second number represents the role number. The calculation process is:

[0019] Rout(i)={r|(i,r,j)∈G}

[0020] Rin(i)={r|(j,r,i)∈G}

[0021] Among them, Rout(i) is the set of outgoing edge relations, Rin(i) is the set of incoming edge relations, G is the query graph, and (i, r, j) is a triple in the query graph.

[0022] After that, the position and role encoding query table is flattened to obtain a vector, and the obtained vector is normalized. A linear transformation is used to map this normalized vector to a continuous embedding space to create a type embedding, thereby capturing the structural context of a given query type.

[0023] For step 4, the pre-trained encoder is extended with additional attention layers and maximum output layers to model different geometric operations such as projection, intersection, and union on the representation space, respectively, to implicitly inject structured modeling into BERT. The SILR model models different geometric operations on the representation space including:

[0024] (1) Projection operation: For queries that only contain relational projection operations, the average of the representations of all target entities is taken as the query representation. The calculation process is:

[0025]

[0026] Where n is the number of target entities, h(v i ) is the representation of the i-th target entity.

[0027] (2) Intersection operation: For queries containing intersection operations, an attention layer is used to model the target entities in multiple subqueries. The input of the attention layer is the representation of the target entity in all subqueries, and the output is the weighted average of the representation of the target entity in each subquery. The calculation process is:

[0028] h{v i} w =α i h{v i}

[0029] Among them, α i is the attention weight, h{v i} is the representation of the target entity in the i-th subquery.

[0030] (3) Union operation: For queries that contain union operations, the maximum layer is used to model the target entities in multiple subqueries. The input of the maximum layer is the representation of the target entity in all subqueries, and the output is the maximum value of all target entity representations. The calculation process is:

[0031] h q =max(h{v i})

[0032] Among them, h q This is the query representation after the union operation.

[0033] For step 5, the graph neural network transfers information through adjacent nodes. The feature vector of each node is updated by aggregating information from adjacent nodes. After repeated information transfer and updating, the feature vector of each node is updated to the embedded representation of the node in the knowledge graph, and then the similarity between the query node and the candidate entity node is calculated to further narrow the answer range.

[0034] The query structure encoding includes:

[0035] (1) Call BERT: By calling the function, the input token ID and other optional parameters are passed into the BERT model. During this process, by executing the forward propagation of the BERT model, the model will calculate the representation of each token based on the input token ID and other information. The BERT model converts the input token ID into an embedding vector, and combines the position encoding and token type encoding to generate the input representation. Since BERT is composed of multiple stacked Transformer encoder layers, each layer performs self-attention calculations and feedforward neural network processing on the input to capture contextual information. Finally, the BERT model outputs the representation of each token, which contains contextual information.

[0036] (2) Pooled output: The output of the last hidden layer is obtained through the BERT model, and then the output is averaged and pooled to obtain the pooled output representing the entire input sequence. Average pooling effectively integrates the information of all tokens in the sequence and captures the overall semantics of the input text. The calculation process is expressed as:

[0037]

[0038] in, is the length of the input sequence, i.e. sequence_length, Indicates The hidden state of tokens has a shape of (batch_size, hidden_size).

[0039] The context of the query structure includes:

[0040] (1) Embedding dimension: Define the embedding and initialize it with a uniform distribution. Dynamically set the dimension of the embedding according to the specific configuration of the model. Combine the embedding with other features through matrix multiplication in the forward propagation to generate the final embedding representation.

[0041] (2) Query graph adjacency: Different types of query graph structures are defined by creating dictionaries. The following are the definitions and execution standards of 14 query types:

[0042] 1p (Path-1): There is a path in the query graph with a path length of 1, indicating that there is a relationship connecting the anchor node and the answer node;

[0043] 2p (Path-2): There is a path in the query graph with a path length of 2, indicating that there are two relations connecting the anchor node and the answer node;

[0044] 3p (Path-3): There is a path in the query graph with a path length of 3, indicating that there are three relationships connecting the anchor node and the answer node;

[0045] pi (Path-1 with Inverse): There is a path in the query graph with a path length of 1, and one of the relations is an inverse relation;

[0046] ip(Inverse-Path-1): There is a path in the query graph with a path length of 1 and the relationship is an inverse relationship;

[0047] 2i (Path-2 with Inverse): There is a path in the query graph with a path length of 2, and one of the relations is an inverse relation;

[0048] 3i (Path-3 with Inverse): There is a path in the query graph with a path length of 3, and one of the relations is an inverse relation;

[0049] pin(Path-1with Intersection): There are two paths in the query graph, both with a path length of 1, and these two paths share an answer node;

[0050] pni (Path-1 with Non-Intersection): There are two paths in the query graph, both with a path length of 1, and these two paths do not share the answer node;

[0051] 2in (Path-2 with Intersection): There are two paths in the query graph, both with a path length of 2, and these two paths share an answer node;

[0052] 3in (Path-3 with Intersection): There are two paths in the query graph, both with a path length of 3, and these two paths share an answer node;

[0053] 2u (Path-2 with Union): There are two paths in the query graph, both with a path length of 2, and the union of the answer node sets of these two paths is the final answer node set;

[0054] 3u (Path-3 with Union): There are two paths in the query graph, both of which have a length of 3, and the union of the answer node sets of these two paths is the final answer node set;

[0055] up(Union with Path): There is a path and an answer node set in the query graph. The union of the path's answer node set and the answer node set is the final answer node set.

[0056] Each structure represents a specific relational pattern. The key of the dictionary is the type of the query graph, and the corresponding value is a dictionary containing the input and output relations.

[0057] (3) Construct relationship head-tail mapping: Determine whether the entity and relationship indexes need to be remapped based on the name of the data set, and load the corresponding entity and relationship mapping files. Obtain the triples of the knowledge graph by reading the training data set, and construct relationship-to-index and entity-to-index mapping dictionaries, and count the number of relationships and entities. By traversing the training data, extract the head entity, relationship, and tail entity, perform index mapping based on the mapping relationship, add the head entity to the head entity set of the corresponding relationship, add the tail entity to the tail entity set, calculate the frequency of the head entity and tail entity of each relationship, and find the maximum frequency and median frequency, so as to better understand the distribution of the relationship. The calculation process is expressed as:

[0058] F h (ri)=count({h|(h, ri, t)∈triples})

[0059] F t (r i )=count({t|(h,r i , t)∈triples})

[0060] F h,max =max(F h )

[0061] F t,max =max(F t )

[0062] F h,median =median(F h )

[0063] F t,median =median(F t )

[0064] Among them, triples represents all triples in the knowledge graph, is the head entity, and is the tail entity.

[0065] (4) Mask calculation: Mask calculation is used to indicate which entities are valid and which are filled, so as to achieve the purpose of correctly processing input data during training. There are two types of mask calculation: without sampling and with sampling. Without sampling, a mapping is constructed based on the number of head entities and tail entities of each relationship, and it is filled to the maximum frequency. With sampling, the head entity and tail entity of each relationship are randomly selected to ensure that the number does not exceed the specified sampling number, and the insufficient part is filled with a special value.

[0066] (5) Query2box: The entities and relations of the encoded query structure are embedded into a collection space, and geometric shapes are used to represent and reason about complex relationships in the knowledge graph. The knowledge graph embedding method is adopted to enhance the model's ability to handle complex queries, thereby improving the accuracy of reasoning.

[0067] (6) Normalization of feature matrices: Traverse each feature matrix in the query type dictionary and calculate the sum of the elements of each matrix. Find the largest element from the list, which is the maximum value in the sum of the feature matrices. Then normalize the feature matrix and fix the element values ​​of all feature matrices between 0 and 1. The normalization operation can prevent certain features from dominating the model training, resulting in a decrease in model performance. At the same time, it can accelerate the convergence speed of the model and improve the efficiency and effect of subsequent model training. The calculation process is expressed as:

[0068]

[0069]

[0070] Among them, query2table[qtype] is the feature matrix of each query type, query2table_norm[qtype] is the normalized feature matrix, To sum the feature matrices for each query type, find the maximum value in the sum of all feature matrices.

[0071] (7) One-hot encoding: By traversing each query type, transposing the dimension of the feature matrix of each query type, flattening the transposed matrix into a one-dimensional tensor, and finally moving it to the specified device. One-hot encoding converts the feature matrix of the query type into a format suitable for machine learning model input, so as to better capture the relationship between features.

[0072] The BERT model includes:

[0073] (1) Transformer model: Apply dropout operation to prevent overfitting, project features into output space, generate logits, and then process the logits. If the last dimension of logits is 1, indicating that there is only one output, this dimension is removed by squeeze for subsequent processing. Finally, the processed logits are returned for classification. The representation of the candidate entity is calculated, the query sequence is encoded, and the mask representation is extracted from it. The calculated logits are converted into probability scores through softmax.

[0074] (2) Masked Language Model: Randomly mask some words in a sentence and then predict these masked words, so that the model learns to infer the missing words from the context.

[0075] (3) Next Sentence Prediction: By training the model to determine whether two given sentences are continuous in the original text, it helps the model learn how to capture the logical relationship and context between sentences.

[0076] (4) Euclidean distance calculation method: By inputting two vectors, calculating the difference between the two vectors and obtaining their L2 norm, the calculated Euclidean distance is returned, which represents the distance between the two input vectors. These two vectors are the feature representations of the query-matched entity and the candidate entity, respectively. During the training process, the Euclidean distance can be used to calculate the loss function to help the model optimize its parameters so that similar inputs are closer in the feature space. During the inference stage, the model may need to select one from multiple candidate answers, and the Euclidean distance can help determine which candidate answer is closest to the query representation. The calculation process is expressed as:

[0077] distamce=rep1-rep2

[0078]

[0079] Among them, rep1 and rep2 represent two feature vectors that need to be compared or calculated, and distance represents the distance measure between two feature representations (or vectors), which is obtained by calculating the Euclidean distance between rep1 and rep2, indicating the similarity or difference between the two vectors in the feature space.

[0080] (5) Cross entropy calculation method: After the softmax function is processed, the probability distribution is obtained. The predicted probability corresponding to each sample is obtained through the index. The natural logarithm of these probabilities is calculated. In order to avoid the case of zero logarithm, a constant 1e-5 is added, and the log likelihood is obtained by taking a negative value. Use torch.sum to calculate the sum of the log likelihoods of all samples. Then calculate the average loss according to the number of samples. The calculation process is expressed as:

[0081]

[0082] Among them, m is the number of samples and is the predicted probability corresponding to the true category of the jth sample.

[0083] (1) Graph Isomorphic Network: Through a powerful node feature update mechanism, nodes can capture local structural information in the graph and distinguish different graph isomorphisms.

[0084] (2) Message passing mechanism: Through an aggregation-based message passing mechanism, the representation of a node is updated by aggregating the features of its neighbors. The calculation process is expressed as:

[0085]

[0086] in Represents the characteristics of the nodes in the layer, Represents the set of neighbor nodes of a node, MLP (l) It is a multi-layer perceptron layer used for feature updating.

[0087] (3) Feature update: The graph isomorphism network uses a multi-layer perceptron as a feature update function to directly concatenate the aggregated features with the node's own features, and then undergoes a nonlinear transformation, so that the model can better learn and distinguish structural information. The calculation process is expressed as:

[0088]

[0089] in It is the sum of the features of all neighboring nodes of a node.

[0090] Graph-level representation: Aggregate and summarize the entire entity and relationship information into a single vector or feature representation. Perform high-level analysis and prediction by extracting global information from complex graph structures.

[0091] The beneficial effects of the present invention are:

[0092] The present invention uses the above-mentioned technical solution, uses the BERT model to encode the linearized query structure and entity, and designs a step-by-step indication to prompt the pre-trained language model to perform geometric operations in the query. The Transformer encoder is used to model the multi-hop relationship projection, and the intersection and union of multiple sub-queries are modeled using the attention layer and the maximum output layer. The query structure and candidate entities are encoded using a graph neural network, and the node neighbors are feature-summarized to update the node representation, thereby further narrowing the answer range. In addition, the location clues and functional roles of the nodes are combined by traversing the adjacency relationship of the query graph, constructing the head-tail mapping of the relationship, mask calculation and unique hot encoding, etc., to understand the overall structure and relationship in the query. In the face of the difficulty of embedding-based methods in capturing multiple dependencies in complex queries, the present invention processes the query statement, and through structured prompts and geometric operation modeling, multi-hop reasoning can be effectively performed, and a more comprehensive knowledge graph is constructed according to the context information of each node in the query structure, thereby greatly improving the accuracy of reasoning. Finally, the final answer is predicted by calculating the similarity between the candidate entity and the query, so as to find the most reasonable answer. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1 It is the overall flow chart of the present invention;

[0094] Figure 2 This is a network structure diagram of the pre-trained language model BERT in the present invention;

[0095] Figure 3 A structure diagram for query structure coding in the present invention;

[0096] Figure 4 It is the overall structure diagram of the present invention;

[0097] Figure 5 This is a query structure context representation diagram in the present invention. DETAILED DESCRIPTION

[0098] The specific operation steps of a multi-hop knowledge graph reasoning method based on query structure encoding of the present invention will be described in more detail below with reference to the accompanying drawings.

[0099] The overall implementation process of the present invention mainly includes three parts, namely, a structure encoding module for query statements, a structure context module, and an embedding training reasoning module.

[0100] The present invention constructs a flow chart as shown Figure 1 As shown, the overall structure is Figure 4 As shown, each step will be described in detail below.

[0101] Step 1: According to the type of query statement, determine whether the query statement is a projection, intersection or union, select the appropriate linearization template, and convert the query structure into a text sequence. For query types consisting of intersection, union and projection, the last relational projection is performed on the intersection and union of the previous triples. Direct flattening queries cannot retain such structural information, so it is necessary to split the query structure and repeatedly connect each triple to the last relational projection, thereby moving the intersection and union operations to the last step.

[0102] Step 2: Introduce step-by-step instructions indicating the order in which geometric operations are performed to prompt the pre-trained encoder with implicit structural information of the query. Each hint consists of two parts: the total number of execution steps and the order of operations in the query, expressed as "total number of steps: order of operations". The structural hint and the linearized query structure are concatenated together as the input of the pre-trained language model.

[0103] Step 3: Analyze the position clues, functional roles and query types of each node in the encoded query structure. Use methods such as node adjacency and relationship head-tail mapping, and flatten the position and role encoding query table to obtain a vector, normalize the obtained vector, and use linear transformation to map this normalized vector to a continuous embedding space to create a type embedding, thereby capturing the structural context of a given query type.

[0104] Step 4: Extend the pre-trained encoder with additional attention layers and max-output layers to model different geometric operations such as projection, intersection, and union on the representation space, respectively, to implicitly inject structured modeling into BERT.

[0105] Step 5: Initialize feature vectors for each node and edge. In graph neural networks, nodes transmit information through adjacent nodes. The feature vector of each node is updated by aggregating information from other adjacent nodes. After continuous message transmission and updating, the feature vector of each node is updated to the embedded representation of the node in the knowledge graph. Then, the similarity between the query node and the candidate entity node is calculated to further narrow the answer range.

[0106] Step 6: Use cross entropy loss to handle the prediction tasks of entities and relationships, assign weighted coefficients to the loss function, and find the most suitable coefficient combination by continuously optimizing the loss function to improve the reasoning accuracy of the model.

[0107] Further, the data processing is described as follows:

[0108] (1) Data preservation: Use the json.dump() method to serialize different types of data lists into JSON format and write them to a file, and encapsulate the data in a unified structure. When calling the json.dump() method, set the indentation level to 2 so that the generated JSON file can have good readability.

[0109] (2) Data reading: Use functions to receive the name and type of the dataset, determine whether to add joint data, and whether to process other types of data. Use the pickle module to load entity mappings and relationship mappings from mapping files. Initialize an empty dictionary to store the mapping between entities and their text descriptions.

[0110] (3) Get entity text: Set the directory of text data and initialize a dictionary to store entities and their corresponding text. Read the entity file line by line. Each line contains an entity and its corresponding text. Remove extra spaces to separate the entity and text, and store the mapping between entity and name in the dictionary. Get the entity name by traversing all entity indexes, get the corresponding text from the dictionary, store it in the dictionary, and save it.

[0111] (4) Get relation text: Use the given relation to get the corresponding text description. If the relation does not exist in the dictionary, try to find the inverse relation of the relation to return the text description of the relation. If dealing with more complex relation strings, first remove the period in the relation string through the function, split the relation string into multiple parts by slash, remove the main identifier of the relation from the split parts, and extract the rest. For the remaining part, split it by underscore to obtain more refined vocabulary. Finally, add all non-empty words to the list, and connect the words in the list with spaces to form a complete text description.

[0112] (5) Query structure encoding: By encoding the query structure, the input query text is converted into a format that the BERT model can understand. According to different query types, a query string of a specific format is constructed. The query string will start with [MASK], followed by the relation marker [rela] and the content of the query text. After the query string is constructed, it is converted into an input ID and an attention mask through a tokenizer, thereby ensuring that the input data meets the requirements of the model and can be correctly recognized by the model. In addition, by generating a position ID, the BERT model can know the position information of each word in the input sequence. Finally, the input ID, attention mask and position ID are encapsulated into a dictionary and returned to the model for training and inference.

[0113] Furthermore, the query structure encoding is described as follows:

[0114] (1) Call BERT: By calling the function, the input token ID and other optional parameters are passed into the BERT model. During this process, by executing the forward propagation of the BERT model, the model will calculate the representation of each token based on the input token ID and other information. The BERT model converts the input token ID into an embedding vector, and combines the position encoding and token type encoding to generate the input representation. Since BERT is composed of multiple stacked Transformer encoder layers, each layer performs self-attention calculations and feedforward neural network processing on the input to capture contextual information. Finally, the BERT model outputs the representation of each token, which contains contextual information.

[0115] (2) Model selection: BERT is a Transformer-based pre-trained language model with strong language understanding capabilities and rich pre-trained knowledge, which enables it to excel in capturing text semantics and adapting to unseen entities. Its efficient training and reasoning performance, strong generalization ability, superior structured knowledge modeling capabilities, and high robustness make BERT perform well in various natural language processing tasks. Compared with other models, BERT not only has higher performance, but also can better handle dynamically changing knowledge graphs, providing great practicality and benefits for research on knowledge graph reasoning. Its structure diagram is shown in the figure below. Figure 2 shown.

[0116] (3) Pooled output: The output of the last hidden layer is obtained through the BERT model, and then the output is averaged and pooled to obtain the pooled output representing the entire input sequence. Average pooling effectively integrates the information of all tokens in the sequence and captures the overall semantics of the input text. The calculation process is expressed as:

[0117]

[0118] in, is the length of the input sequence, i.e. sequence_length, Indicates The hidden state of tokens has a shape of (batch_size, hidden_size).

[0119] Furthermore, the training inference model is described as follows:

[0120] (1) Model training: The model parameters are divided into parameters that require weight decay and parameters that do not require weight decay. This method is used to control regularization during the optimization process. The classified parameters, learning rate, epsilon and other parameters are passed in through the Adam optimizer.

[0121] Initialize variables to track training progress, including global steps, batch counts, best scores, and best rounds. Record loss values, learning rates, and other information during training, and perform evaluations after reaching a specific number of steps or rounds, while recording the current training loss, so as to monitor the training process.

[0122] (2) Data processing: Read and clean the data, segment and encode the text, and convert the text into the format required for model input. During the training process, the data set is divided into small batches, and multi-threading is used to load the data, and the data is placed on the GPU to accelerate calculation. During the evaluation phase, the data processing logic will be called to ensure that each batch of data can be processed and formatted.

[0123] (3) Query structure linearization: Select an appropriate linearization method based on the query type. For each triple in the query, it is represented as [anchor]t(e i )[relation]t(r i ). The structural hint not only inputs the linearized query structure into the pre-trained language model PLM, but also introduces step-by-step instructions indicating the order in which geometric operations are performed. The instructions include two parts: the total number of execution steps and the order of operations in the query. The structural hint s and the linearized query structure t(q) are connected to form an input format [CLS][qtype]s[SEP]t(q) which is input into the training encoder to obtain the output hidden state H = (h1, h1, ..., h |t| ).

[0124] Among them, [CLS] is a special marker used to indicate the beginning of the input sequence, [qtype] is used to indicate the type or category of the query, and [SEP] stands for "separator" and is used to separate different parts of the input sequence. The model diagram is as follows Figure 3 shown.

[0125] (4) Model evaluation: When testing the performance of a trained model, use an evaluation dataset. This is done by setting the model to evaluation mode, disabling gradient calculation, processing the evaluation data batch by batch, and calculating the corresponding evaluation metrics. Finally, the results are recorded and output.

[0126] Furthermore, the reasoning task is described as follows:

[0127] (1) Central intersection layer: A neural network layer is defined to process the embedded representation and generate weighted output. The input is transformed through two linear layers, and the weights of the linear layers are initialized using Xavier uniform distribution to promote training stability and convergence speed.

[0128] (2) Forward propagation method: The input is linearly transformed through two linear layers, and the ReLU activation function and Softmax function are applied to obtain the attention weights, and the sum of the attention weights is ensured to be 1. All embeddings are weighted and summed to generate a new embedding representation, and finally the weighted embedding representation is returned.

[0129] (3) Xavier: Ensure that the variance of the signal can be kept within a reasonable range during forward propagation and back propagation, thereby avoiding the problem of gradient vanishing or gradient exploding. Xavier initialization sets the initial value of the weight by considering the number of inputs and outputs, so that the output variance of each layer is close to the input variance. Through reasonable weight initialization, the model can converge faster in the early stage of training and reduce training time. The calculation process is expressed as:

[0130]

[0131] Among them, fan in is the number of input units, fan out is the number of output units, and the weights are initialized to be uniformly distributed in the range [-a, a].

[0132] Further, the context description for the query structure is as follows:

[0133] (1) Embedding dimension: Define the embedding and initialize it with a uniform distribution. Dynamically set the dimension of the embedding according to the specific configuration of the model. Combine the embedding with other features through matrix multiplication in the forward propagation to generate the final embedding representation.

[0134] (2) Query graph adjacency relations: Different types of query graph structures are defined by creating dictionaries. Each structure represents a specific relationship pattern. The key of the dictionary is the type of query graph, and the corresponding value is a dictionary containing input and output relations.

[0135] (3) Construct relationship head-tail mapping: Determine whether the entity and relationship indexes need to be remapped based on the name of the data set, and load the corresponding entity and relationship mapping files. Obtain the triples of the knowledge graph by reading the training data set, and construct relationship-to-index and entity-to-index mapping dictionaries, and count the number of relationships and entities. By traversing the training data, extract the head entity, relationship, and tail entity, perform index mapping based on the mapping relationship, add the head entity to the head entity set of the corresponding relationship, add the tail entity to the tail entity set, calculate the frequency of the head entity and tail entity of each relationship, and find the maximum frequency and median frequency, so as to better understand the distribution of the relationship. The calculation process is expressed as:

[0136] F h (r i)=count({h|(h,r i , t)∈triples})

[0137] F t (r i )=count({t|(h,r i , t)∈triples})

[0138] F h,max =max(F h )

[0139] F t,max =max(F t )

[0140] F h,median =median(F h )

[0141] F t,median =median(F t )

[0142] Among them, triples represents all triples in the knowledge graph, h is the head entity, and t is the tail entity.

[0143] (4) Mask calculation: Mask calculation is used to indicate which entities are valid and which are filled, so as to achieve the purpose of correctly processing input data during training. There are two types of mask calculation: without sampling and with sampling. Without sampling, a mapping is constructed based on the number of head entities and tail entities of each relationship, and it is filled to the maximum frequency. With sampling, the head entity and tail entity of each relationship are randomly selected to ensure that the number does not exceed the specified sampling number, and the insufficient part is filled with a special value.

[0144] (5) Query2box: The entities and relations of the encoded query structure are embedded into a collection space, and geometric shapes are used to represent and reason about complex relationships in the knowledge graph. The knowledge graph embedding method is adopted to enhance the model's ability to handle complex queries, thereby improving the accuracy of reasoning.

[0145] (6) Normalization of feature matrices: Traverse each feature matrix in the query type dictionary and calculate the sum of the elements of each matrix. Find the largest element from the list, which is the maximum value in the sum of the feature matrices. Then normalize the feature matrix and fix the element values ​​of all feature matrices between 0 and 1. The normalization operation can prevent certain features from dominating the model training, resulting in a decrease in model performance. At the same time, it can accelerate the convergence speed of the model and improve the efficiency and effect of subsequent model training. The calculation process is expressed as:

[0146]

[0147]

[0148] Among them, query2table[qtype] is the feature matrix of each query type, query2table_norm[qtype] is the normalized feature matrix, To sum the feature matrices of each query type, find the maximum value of the sum of all feature matrices. The model diagram is as follows Figure 5 shown.

[0149] (7) One-hot encoding: By traversing each query type, transposing the dimension of the feature matrix of each query type, flattening the transposed matrix into a one-dimensional tensor, and finally moving it to the specified device. One-hot encoding converts the feature matrix of the query type into a format suitable for machine learning model input, so as to better capture the relationship between features.

[0150] It should be noted that the above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art of the present invention, several equivalent substitutions or obvious variations can be made without departing from the concept of the present invention, and the performance or use is the same, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A multi-hop knowledge graph reasoning method based on query structure encoding, characterized in that: Here are the steps: Step 1: According to the type of query statement, select a suitable linearization template to convert the query structure into a text sequence, thereby retaining the structural information of the query statement; Step 2: Design structured hints for each query type to indicate the order in which geometric operations are performed in the query, thereby helping the pre-trained language model capture the structural information of the query; Step 3: Use position embedding, role embedding, and type embedding to learn structural context and fuse it with the model; Step 4: Use the pre-trained encoder, attention layer, and maximum layer to model the three geometric operations of projection, intersection, and union in the representation space respectively; Step 5: Use graph neural network to encode query and candidate entities respectively and calculate the similarity between them; Step 6: Optimize the contrast loss and classification loss for training. During inference, use the query structure that incorporates context information for query reasoning.

2. The multi-hop knowledge graph reasoning method based on query structure encoding as claimed in claim 1, characterized in that: For steps 1 and 2, for each triple in the query, denote it as [anchor]t(ei)[relation]t(r i ), introduces step-by-step instructions that indicate the order in which geometric operations are performed. The instructions include two parts: the total number of execution steps and the order of operations in the query. The structure hint s and the linearized query structure t(q) are connected to form the input format [CLS][qtype]s[SEP]t(q) and input to the training encoder.

3. The multi-hop knowledge graph reasoning method based on query structure encoding as claimed in claim 1, characterized in that: For step 3, an adjacency matrix is ​​constructed, where the rows and columns correspond to the nodes in the query graph. If there is an edge between node i and node j, the element in the i-th row and j-th column of the matrix is ​​1, otherwise it is 0, indicating that there is no edge between node i and node j. The adjacency matrix can effectively represent the connection relationship between nodes, thereby obtaining the set of neighbor nodes of a node. The calculation process is: N(i)={j|A(i,j)} Among them, A(i,j) represents the element in the i-th row and j-th column of the adjacency matrix. Using the head-to-tail mapping method, a mapping table is established. The key of the table is the relationship number, and the value is a set of two tuples. Each node is assigned a position number and a role number. The first number of the tuple of each node in the query graph represents the position number, and the second number represents the role number. The calculation process is: Rout(i)={r|(i,r,j)∈G} Rin(i)={r|(j,r,i)∈G} Among them, Rout(i) is the set of outgoing edge relations, Rin(i) is the set of incoming edge relations, G is the query graph, and (i, r, j) is a triple in the query graph. After that, the position and role encoding query table is flattened to obtain a vector, and the obtained vector is normalized. A linear transformation is used to map this normalized vector to a continuous embedding space to create a type embedding, thereby capturing the structural context of a given query type.

4. The multi-hop knowledge graph reasoning method based on query structure encoding as claimed in claim 1, characterized in that: For step 4, the pre-trained encoder is extended with additional attention layers and maximum output layers to model different geometric operations such as projection, intersection, and union on the representation space, respectively, to implicitly inject structured modeling into BERT. The SILR model models different geometric operations on the representation space including: (1) Projection operation: For queries that only contain relational projection operations, the average of the representations of all target entities is taken as the query representation. The calculation process is: Where n is the number of target entities, h(v i ) is the representation of the i-th target entity. (2) Intersection operation: For queries containing intersection operations, an attention layer is used to model the target entities in multiple subqueries. The input of the attention layer is the representation of the target entity in all subqueries, and the output is the weighted average of the representation of the target entity in each subquery. The calculation process is: h{v i } w =α i h{v i } Among them, α i is the attention weight, h{v i } is the representation of the target entity in the i-th subquery. (3) Union operation: For queries that contain union operations, the maximum layer is used to model the target entities in multiple subqueries. The input of the maximum layer is the representation of the target entity in all subqueries, and the output is the maximum value of all target entity representations. The calculation process is: h q =max(h{v i }) Among them, h q This is the query representation after the union operation.

5. The multi-hop knowledge graph reasoning method based on query structure encoding as claimed in claim 1, characterized in that: For step 5, the graph neural network transfers information through adjacent nodes. The feature vector of each node is updated by aggregating information from adjacent nodes. After repeated information transfer and updating, the feature vector of each node is updated to the embedded representation of the node in the knowledge graph, and then the similarity between the query node and the candidate entity node is calculated to further narrow the answer range.

6. The multi-hop knowledge graph reasoning method based on query structure encoding as claimed in claim 2, characterized in that: The query structure encoding includes: (1) Call BERT: By calling the function, the input token ID and other optional parameters are passed into the BERT model. During this process, by executing the forward propagation of the BERT model, the model will calculate the representation of each token based on the input token ID and other information. The BERT model converts the input token ID into an embedding vector, and combines the position encoding and token type encoding to generate the input representation. Since BERT is composed of multiple stacked Transformer encoder layers, each layer performs self-attention calculations and feedforward neural network processing on the input to capture contextual information. Finally, the BERT model outputs the representation of each token, which contains contextual information. (2) Pooled output: The output of the last hidden layer is obtained through the BERT model, and then the output is averaged and pooled to obtain the pooled output representing the entire input sequence. Average pooling effectively integrates the information of all tokens in the sequence and captures the overall semantics of the input text. The calculation process is expressed as: in, is the length of the input sequence, i.e., sequence_length, output :,i,: Represents the hidden state of the i-th token, with a shape of (batch_size, hidden_size).

7. The multi-hop knowledge graph reasoning method based on query structure encoding as claimed in claim 3, characterized in that: The context of the query structure includes: (1) Embedding dimension: Define the embedding and initialize it with a uniform distribution. Dynamically set the dimension of the embedding according to the specific configuration of the model. Combine the embedding with other features through matrix multiplication in the forward propagation to generate the final embedding representation. (2) Query graph adjacency: Different types of query graph structures are defined by creating dictionaries. The following are the definitions and execution standards of 14 query types: 1p (Path-1): There is a path in the query graph with a path length of 1, indicating that there is a relationship connecting the anchor node and the answer node; 2p (Path-2): There is a path in the query graph with a path length of 2, indicating that there are two relations connecting the anchor node and the answer node; 3p (Path-3): There is a path in the query graph with a path length of 3, indicating that there are three relationships connecting the anchor node and the answer node; pi (Path-1 with Inverse): There is a path in the query graph with a path length of 1, and one of the relations is an inverse relation; ip(Inverse-Path-1): There is a path in the query graph with a path length of 1 and the relationship is an inverse relationship; 2i (Path-2 with Inverse): There is a path in the query graph with a path length of 2, and one of the relations is an inverse relation; 3i (Path-3 with Inverse): There is a path in the query graph with a path length of 3, and one of the relations is an inverse relation; pin(Path-1with Intersection): There are two paths in the query graph, both with a path length of 1, and these two paths share an answer node; pni (Path-1 with Non-Intersection): There are two paths in the query graph, both with a path length of 1, and these two paths do not share the answer node; 2in (Path-2 with Intersection): There are two paths in the query graph, both with a path length of 2, and these two paths share an answer node; 3in (Path-3 with Intersection): There are two paths in the query graph, both with a path length of 3, and these two paths share an answer node; 2u (Path-2 with Union): There are two paths in the query graph, both with a path length of 2, and the union of the answer node sets of these two paths is the final answer node set; 3u (Path-3 with Union): There are two paths in the query graph, both of which have a length of 3, and the union of the answer node sets of these two paths is the final answer node set; up(Union with Path): There is a path and an answer node set in the query graph. The union of the path's answer node set and the answer node set is the final answer node set. Each structure represents a specific relational pattern. The key of the dictionary is the type of the query graph, and the corresponding value is a dictionary containing the input and output relations. (3) Construct relationship head-tail mapping: Determine whether the entity and relationship indexes need to be remapped based on the name of the data set, and load the corresponding entity and relationship mapping files. Obtain the triples of the knowledge graph by reading the training data set, and construct relationship-to-index and entity-to-index mapping dictionaries, and count the number of relationships and entities. By traversing the training data, extract the head entity, relationship, and tail entity, perform index mapping based on the mapping relationship, add the head entity to the head entity set of the corresponding relationship, add the tail entity to the tail entity set, calculate the frequency of the head entity and tail entity of each relationship, and find the maximum frequency and median frequency, so as to better understand the distribution of the relationship. The calculation process is expressed as: F h (r i )=count({h|(h,r i ,t)∈triples}) F t (r i )=count({t|(h,r i ,t)∈triples}) F h,max =max(F h ) F t,max =max(F t ) F h,median =median(F h ) F t,median =median(F t ) Among them, triples represents all triples in the knowledge graph, is the head entity, and is the tail entity. (4) Mask calculation: Mask calculation is used to indicate which entities are valid and which are filled, so as to achieve the purpose of correctly processing input data during training. There are two types of mask calculation: without sampling and with sampling. Without sampling, a mapping is constructed based on the number of head entities and tail entities of each relationship, and it is filled to the maximum frequency. With sampling, the head entity and tail entity of each relationship are randomly selected to ensure that the number does not exceed the specified sampling number, and the insufficient part is filled with a special value. (5) Query2box: The entities and relations of the encoded query structure are embedded into a collection space, and geometric shapes are used to represent and reason about complex relationships in the knowledge graph. The knowledge graph embedding method is adopted to enhance the model's ability to handle complex queries, thereby improving the accuracy of reasoning. (6) Normalization of feature matrices: Traverse each feature matrix in the query type dictionary and calculate the sum of the elements of each matrix. Find the largest element from the list, which is the maximum value in the sum of the feature matrices. Then normalize the feature matrix and fix the element values ​​of all feature matrices between 0 and 1. The normalization operation can prevent certain features from dominating the model training, resulting in a decrease in model performance. At the same time, it can accelerate the convergence speed of the model and improve the efficiency and effect of subsequent model training. The calculation process is expressed as: Among them, query2table[qtype] is the feature matrix of each query type, query2table_norm[qtype] is the normalized feature matrix, ∑ i,j query2table[qtype] i,j To sum the feature matrices for each query type, find the maximum value in the sum of all feature matrices. (7) One-hot encoding: By traversing each query type, transposing the dimension of the feature matrix of each query type, flattening the transposed matrix into a one-dimensional tensor, and finally moving it to the specified device. One-hot encoding converts the feature matrix of the query type into a format suitable for machine learning model input, so as to better capture the relationship between features.

8. The multi-hop knowledge graph reasoning method based on query structure encoding as claimed in claim 4, characterized in that: The BERT model includes: (1) Transformer model: Apply dropout operation to prevent overfitting, project features into output space, generate logits, and then process logits. If the last dimension of logits is 1, indicating that there is only one output, this dimension is removed by squeezing for subsequent processing. Finally, the processed logits are returned for classification. The representation of candidate entities is calculated, the query sequence is encoded, and the mask representation is extracted from it. The calculated logits are converted into probability scores through softmax. (2) Masked Language Model: Randomly mask some words in a sentence and then predict these masked words, so that the model learns to infer the missing words from the context. (3) Next Sentence Prediction: By training the model to determine whether two given sentences are continuous in the original text, it helps the model learn how to capture the logical relationship and context between sentences. (4) Euclidean distance calculation method: By inputting two vectors, calculating the difference between the two vectors and obtaining their L2 norm, the calculated Euclidean distance is returned, which represents the distance between the two input vectors. These two vectors are the feature representations of the query-matched entity and the candidate entity, respectively. During the training process, the Euclidean distance can be used to calculate the loss function to help the model optimize its parameters so that similar inputs are closer in the feature space. During the inference stage, the model may need to select one from multiple candidate answers, and the Euclidean distance can help determine which candidate answer is closest to the query representation. The calculation process is expressed as: distance = rep1 - rep2 Among them, rep1 and rep2 represent two feature vectors that need to be compared or calculated, and distance represents the distance measure between two feature representations (or vectors), which is obtained by calculating the Euclidean distance between rep1 and rep2, indicating the similarity or difference between the two vectors in the feature space. (5) Cross entropy calculation method: After the softmax function is processed, the probability distribution is obtained. The predicted probability corresponding to each sample is obtained through the index. The natural logarithm of these probabilities is calculated. In order to avoid the case of zero logarithm, a constant 1e-5 is added, and the log likelihood is obtained by taking a negative value. Use torch.sum to calculate the sum of the log likelihoods of all samples. Then calculate the average loss by the number of samples. The calculation process is expressed as: Among them, m is the number of samples and is the predicted probability corresponding to the true category of the jth sample.

9. The multi-hop knowledge graph reasoning method based on query structure encoding as claimed in claim 5, characterized in that Graph neural networks include: (1) Graph Isomorphic Network: Through a powerful node feature update mechanism, nodes can capture local structural information in the graph and distinguish different graph isomorphisms. (2) Message passing mechanism: Through an aggregation-based message passing mechanism, the representation of a node is updated by aggregating the features of its neighbors. The calculation process is expressed as: in Represents the characteristics of the nodes in the layer, Represents the set of neighbor nodes of a node, MLP (l) It is a multi-layer perceptron layer used for feature updating. (3) Feature update: The graph isomorphism network uses a multi-layer perceptron as a feature update function to directly concatenate the aggregated features with the node's own features, and then undergoes a nonlinear transformation, so that the model can better learn and distinguish structural information. The calculation process is expressed as: in It is the sum of the features of all neighboring nodes of a node. (4) Graph-level representation: Aggregate and summarize the entire entity and relationship information into a single vector or feature representation. By extracting global information from complex graph structures, high-level analysis and prediction can be performed.

Citation Information

Cited By

  • Coal rock knowledge graph-oriented relation perception gating neural network link prediction method

    CN120781945A

  • Maritime photovoltaic information management system based on knowledge graph

    CN122286701A