Lightweight knowledge reasoning method for large knowledge graph, electronic device and computer readable medium
Patent Information
- Application Number
- CN202510365914.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]本发明提供一种面向大型知识图谱的轻量级知识推理方法、电子设备及计算机可读介质,以解决了现有技术中计算开销过高和推理效率低下的问题
[0032]1.本发明通过倒排索引技术快速提取目标子图,显著减少了大规模知识图谱的计算复杂度。结合多层知识蒸馏与动态int8量化方法,模型参数量和内存占用大幅降低,适用于资源受限的边缘计算场景。
Smart Images

Figure CN122840211A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of natural language processing and information security technology, specifically to a lightweight knowledge reasoning method, electronic device, and computer-readable medium for large knowledge graphs. Background Technology
[0002] Knowledge reasoning aims to infer new or potential facts from existing facts or knowledge, thereby improving the knowledge system or assisting decision-making. In the field of cybersecurity, with the increasing diversification and complexity of cyberattacks, the collection and management of large-scale security data has become crucial. By constructing knowledge graphs from threat intelligence, vulnerability information, attack samples, and log data, knowledge reasoning can help identify potential attack chains, predict threat intelligence, and assist in security decision-making. However, knowledge reasoning methods based on graph neural networks (GNNs) still suffer from high computational overhead, data noise and uncertainty, insufficient integration of domain knowledge, and difficulty in balancing real-time performance and scalability in practical applications. Summary of the Invention
[0003] This invention provides a lightweight knowledge reasoning method, electronic device, and computer-readable medium for large knowledge graphs, which solves the problems of high computational overhead and low reasoning efficiency in the prior art.
[0004] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0005] Design a lightweight knowledge reasoning method for large knowledge graphs, including the following steps:
[0006] Step S1: Obtain the knowledge graph K = {(h,r,t)|h,t∈ε,r∈R}, where ε is the entity set, R is the relation set, and the triple (h,r,t) represents the head entity h connected to the tail entity t through relation r; input the target entity e = {w1,w2,w3,...w n}, where w i For each character in the target entity string, n is the length of the target entity string. Using the multilevel alphabetic indexing (MLI), find the inverted index I(e) = {(h,r,t)∈K|h=e∪t=e} that stores all triples directly connected to e. Then, extract the triples (e,r,e...). o ) or (e o Transform the graph structure (r, e) into a graph structure, with the target entity e as the central node, the relations in the triples as edges, and the other entity e o Construct the target subgraph G for the other nodes connected to e. e = (V, E), where V is all nodes in the subgraph and E is all edges in the subgraph;
[0007] Step S2: Initialize the target subgraph G e Embedding of each node and edge in the knowledge graph K n Embedding r , target subgraph G e And the node list V, edge list E, and embedding of the knowledge graph K. n and Embedding r The contiguous matrix M is input into the GAT model to obtain the student model θ1 and the teacher model θ2, respectively.
[0008] Step S3: Define the scoring function To define the loss function At the same time, the scoring function is also used for the final prediction, where, e s and r q Represent the embeddings of the given query entity and relation, respectively, e t For candidate entity embedding, v e and v r b represents the entity weight matrix and the relation weight matrix, respectively. c b p These are the combined bias and the mapped bias, respectively; y is a binary vector, y i =1 indicates that the i-th candidate entity is a positive label, otherwise it is a negative label;
[0009] Step S4: Perform multi-layer knowledge distillation on the embedding layer, hidden layer, attention matrix, and prediction layer of the student model θ1 to achieve the goal of the student model θ1 learning the knowledge of the teacher model θ2. The loss of each layer of the student model θ1 is L. embedding L attention L hidden L prediction Based on the Loss function defined in step S3 and the loss functions of each layer mentioned above, construct the objective function for the student model θ1:
[0010]
[0011] Where m represents the number of layers in the model.
[0012] Furthermore, step S5 involves using a dynamic int8 quantization method to quantize the original float32 model parameters into integers q = round(sx + z), where s is the scaling factor and λ is the zero-point offset, in order to reduce model storage requirements and improve inference efficiency.
[0013] Furthermore, in step S1, the multi-level alphabetic index (MLI) uses a 4-level index to match the first 4 characters of the entity string, in order to save storage space.
[0014] Furthermore, step S2 includes: step S2.1, initializing the target subgraph G using TransE as a pre-embedded model. e Embedding of each node and edge n Embedding r , will G e = (V, E) Node list V, edge list E, embedding n and Embedding r The contiguous matrix M is input into the GAT model;
[0015] Step S2.2: Student model θ1 and teacher model θ2 learn the node representation through an attention mechanism. The representation of the i-th node at layer l+1 is:
[0016]
[0017] Among them, M i Let σ(·) represent the set of all nodes connected to i, and let ReLU = max(0,·) be the activation function. Let g(·) represent the information updated from the i-th layer. j w represents the attention score between node i and node j. a These are the parameters that need to be trained.
[0018] Furthermore, step S3 includes: step S3.1, performing an entity prediction task (e s ,r q After learning through an attention mechanism, using a given entity and relation e s r q For each candidate entity e t Rank the results and obtain the scoring function.
[0019] Furthermore, step S4 includes: S4.1, after the student model θ1 and teacher model θ2 have undergone the above steps, multi-layer knowledge distillation is performed on the embedding layer, hidden layer, attention matrix, and prediction layer of the student model θ1, and the loss of each layer is:
[0020] L embedding =MSE(E s W e E T )
[0021] L attention =MSE(A s A T )
[0022] L hidden =MSE(H s W hH T )
[0023] L prediciton =MSE(Z) T Z S )
[0024] Where MSE() represents the mean squared error, E S and E T W represents the embedding matrix of the embedding layers for θ1 and θ2. e It is the learnable weight matrix of the embedding layer; A S and A T H represents the attention weight matrix for θ1 and θ2; S and H T W represents the embedding matrix of the hidden layers of θ1 and θ2. h Z is the learnable weight matrix of the hidden layer; T and Z S This represents the logistic vector generated by the teacher model and the student model in the prediction layer.
[0025] Furthermore, in step S5, the trained student model θ1 quantizes the original float32 model parameters into integers q, using the following quantization formula:
[0026]
[0027] q ii =(w ii / scale i )+zero_point i
[0028] Among them, w i ={w i1 ,w i2 ....,w in} represents the parameter of the i-th layer, zero_polint i Zero-point offset, scale i This is the scaling factor.
[0029] A second aspect of the present invention provides an electronic device comprising: at least one processor; and a memory storing at least one program, wherein when the at least one program is executed by the at least one processor, it implements any of the above-described lightweight knowledge reasoning methods for large knowledge graphs.
[0030] A third aspect of the present invention is to provide a computer-readable medium on which a computer program is stored, wherein the computer program, when executed by a processor, implements any of the above-described lightweight knowledge reasoning methods for large knowledge graphs.
[0031] Compared with the prior art, the beneficial technical effects of the present invention are as follows:
[0032] 1. This invention rapidly extracts target subgraphs using inverted indexing technology, significantly reducing the computational complexity of large-scale knowledge graphs. Combined with multi-layer knowledge distillation and dynamic int8 quantization, the number of model parameters and memory usage are greatly reduced, making it suitable for resource-constrained edge computing scenarios.
[0033] 2. This invention addresses the trade-off between real-time performance and scalability faced by traditional methods when processing ultra-large-scale knowledge graphs through lightweight design and an efficient inference mechanism. The model can rapidly respond to dynamically updated knowledge graph data while maintaining inference performance, making it suitable for scenarios with high real-time requirements, such as cybersecurity threat analysis and vulnerability correlation inference.
[0034] 3. This invention is not only applicable to the field of information security, but can also be extended to other application scenarios that require large-scale knowledge reasoning, such as intelligent question answering, recommendation systems, and medical diagnosis, providing an efficient and lightweight solution for knowledge graph applications in various fields.
[0035] 4. This invention is applicable to large-scale knowledge reasoning tasks in the field of network security, and can efficiently process complex data such as threat intelligence, vulnerability information, and attack samples. Through a lightweight reasoning model, this invention can quickly identify potential attack chains, predict threat intelligence, and assist in security decision-making, significantly improving the real-time performance and accuracy of network security protection. Simultaneously, its low resource consumption allows it to be deployed on edge computing devices, providing efficient and reliable technical support for network security monitoring and response. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the overall model of the present invention.
[0037] Figure 2 This is a schematic diagram of the overall process of the present invention.
[0038] Figure 3 This is a schematic diagram illustrating the process of finding the target entity-related triples in this invention.
[0039] Figure 4 This diagram illustrates the comparison of memory usage and performance between the model in this invention and the GAT model. Detailed Implementation
[0040] The technical solution of the present invention will be described in detail below, but the scope of protection of the present invention is not limited to the embodiments described.
[0041] Example 1: With the widespread application of knowledge graphs in fields such as network security, intelligent question answering, and medical diagnosis, their scale and complexity are constantly increasing. Traditional graph neural network methods often face challenges of computational bottlenecks and excessive memory consumption when processing massive numbers of nodes and edges, making it difficult to meet the requirements of real-time performance and scalability. At the same time, too many model parameters and complex computational processes also make deployment in resource-constrained environments such as edge devices difficult, thus restricting the practical application of large-scale knowledge reasoning technology.
[0042] To address the aforementioned issues, the applicant proposes that for large-scale knowledge graphs, multi-level alphabetic indexing and inverted indexing techniques can be employed to find triples related to target entities and construct subgraphs to focus on the target entity and its directly associated local information, thereby significantly reducing the computational complexity and memory consumption during the overall processing of large-scale knowledge graphs. The retrieval process is as follows: Figure 3 As shown, the target entity is first input as a string into multiple threads. Each thread simultaneously searches for the "WordID" associated with the target entity, and then finds the triplet using the inverted index of the "WordID".<triple number,position> In this model, "triple number" represents the sequence number of the triple, and "position" represents the position of the target entity's word within the triple. A multi-layer knowledge distillation mechanism is also introduced, enabling the lightweight student model to mimic the key features of the teacher model at the embedding, attention, hidden, and prediction layers, resulting in a significant reduction in parameters and computational complexity. However, some semantic information may be lost during data extraction and model compression, leading to some limitations in inference accuracy when handling complex scenarios.
[0043] To further compensate for information loss and insufficient semantic detail, the applicant proposes using dynamic int8 quantization on the trained student model to represent model parameters in low-bit form. This compresses model parameters and reduces memory usage while maintaining inference performance, improving deployment efficiency and practical application effectiveness in resource-constrained environments. A comparison of the resource consumption and performance metrics of the proposed model with the classic GAT is shown below. Figure 4 As shown.
[0044] Based on the above concept, this invention designs a lightweight knowledge reasoning method for large-scale knowledge graphs, see [link to relevant documentation]. Figure 1 and Figure 2 This includes the following steps:
[0045] Step S1: Obtain the knowledge graph K = {(h,r,t)|h,t∈ε,r∈R}, where ε is the entity set, R is the relation set, and the triple (h,r,t) represents the head entity h connected to the tail entity t through relation r. Input target entity e = {w1,w2,w3,...wn}, where w i For each character in the string, and n as the length of the target entity string, the corresponding inverted index I(e) is found using Multi-Level Indices (MLI). To save storage space, MLI uses a 4-level index. The inverted index I(e) = {(h,r,t)∈K|h=e∪t=e} is found by matching the first 4 characters of the entity string. That is, the inverted index I(e) stores all triples directly connected to e. The extracted triples (e,r,e) are then... o ) or (e o Transform the graph structure (r, e) into a graph structure, with the target entity e as the central node, the relations in the triples as edges, and the other entity e o Construct the target subgraph G for the other nodes connected to e. e = (V, E), where V is all nodes (entities) in the subgraph and E is all edges (relationships) in the subgraph. This method of dividing a large-scale graph into several subgraphs reduces the computational burden.
[0046] Step S2: Combine the node list V, edge list E, and embedding of the target subgraph and the knowledge graph. n and Embedding r The adjoint matrix M is input into the GAT model to obtain the student model θ1 and the teacher model θ2, and the representation of learning node i in the student model θ1 and the teacher model θ2. Among them, M i Let σ(·) represent the set of all nodes connected to i, with the activation function ReLU = max(0,·), and g(·) represent the information update from the i-th layer. α j The attention score between node i and node j is represented by α. j This represents the correlation between node i and node j, where w a These are the parameters that need to be trained.
[0047] Step S2.1: Select TransE as the pre-embedded model and initialize the embedding of each node and edge in the subgraph. n Embedding r , will G e = (V, E) Node list V, Edge list E, Embedding n and Embedding r The contiguous matrix is input into the GAT model θ1.
[0048] Step S2.2: The model learns the node representation through an attention mechanism. The representation of the i-th node at layer l+1 is:
[0049]
[0050] Here, M i Let σ(·) represent the set of all nodes connected to i, and let ReLU = max(0,·) be the activation function. Let g(·) represent the information updated from the i-th layer. j w represents the attention score between node i and node j. a These are the parameters that need to be trained.
[0051] Step S3: Based on the scoring function Define loss function At the same time, the scoring function is also used for the final prediction.
[0052] Step S3.1: Assume the entity prediction task is (e s ,r q After learning through the attention mechanism, given an entity and relation e, ... s r q For each candidate entity e t To rank the players, the scoring function is:
[0053]
[0054] Among them, e s and r q Represent the embeddings of the given query entity and relation, respectively, e t For candidate entity embedding, v e and v r b represents the entity weight matrix and the relation weight matrix, respectively. c b p These are combined bias and mapped bias, respectively.
[0055] Step S3.2: Set the loss function Loss:
[0056]
[0057] Where y is a binary vector, y i =1 indicates that the i-th candidate entity is a positive label, otherwise it is a negative label.
[0058] Step S4: Perform multi-layer knowledge distillation in the embedding layer, hidden layer, attention matrix, and prediction layer of the teacher model and the academic model to achieve the goal of the student model learning the knowledge of the teacher model. The loss of each layer is L. embedding L attention L hidden L prediction Finally, the objective function of the student model is constructed based on the defined loss function and the loss functions of each layer.
[0059] Step S4.1: After the student model θ1 and teacher model θ2 have undergone the above steps, multi-layer knowledge distillation is performed on the embedding layer, hidden layer, attention matrix, and prediction layer of the student model θ1. The loss of each layer is:
[0060] L embedding =MSE(E s W e E T )
[0061] L attention =MSE(A s A T )
[0062] L hidden =MSE(H s W h H T )
[0063] L prediciton =MSE(Z) T Z S )
[0064] MSE() represents the mean squared error, E S and E T W represents the embedding matrix of the embedding layers for θ1 and θ2. e It is the learnable weight matrix of the embedding layer; A S and A T H represents the attention weight matrix for θ1 and θ2; S and H T W represents the embedding matrix of the hidden layers of θ1 and θ2. h Z is the learnable weight matrix of the hidden layer; T and Z S This represents the logistic vector generated by the teacher model and the student model in the prediction layer.
[0065] Step S4.2: Construct the loss function for the student model θ1 The loss function expression for training the student model is as follows:
[0066]
[0067] Step S5: Quantize the trained student model θ1 by converting the original float32 model parameters into integers q. This allows the quantized model to maintain high accuracy while significantly reducing storage requirements and computational complexity. The quantization formula is:
[0068]
[0069] q ii =(w ii / scale i )+zero_point i
[0070] Here, w i ={w i1 ,w i2 ....,w in} represents the parameter of the i-th layer, zero_polint i Zero offset serves to ensure that the zero value of a floating-point number is accurately mapped to the zero value of an integer, handle asymmetric distributions and nonlinear characteristics, and reduce quantization errors. i The scaling factor maps the dynamic range of floating-point numbers to the representation range of integers, controlling the precision and computational efficiency of quantization. Together, these factors allow the quantized model to maintain high precision while significantly reducing storage requirements and computational complexity.
[0071] Example 2: The present invention also provides an electronic device, including: at least one processor and a memory, wherein at least one program is stored in the memory, and when the program is executed by the processor, it implements the lightweight knowledge reasoning method for large knowledge graphs in Example 1.
[0072] Example 3: The present invention also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the lightweight knowledge reasoning method for large knowledge graphs in Example 1.
[0073] Currently, existing graph attention networks, when processing large-scale and complex knowledge graphs, perform computations directly on the entire graph, often resulting in slow training speeds, excessive memory consumption, and difficulty in meeting real-time requirements. Furthermore, their high resource dependence makes these technologies difficult to deploy efficiently in resource-constrained environments such as edge computing. This invention proposes a lightweight knowledge reasoning technique for large-scale knowledge graphs. First, it utilizes a reverse index query method to efficiently extract target subgraphs from large-scale knowledge graphs, significantly reducing the size of the data to be processed, thereby reducing computational complexity and memory burden. Second, it employs GAT for node representation learning on the target subgraph, and through a multi-layer knowledge distillation mechanism, enables the lightweight student model to effectively mimic the teacher model in the embedding, attention, hidden, and prediction layers, thus maintaining high inference performance while significantly reducing parameters. Finally, it combines dynamic int8 quantization technology to perform low-bit quantization of model parameters, further reducing memory usage and operating costs, and improving deployment efficiency in resource-constrained environments. In summary, this invention not only solves the problems of high computational cost and low inference efficiency in large-scale knowledge graph processing, but also takes into account the requirements of model accuracy and lightweight design, making it suitable for various practical application scenarios such as network security, intelligent question answering, and recommendation systems.
Claims
1. A lightweight knowledge reasoning method for large-scale knowledge graphs, characterized in that, Includes the following steps: Step S1: Obtain the knowledge graph K = {(h,r,t)|h,t∈ε,r∈R}, where ε is the entity set, R is the relation set, and the triple (h,r,t) represents the head entity h connected to the tail entity t through relation r; input the target entity e = {w1,w2,w3,...w n }, where w i For each character in the target entity string, n is the length of the target entity string. Using the multilevel alphabetic indexing (MLI), find the inverted index I(e) = {(h,r,t)∈K|h=e∪t=e} that stores all triples directly connected to e. Then, extract the triples (e,r,e...). o ) or (e o Transform the graph structure (r, e) into a graph structure, with the target entity e as the central node, the relations in the triples as edges, and the other entity e o Construct the target subgraph G for the other nodes connected to e. e = (V, E), where V is all nodes in the subgraph and E is all edges in the subgraph; Step S2: Initialize the target subgraph G e Embedding of each node and edge in the knowledge graph K n Embedding r , target subgraph G e And the node list V, edge list E, and embedding of the knowledge graph K. n and Embedding r The contiguous matrix M is input into the GAT model to obtain the student model θ1 and the teacher model θ2, respectively. Step S3: Define the scoring function To define the loss function At the same time, the scoring function is also used for the final prediction, where, e s and r q e represents the embedding of a given query entity and relation, respectively. t For candidate entity embedding, v e and v r b represents the entity weight matrix and the relation weight matrix, respectively. c b p These are the combined bias and the mapped bias, respectively; y is a binary vector, y i =1 indicates that the i-th candidate entity is a positive label, otherwise it is a negative label; Step S4: Perform multi-layer knowledge distillation on the embedding layer, hidden layer, attention matrix, and prediction layer of the student model θ1 to achieve the goal of the student model θ1 learning the knowledge of the teacher model θ2. The loss of each layer of the student model θ1 is L. embedding L attention L hidden L prediction Based on the Loss function defined in step S3 and the loss functions of each layer mentioned above, construct the objective function for the student model θ1: Where m represents the number of layers in the model.
2. The lightweight knowledge reasoning method for large-scale knowledge graphs according to claim 1, characterized in that, Also includes: Step S5: Use dynamic int8 quantization to quantize the original float32 model parameters into integers q = round(sx + z), where s is the scaling factor and z is the zero offset, in order to reduce model storage requirements and improve inference efficiency.
3. The lightweight knowledge reasoning method for large-scale knowledge graphs according to claim 1, characterized in that, In step S1, the multi-level alphabetic index (MLI) uses a 4-level index to match the first 4 characters of the entity string in order to save storage space.
4. The lightweight knowledge reasoning method for large-scale knowledge graphs according to claim 1, characterized in that, Step S2 includes: Step S2.1: Initialize the target subgraph G using TransE as the pre-embedded model. e Embedding of each node and edge n Embedding r , will G e = (V, E) Node list V, edge list E, embedding n and Embedding r The contiguous matrix M is input into the GAT model; Step S2.2: Student model θ1 and teacher model θ2 learn the node representation through an attention mechanism. The representation of the i-th node at layer l+1 is: Among them, M i Let σ(·) represent the set of all nodes connected to i, and let ReLU = max(0,·) be the activation function. Let g(·) represent the information updated from the i-th layer. j w represents the attention score between node i and node j. a These are the parameters that need to be trained.
5. The lightweight knowledge reasoning method for large-scale knowledge graphs according to claim 1, characterized in that, Step S3 includes: Step S3.1: Perform entity prediction task (e s ,r q After learning through an attention mechanism, using a given entity and relation e s r q For each candidate entity e t Rank the results and obtain the scoring function.
6. The lightweight knowledge reasoning method for large-scale knowledge graphs according to claim 1, characterized in that, Step S4 includes: S4.1 After the above steps are performed on both the student model θ1 and the teacher model θ2, multi-layer knowledge distillation is performed on the embedding layer, hidden layer, attention matrix, and prediction layer of the student model θ1. The loss of each layer is: L embedding =MSE(E s W e ,E T ) L attention =MSE(A s ,A T ) L hidden =MSE(H s W h ,H T ) L prediciton =MSE(Z T ,Z S ) Where MSE() represents the mean squared error, E S and E T W represents the embedding matrix of the embedding layers for θ1 and θ2. e It is the learnable weight matrix of the embedding layer; A S and A T H represents the attention weight matrix for θ1 and θ2; S and H T W represents the embedding matrix of the hidden layers of θ1 and θ2. h Z is the learnable weight matrix of the hidden layer; T and Z S This represents the logistic vector generated by the teacher model and the student model in the prediction layer.
7. The lightweight knowledge reasoning method for large-scale knowledge graphs according to claim 2, characterized in that, The quantized student model θ1 converts the original float32 model parameters into integers q. The quantization formula is as follows: q ii =(w ii / scale i )+zero_point i Among them, w i ={w i1 ,w i2 ....,w in } represents the parameter of the i-th layer, zero_polint i Zero-point offset, scale i This is the scaling factor.
8. An electronic device, comprising: At least one processor; A memory storing at least one program that, when executed by the at least one processor, implements the lightweight knowledge reasoning method for large knowledge graphs as described in any one of claims 1-7.
9. A computer-readable medium storing a computer program that, when executed by a processor, implements the lightweight knowledge reasoning method for large knowledge graphs as described in any one of claims 1-7.