Reinforced learning sparse graph reasoning method and system based on fusion reward
By adopting reinforcement learning methods based on fusion rewards in the dam emergency response system, the semantic information and graph structure information of the knowledge graph are integrated, and the rule-based reward shaping mechanism is introduced, the inconsistency or lengthy problems caused by sparse knowledge graphs are solved, and efficient and accurate emergency decision support is achieved.
Patent Information
- Application Number
- CN202510236390.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-02
AI Technical Summary
The sparse knowledge graph in the dam emergency response system results in disconnected or lengthy inference paths, affecting the accuracy and speed of decisions.
Reinforcement learning method based on fusion rewards is adopted to integrate semantic information and graph structure information by constructing fusion embedding modules, and a rule-based reward shaping mechanism is introduced to alleviate the problem of reward sparseness and optimize the reasoning path.
It significantly improves the accuracy and efficiency of sparse knowledge graph reasoning tasks, ensures the timeliness and interpretability of emergency decisions, and supports the applications of knowledge graph completion, disaster emergency decision-making and intelligent early warning.
Smart Images

Figure CN119918648A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sparse knowledge graph reasoning method and system based on fusion reward reinforcement learning, which belongs to the field of knowledge engineering and artificial intelligence technology and is mainly used for reasoning tasks in dam emergency response systems. The method can process sparse emergency response knowledge graphs and improve the accuracy and efficiency of emergency decision-making. Its application scope covers dam emergency response, knowledge graph completion, intelligent decision-making, disaster warning and other fields. Background Art
[0002] With the widespread application of knowledge graphs, their advantages in representing and organizing structured knowledge are gradually emerging, especially in dam emergency response systems, where knowledge graphs can help the system efficiently organize and reason about complex emergency response information. As a network data structure composed of entities and relationships, knowledge graphs can be used to characterize complex relationship networks in the real world, especially for scenarios such as dam monitoring and disaster prediction. However, since the construction process of knowledge graphs is usually limited by factors such as data acquisition, annotation costs, and domain coverage, the sparsity problem of knowledge graphs in dam emergency response systems is particularly significant. In sparse knowledge graphs, there are a large number of unlabeled or missing entity relationships, which directly affects the accuracy and coverage of emergency response decisions.
[0003] Existing knowledge reasoning methods can be mainly divided into three categories: rule-based methods, embedded representation-based methods, and reinforcement learning-based path reasoning methods. Rule-based reasoning methods use logical rules to perform knowledge reasoning, which has good interpretability and is suitable for scenarios with clear rules in dam emergency response systems (such as "water level exceeds the warning line → start the flood discharge port"). However, this type of method relies on expert knowledge to construct a rule set, resulting in insufficient rule coverage and poor generalization ability. Embedded representation-based reasoning methods map entities and relationships in knowledge graphs to low-dimensional vector spaces and use vector operations to capture potential semantic associations. They have good model performance, but lack interpretability of the reasoning process. In sparse knowledge graphs, their reasoning performance is severely limited due to insufficient training samples. Path reasoning methods based on reinforcement learning explore reasoning paths and learn optimal decision strategies through the interaction between agents and knowledge graphs. This method can generate interpretable reasoning paths, taking into account both reasoning performance and interpretability. However, in sparse knowledge graphs, agents face the challenge of insufficient information, resulting in a lack of positive feedback signals in the reasoning path, which seriously affects the learning efficiency and reasoning ability of the model. Especially in emergency response systems, sparse graphs may make the reasoning path disconnected or lengthy, resulting in affected decision-making speed and accuracy. Existing sparse knowledge graph reasoning methods have a significant impact on dam emergency response. The dam emergency response system needs to make decisions in a timely and accurate manner, while the missing information and incomplete relationships in the sparse knowledge graph will cause the reasoning path to be disconnected or wrong, thus affecting the speed and accuracy of decision-making. For example, if the graph lacks a clear relationship between "water level exceeding the warning line" and "activating the flood discharge port", it may delay the initiation of emergency measures, causing dam managers to miss the best time for emergency response, thereby affecting the safety of the dam and the stability of the surrounding environment.
[0004] Existing reasoning methods based on embedding representation have technical defects in dam emergency response, especially in the application of sparse knowledge graphs. Due to the sparsity problem, there are a large number of missing relations in the knowledge graph, which makes the embedding representation-based methods perform poorly in the reasoning process. The lack of sufficient training samples leads to a decrease in the accuracy and reliability of the reasoning results. In these cases, the system often cannot handle complex emergency scenarios, especially when the relationship chain is long or multiple relationships are inferred in parallel, which is prone to path breakage or lengthy paths, further affecting the efficiency of decision-making and the accuracy of emergency response. Summary of the invention
[0005] Purpose of the invention: In view of the above problems, the present invention provides a sparse graph reasoning method and system for reinforcement learning based on fused rewards, which can effectively utilize the semantic information of dam emergency response data in sparse knowledge graphs and alleviate the knowledge reasoning method with sparse rewards. The present invention utilizes a graph embedding module to fuse the semantic information and graph structure information in the sparse knowledge graph, and introduces a rule-based reward shaping mechanism to integrate the prior knowledge provided by the rule set into the reinforcement learning process, thereby effectively alleviating the problem of sparse rewards and improving the reasoning performance of the model. In the dam emergency response system, this method can combine real-time data and historical emergency response rules to generate efficient and explainable reasoning paths, providing theoretical and practical support for emergency decision-making.
[0006] Technical solution: A sparse graph reasoning method based on reinforcement learning with fusion rewards, by introducing a set of rules and a fusion reward mechanism, effectively alleviates the reward sparsity problem caused by insufficient information in sparse knowledge graphs, improves reasoning performance and efficiency, and ensures accurate and timely responses, especially in emergency response decisions. At the same time, a fusion embedding module is constructed to jointly represent the semantic association information and graph structure information in the dam emergency response knowledge graph, providing a more complete and accurate knowledge expression for the agent's reasoning, and capable of coping with complex emergency response scenarios. In addition, through a rule-based reward shaping mechanism, combined with terminal rewards to realize a fusion reward function, the reinforcement learning model is further guided to generate high-quality and explainable reasoning paths, ensuring that the generated decision path can be quickly understood and executed in emergency situations. This method significantly improves the accuracy and practicality of sparse knowledge graph reasoning tasks in the dam emergency response system, and can effectively support applications in fields such as knowledge graph completion, disaster emergency decision-making, and intelligent early warning.
[0007] The method comprises the following steps:
[0008] (1) Constructing a fusion embedding module: By combining the semantic association embedding representation based on TransE and the graph structure embedding representation based on GCN, a fusion embedding representation of the sparse knowledge graph in the dam emergency response system is generated. This module effectively integrates the semantic information (such as the relationship between "water level exceeds the warning line" and "starting the flood discharge port") and structural information (such as the connection relationship between various devices) in the dam emergency response knowledge graph, providing high-quality knowledge representation for the subsequent reasoning process, ensuring that the intelligent agent can make decisions based on complete information.
[0009] (2) Introducing rule sets: The AnyBURL method is used to summarize the sparse knowledge graph in the dam emergency response system and extract the logical rule set related to emergency response (such as "water level exceeds the warning line → start the backup pump station"). The support and confidence of the rules are evaluated to ensure the effectiveness and accuracy of the rules, provide rule-based prior knowledge for the agent reasoning, and optimize the reasoning path.
[0010] (3) Design of fusion reward function: A fusion reward function is constructed by combining rule-based rewards and terminal rewards. Rule-based rewards are calculated by the support and confidence of the rules, and terminal rewards are used to motivate the agent to correctly reach the target entity in the dam emergency response system (such as activating the flood discharge port, activating the backup pump station, etc.). The fusion reward function can effectively alleviate the reward sparsity problem in reinforcement learning, optimize the reasoning process, and ensure the accuracy and timeliness of emergency response.
[0011] (4) Path reasoning based on reinforcement learning: By using the interaction between the reinforcement learning agent and the dam emergency response knowledge graph environment, the optimal action is selected through the policy network, and the reasoning path is gradually generated to finally implement the emergency response measures. During the reasoning process, the agent will optimize according to the fusion reward signal, thereby improving the efficiency and accuracy of the reasoning task and ensuring the correct implementation of the emergency response measures in complex situations.
[0012] Furthermore, the specific steps of constructing the fusion embedding module in step (1) are as follows:
[0013] (1.1) Based on the TransE method, the semantic association information of the dam emergency response knowledge graph is embedded. For the relationship triple (h, r, t) in the dam emergency response system, the goal of TransE is to make the vector of the head entity plus the vector of the relationship equal to the vector of the tail entity as much as possible. The target formula is as follows:
[0014] h+r≈t
[0015] Among them, h, r, and t represent the embedding vectors of the head entity, relation, and tail entity, respectively.
[0016] The consistency of the triples is measured using the following scoring function:
[0017] Score(h,r,t)=-||h+rt|| p
[0018] where ||·|| p The scoring function uses the norm l1 or l2 to measure the distance between embedded vectors. By minimizing the scoring function, the semantic association embedding representation of entities and relations is learned, providing accurate semantic expression for reasoning tasks in the dam emergency response system.
[0019] (1.2) Based on the GCN method, the graph structure information of the dam emergency response knowledge graph is encoded. For the relation triple (h, r, t), the goal of GCN is to predict the representation of the tail entity t by aggregating the neighbor information of the head entity h and the representation of the relation r. The graph convolution layer forward propagation formula of the entity node update process is expressed as follows:
[0020]
[0021] where σ is the activation function, is the result of adding self-connection to matrix A, is the matrix of measurement, which is connected by the matrix The sum of each row of W is obtained. (l) It represents the weight matrix of the lth layer.
[0022] (1.3) Fusion of semantic association information and graph structure information: The semantic association information and graph structure information in the dam emergency response knowledge graph are extracted by TransE and GCN respectively. Based on the above operations, the semantic association representation and graph structure information of the graph are encoded to obtain their vector representations respectively. The two vectors are concatenated to obtain the concatenated vector E. cat , the specific formula is shown as follows:
[0023]
[0024] Among them, E TransE represents the entity semantic association information representation of the graph by the TransE model, E GCN Represents the entity graph structure information representation based on GCN.
[0025] Then the three vectors E TransE 、E GCN and E cat Through a fully connected layer, the weighted fusion method is used to map them into the same semantic space. The fused vector retains both the semantic association information and the graph structure information of the graph.
[0026] E=w1·E TransE +w2·E GCN +w3·E cat
[0027] Among them, E represents the entity feature vector embedded based on the fusion of image data information associated with the entity, w is the weight calculated by the fully connected layer, w1 corresponds to the weight of TransE embedding; w2 corresponds to the weight of GCN embedding; w3 corresponds to the weight of the concatenated embedding.
[0028] Furthermore, the specific steps of constructing the rule set in step (2) are as follows:
[0029] (2.1) Logical rules are mined from the dam emergency response knowledge graph using the AnyBURL method. Each rule in the rule set is formally represented as a Horn rule.
[0030] H r(X,Y)←b1(X,A1)∧...∧b n (A n ,Y)
[0031] The obtained Horn rule is represented by H, X, Y, A n Represent different entities respectively. The entity association in the rule is represented by b n (...) indicates that. For the constructed Horn rule H r (e i ,e j ), which is equivalent to the fact triple (e i ,r,e j ). Each Horn rule is used to describe the logical relationship between different entities in the dam emergency response system. For example, the rule can be expressed as "water level exceeds the warning line → start the flood discharge port", where the entities "water level exceeds the warning line" and "start the flood discharge port" are connected by a logical relationship. For the constructed Horn rule, it is equivalent to a fact triple, which is used to guide the emergency decision-making of the system.
[0032] (2.2) AnyBURL evaluates each rule in the rule set through the inference results in pre-training. The support of rule r indicates how many triples in a given knowledge graph the rule can match in the dam emergency response task. The formula is as follows:
[0033]
[0034] in, Indicates the number of correct triples that satisfy the rule.
[0035] (2.3) The confidence of the rule is evaluated and determined by calculating the ratio of the correct facts of the rule's head node prediction. The confidence of rule r is expressed as follows:
[0036]
[0037] in, Indicates the number of all triples matched by the rule. In the dam emergency response system, the confidence of the rule can be used to measure the reliability of the rule in the actual emergency response and ensure that the system makes decisions based on the rules with high confidence.
[0038] Furthermore, the specific steps of designing the fusion reward function in step (3) are as follows:
[0039] (3.1) Design a rule reward function and evaluate the quality of rules by rule support and confidence: First, design a rule reward function based on the rule set constructed by AnyBURL. The rule reward function evaluates the quality of rules based on the support and confidence of the rules. Specifically, the support measures the matching of the rules in the dam emergency response system, and the confidence reflects the predictive ability of the rules. Based on these two indicators, we can assign a corresponding reward value to each rule, so as to better guide the decision-making of the intelligent agent during the reasoning process.
[0040] R r (Q) = ∑ r∈Q support(r)×conf(r)
[0041] Among them, Q is the rule set constructed by AnyBURL, R r (Q) is the definition of the reward function corresponding to rule-based reward shaping.
[0042] (3.2) Definition of terminal reward: The terminal reward is given only when the agent reaches the target entity correctly, and the value is 1; otherwise the reward is 0. The terminal reward ensures that the agent's reasoning can focus on the correct target entity, thereby completing the correct emergency response. The formula is as follows:
[0043]
[0044] Among them, s T Indicates the end state of the search process, e s It represents the starting entity of the search process, r q Indicates the relationship of the query, e t Represents the target entity of the search process.
[0045] (3.3) The rule reward is calculated based on the support and confidence provided by the rule set, and finally the rule-based reward and hit reward are fused. The rule reward is calculated based on the support and confidence provided by the rule set. Then, the rule reward and the terminal reward are weighted and fused to form the final fusion reward function to optimize the learning process of the intelligent agent. The final fusion reward function is:
[0046] R total =λR r +(1-λ)R h
[0047] Among them, R h is the hit reward function after reward shaping, R r is a rule-based reward function, and λ is a reward balancing factor.
[0048] (3.4) Application of fusion rewards: Guided by rule rewards, the agent can more efficiently explore potential relationship paths in the dam emergency response knowledge graph, ensuring that the emergency response path is found faster. The introduction of terminal rewards ensures that the agent's path reasoning can focus on the correct target entity, thereby improving the accuracy and efficiency of reasoning and ensuring that the decision-making of the dam emergency response system in actual applications can be executed correctly and timely.
[0049] Furthermore, the specific steps of the path reasoning based on reinforcement learning in step (4) are as follows:
[0050] (4.1) Path reasoning strategy modeling: divided into state definition and action definition. State definition: the state of the agent in the dam emergency response knowledge graph is composed of the current entity, query relationship and historical path; action definition: the action space is composed of all connected relationships of the current entity; reward feedback: according to the fusion reward function, the feedback value of the agent's behavior is calculated, combined with the rule reward and the terminal reward, the reasoning process is optimized to ensure that the agent gradually makes correct emergency response decisions.
[0051] (4.2) Strategy network design: Use the long short-term memory network (LSTM) to encode the agent's historical path, generate path features, and output the probability distribution of the current action space through the feedforward neural network to guide the agent to choose the optimal action. In the dam emergency response system, LSTM can effectively process historical emergency response data and help the agent predict the optimal decision path in the future.
[0052] (4.3) Reasoning path output: The agent starts from the starting entity, gradually selects relationships and entities through multi-step reasoning, and finally reaches the target entity). The path generated by the reasoning process can provide decision makers with an explainable reasoning path, helping them understand the decision logic of the agent and ensure the accuracy and timeliness of emergency response measures.
[0053] A sparse knowledge graph reasoning system based on fusion reward, including the following four modules:
[0054] (1) Constructing a fusion embedding module: The knowledge graph of dam emergency response is fully represented by fusing semantic association information with graph structure information. This module specifically includes the following contents: Semantic association information extraction: Based on the TransE embedding method, the semantic association between entities and relationships in the dam emergency response system is captured through vector offset. Graph structure information encoding: Based on the graph convolutional network (GCN), the contextual structure information in the dam emergency response system is extracted to help the intelligent agent understand the relationship between entities. Fusion representation generation: Through the weighted fusion method, the output feature vectors of TransE and GCN are combined to construct a graph fusion embedding vector.
[0055] (2) Rule construction module: The rule-based AnyBURL (Anytime Bottom-Up Rule Learning) method is used to summarize the rule set from the sparse knowledge graph of the dam emergency response system. Through path sampling and inductive logic programming techniques, the Horn rule set is constructed to capture the potential relationship rules in the dam emergency response. The support and confidence of the obtained rules are calculated and evaluated to assess the quality of the rules and provide powerful prior knowledge for the agent reasoning.
[0056] (3) Reward shaping module: Using rule sets and rule-based reward shaping methods, we design a fusion reward function to alleviate the reward sparsity problem in reinforcement learning. This module specifically includes: Rule-based rewards: Calculate the quality of rules based on rule support and confidence, and inject rule rewards into the reinforcement learning process as prior knowledge. Fusion reward function: Fusion the rule reward with the traditional terminal reward according to the set weights to generate the final reward function and optimize the agent's reasoning learning strategy.
[0057] (4) Reinforcement learning reasoning module: A reinforcement learning-based reasoning strategy network is used for path reasoning, including: Strategy network design: The long short-term memory network (LSTM) is used to encode the historical path of the agent and generate a path feature vector. Action selection: The strategy network selects the optimal action in the action space according to the current state, and updates the reasoning path until the target entity is reached. Path generation: The final reasoning path is generated through multi-hop reasoning, and the target entity and path results are output to provide an explainable reasoning path for emergency response, ensuring the transparency and accuracy of decision-making.
[0058] The implementation process of the system and method is the same and will not be repeated here.
[0059] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the sparse knowledge graph reasoning method based on fusion reward as described above is implemented.
[0060] A computer-readable storage medium stores a computer program for executing the sparse knowledge graph reasoning method based on fusion reward as described above.
[0061] Beneficial effects: The present invention introduces considerations of dynamic prediction functions and rule length, and can give priority to short path rules, thereby significantly improving reasoning efficiency and ensuring faster decision-making in emergency situations. In the dam emergency response knowledge graph, due to insufficient data or real-time update restrictions, some emergency measures or equipment information may be missing, resulting in sparse graphs. The present invention can dynamically predict the potential relationships of the current entity through dynamic completion and replacement strategies, thereby generating additional action spaces. This method effectively alleviates the problem of missing paths in sparse graphs, and can flexibly adjust the decision space during the reasoning process, avoiding fixed path searches. The system dynamically adjusts decisions according to the current emergency state and the existing information in the graph, optimizes the reasoning path, and ensures the timely implementation of emergency response measures. Through dynamic prediction and dynamic completion, the sparsity, lengthy paths, and missing information problems in the dam emergency response knowledge graph are alleviated. This method not only improves the efficiency of the emergency response knowledge graph reasoning system, but also improves the accuracy and flexibility of the system in different emergency situations, providing strong support for the safe operation of the dam and the response to sudden disasters. The system can respond quickly in an emergency, minimize decision delays, and contribute to ensuring the safety of the dam and the surrounding environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a flow chart of a method according to an embodiment of the present invention;
[0063] Figure 2 It is the graph information fusion embedding representation of the embodiment of the present invention;
[0064] Figure 3 A schematic diagram of the interaction between the strategy network and the reward signal according to an embodiment of the present invention;
[0065] Figure 4 This is an overall framework diagram of the sparse graph reasoning method based on reinforcement learning with fusion rewards according to an embodiment of the present invention. DETAILED DESCRIPTION
[0066] The present invention is further explained below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0067] like Figure 1 As shown in the figure, a sparse graph reasoning method based on reinforcement learning with fusion reward specifically includes the following steps:
[0068] Step (1) embeds the semantic association information of the dam emergency response knowledge graph, such as Figure 2 :
[0069] (1.1) Extracting semantic association information based on TransE: In the dam emergency response knowledge graph, each triple (h, r, t) is represented as "head entity + relationship = tail entity", for example: h = "water level sensor", r = "monitoring", t = "water level exceeds the warning line". The target formula of TransE is:
[0070] h+r≈t
[0071] By minimizing the scoring function:
[0072] Score(h,r,t)=-||h+rt|| p
[0073] The embedding representation of each entity and relationship can be learned to capture the semantic association of the knowledge graph. For example, the semantics of "water level sensor monitoring water level exceeding the warning line" is embedded as a vector representation. By minimizing the scoring function, the model learns the vector representation of each entity and relationship, thereby capturing the semantic association in the dam emergency response knowledge graph.
[0074] (1.2) Extracting graph structure information based on GCN: GCN aggregates the information of "water level sensor" and its related nodes (such as "standby pump station") through the adjacency matrix to generate the graph structure embedding of the node. The specific formula is:
[0075]
[0076] For example, the graph structure embedding of “water level sensor” will integrate the information of its neighboring nodes such as “water level exceeds warning line” and “backup pump station” to reflect its contextual structure and generate an embedded representation of the graph structure, such as the structural relationship between the water level sensor and the backup pump station and flood outlet.
[0077] (1.3) Fusion of semantic and structural information: Concatenate the embedding representations of TransE and GCN to obtain a comprehensive representation:
[0078] E fusion =ω1E TransE +ω2E GCN
[0079] Among them, w1 is the weight associated with the TransE embedding representation, which controls the contribution of the TransE embedding to the final fusion; w2 is the weight associated with the GCN embedding representation, which controls the contribution of the GCN embedding to the final fusion. Through this embedding representation, the embedding of "water level sensor" contains both semantic information (such as "monitoring") and reflects its structured relationship in the graph, helping the agent to make more accurate decisions during the reasoning process. Specifically, when making decisions, the agent first obtains the semantic and structural information in the knowledge graph through the two methods of TransE and GCN, and then fuses these two types of information according to the weights, thereby providing the agent with a complete perspective to help it choose the optimal path in complex emergency response scenarios. By constantly adjusting these weights, the agent can better balance different types of information when making decisions, optimize the final decision-making process, and improve the efficiency and accuracy of the dam emergency response.
[0080] Step (2) Construct a rule set:
[0081] (2.1) Mining rules through AnyBURL: Induce Horn rules related to the water conservancy field from the knowledge graph. For example:
[0082] H("flood discharge outlet", "start")←"water level exceeds warning line"∧"rainfall is too heavy".
[0083] H("Backup pump station", "Start")←"Spillway failure"∧"High drainage demand".
[0084] These Horn rules indicate that the dam emergency response equipment (such as spillways, backup pumping stations) will be activated under certain conditions. For example, when the water level exceeds the warning line and the rainfall is too heavy, the spillway is activated; when the spillway fails and the drainage demand is high, the backup pumping station is activated. These rules help the reasoning system make emergency response decisions based on historical data and actual conditions.
[0085] (2.2) Support calculation: For a rule such as “water level exceeds warning line → start flood discharge outlet”, calculate the support:
[0086]
[0087] For example, the rule “water level exceeds warning line → activate flood discharge outlet” matched 100 correct triples in the reasoning task.
[0088] (2.3) Confidence calculation: The confidence is calculated by the ratio of the correct facts of the rule prediction head node:
[0089]
[0090] For example, if the rule matches 150 triples, 100 of which are correct, the confidence is 100 / 150 = 0.67. Rules with higher confidence can better guide the reasoning process and help the agent make more accurate and reliable decisions.
[0091] Step (3) Design the fusion reward function:
[0092] (3.1) Design rule reward function: Calculate the rule reward based on the support and confidence of the rule. For example, for the rule "water level exceeds the warning line → start the flood discharge port", assuming that the support of the rule is 100 and the confidence is 0.67, the rule reward can be calculated as:
[0093] R r (Q) = ∑ r∈Q support(r)×conf(r)
[0094] For example, for the rule “water level exceeds warning line → activate flood discharge outlet”, the rule reward is 100·0.67=67.
[0095] (3.2) Terminal reward: After the reasoning is completed, if the agent reaches the target entity (such as the “flood discharge port”), the terminal reward is 1, otherwise the reward is 0. The terminal reward ensures that the agent’s reasoning focuses on the correct target entity and promotes the achievement of the final goal of the reasoning process.
[0096] (3.3) Fusion reward: Combine the rule reward and terminal reward to generate the final reward function:
[0097] R total =λR r +(1-λ)R h
[0098] For example, when the agent reaches the “flood discharge port”, the reward is 0.5·1+0.5·67=34. This fused reward function will help the agent balance the effectiveness of the rules and the accuracy of the goals during the reasoning process, ensuring that the system can make correct and timely decisions during the emergency response process.
[0099] Step (4) is path reasoning based on reinforcement learning, such as Figure 3 :
[0100] (4.1) Path reasoning strategy modeling: In the dam emergency response system, the state of the agent is defined as "current entity (water level sensor) + query relationship (monitoring) + historical path", and the action is to select the next hop relationship or entity (such as jumping from "water level sensor" to "flood discharge port" or "standby pump station"). The agent decides the next reasoning action based on the current state and the accumulated information of the historical path, and explores an effective emergency response path.
[0101] (4.2) Strategy network design: A long short-term memory network (LSTM) is used to encode the agent’s historical path to capture the impact of past decisions and generate a path feature vector. The probability distribution of the current action space is output through a feedforward neural network to guide the agent to choose the optimal action. Specifically, starting from the “water level sensor”, the agent may choose the following path:
[0102] (“Monitoring”, “Water level monitoring line”).
[0103] ("Connection", "Backup Pumping Station").
[0104] For example, the agent can infer water level changes based on historical information and decide whether to take emergency measures, such as activating a backup pump station or opening a flood discharge outlet.
[0105] (4.3) Reasoning path output: The agent selects a path based on the fusion reward signal (combining rule rewards and terminal rewards), and finally generates a reasoning result to help decision makers respond in emergency situations. The path generated by the reasoning result is interpretable, making the decision process more transparent, such as Figure 4 .like:
[0106] Path 1: “Water level sensor” → “Monitoring” → “Water level exceeds warning line” → “Trigger” → “Flood discharge outlet”.
[0107] Path 2: "Water level sensor" → "Connect" → "Backup pump station" → "Start".
[0108] These reasoning paths show how the agent can start from the starting entity and gradually select appropriate emergency response measures to ensure the safety of the dam. For example, if the water level exceeds the warning line, the system automatically chooses to trigger the flood discharge. If there is an equipment failure or other emergency needs, the agent may choose to start a backup pump station.
[0109] A sparse knowledge graph reasoning system based on fusion reward, including the following four modules:
[0110] (1) Embedding representation module: By embedding the entities and relations in the dam emergency response knowledge graph, a unified fusion embedding is generated. First, the TransE model is used to extract the semantic association information in the knowledge graph to capture the semantic offset features between entities (such as "water level sensor") and relations (such as "monitoring") in the dam emergency response system; then, GCN (graph convolutional network) is used to extract the graph structure information of the knowledge graph and aggregate the contextual features of the entity neighborhood (such as "backup pump station"). Finally, through weighted fusion, the semantic information and structural information are integrated into a unified embedding representation, providing comprehensive feature support for subsequent rule construction and reasoning, ensuring that the intelligent agent can make efficient decisions based on rich knowledge.
[0111] (2) Rule construction module: Based on the embedded representation, the AnyBURL (Anytime Bottom-Up Rule Learning) algorithm is used to perform path sampling and rule induction on the dam emergency response knowledge graph to generate a rule set. The quality of the rules is evaluated by calculating the support and confidence of the rules, and high-quality rules are screened to form an optimized rule set. This rule set provides prior knowledge for the reward mechanism and guides the reinforcement learning agent to prioritize high-quality reasoning paths, thereby improving the accuracy and timeliness of emergency response.
[0112] (3) Fusion reward shaping module: Based on the rule set and the task objectives of reinforcement learning, a fusion reward mechanism is designed. The rule reward is calculated by rule support and confidence to guide the agent to prioritize high-quality paths; combined with the terminal reward, positive feedback is given when the agent completes the target task (reaching the target entity, activating emergency equipment). The two rewards are fused through the balance factor to generate the final reward signal, which is used to optimize policy learning in reinforcement learning and alleviate the reward sparsity problem in sparse knowledge graphs.
[0113] (4) Reinforcement learning path reasoning module: The path reasoning task is implemented based on the reinforcement learning framework, and the dam emergency response knowledge graph reasoning is modeled as a Markov decision process (MDP). The agent selects the path action through the policy network combined with the current state (including the starting entity, query relationship and target entity) to generate the reasoning path. During the training process, the agent optimizes the policy network according to the fused reward signal and gradually improves the path selection. After the reasoning is completed, the output path result is interpretable and can effectively predict the target entity relationship, providing support for knowledge graph completion and other downstream tasks (such as automatic equipment startup, flood discharge, etc.).
[0114] Obviously, those skilled in the art should understand that the various steps of the sparse knowledge graph reasoning method based on fusion reward or the various modules of the sparse knowledge graph reasoning system based on fusion reward of the above-mentioned embodiment of the present invention can be implemented by a general computing device. These modules can be concentrated on a single computing device, or distributed in a network composed of multiple computing devices to work together. Optionally, they can be implemented by executable program codes of the computing device, so that the program codes can be stored in a storage device and executed by the computing device. In addition, in some cases, the steps of the method can be performed in an order different from that described or shown herein, or they can be made into separate integrated circuit modules, or multiple modules or steps therein can be combined to form a single integrated circuit module for implementation.
[0115] For example, the embedding representation module can be distributed in multi-core processors through parallel computing to improve the embedding computing efficiency of the knowledge graph; the rule construction module and the fusion reward shaping module can be implemented through a deep learning framework, such as TensorFlow or PyTorch, using existing hardware acceleration technologies (such as GPU or TPU) to optimize the efficiency of rule extraction and reward calculation; the reinforcement learning path reasoning module can be implemented through a reinforcement learning framework, and the decision network of the intelligent agent can be trained online or offline; the reasoning path generation module can be implemented as a separate software module for post-processing and outputting the reasoning results. In this way, the embodiments of the present invention are neither limited to any specific hardware architecture nor to a specific software implementation.
Claims
1. A sparse graph reasoning method based on reinforcement learning with fusion rewards, characterized in that: The following steps are involved: (1) Constructing a fusion embedding module: By combining the semantic association embedding representation based on TransE and the graph structure embedding representation based on GCN, a fusion embedding representation of the sparse knowledge graph in the dam emergency response system is generated; (2) Introducing rule sets: Using the AnyBURL method, the sparse knowledge graph in the dam emergency response system is summarized and the logical rule sets related to emergency response are extracted; And evaluate the support and confidence of the rules; (3) Designing a fusion reward function: Constructing a fusion reward function by combining rule-based rewards and terminal rewards; (4) Path reasoning based on reinforcement learning: By utilizing the interaction between the reinforcement learning agent and the dam emergency response knowledge graph environment, the optimal action is selected through the policy network, and the reasoning path is gradually generated to ultimately realize the execution of emergency response measures.
2. The sparse knowledge graph reasoning method based on fusion reward for reinforcement learning according to claim 1 is characterized in that: The specific steps of constructing the fusion embedding module in step (1) are as follows: (1.1) Based on the TransE method, the semantic association information of the dam emergency response knowledge graph is embedded. For the relationship triple (h, r, t) in the dam emergency response system, the goal of TransE is to make the vector of the head entity plus the vector of the relationship equal to the vector of the tail entity as much as possible. The target formula is as follows: h+r≈t Among them, h, r, and t represent the embedding vectors of the head entity, relation, and tail entity, respectively; The consistency of the triples is measured using the following scoring function: Score(h,r,t)=-||h+r-t|| p where ||·|| p The scoring function uses the norm l1 or l2 to measure the distance between embedded vectors. By minimizing the scoring function, the semantic association embedding representation of entities and relations is learned, providing accurate semantic expression for reasoning tasks in the dam emergency response system. (1.2) Based on the GCN method, the graph structure information of the dam emergency response knowledge graph is encoded. For the relation triple (h, r, t), the goal of GCN is to predict the representation of the tail entity t by aggregating the neighbor information of the head entity h and the representation of the relation r; the graph convolution layer forward propagation formula of its entity node update process is expressed as follows: where σ is the activation function, is the result of adding self-connection to matrix A, is the matrix of measurement, which is connected by the matrix Sum each row of ; (1.3) Fusion of semantic association information and graph structure information: The semantic association information and graph structure information in the dam emergency response knowledge graph are extracted by TransE and GCN respectively. Based on the above operations, the semantic association representation and graph structure information of the graph are encoded to obtain their vector representations respectively. The two vectors are concatenated to obtain the concatenated vector E. cat , the specific formula is shown as follows: Among them, E TransE represents the entity semantic association information representation of the graph by the TransE model, E GCN Represents the entity graph structure information representation based on GCN; Then the three vectors E TransE 、E GCN and E cat Through a fully connected layer, the weighted fusion method is used to map them into the same semantic space. The fused vector retains both the semantic association information and the graph structure information of the graph. E=w1·E TransE +w2·E GCN +w3·E cat Among them, E represents the entity feature vector embedded based on the fusion of image information, and w is the weight calculated by the fully connected layer.
3. The sparse knowledge graph reasoning method based on fusion reward for reinforcement learning according to claim 1 is characterized in that: The specific steps of constructing the rule set in step (2) are as follows: (2.1) Logical rules are mined from the dam emergency response knowledge graph using the AnyBURL method; each rule in the rule set is formally represented as a Horn rule. H r (X,Y)←b1(X,A1)∧...∧b n (A n ,Y) The obtained Horn rule is represented by H, X, Y, A n Represent different entities respectively. The entity association in the rule is represented by b n (...) indicates that for the constructed Horn rule H r (e i ,e j ), which is equivalent to the fact triple (e i ,r,e j ); Each Horn rule is used to describe the logical relationship between different entities in the dam emergency response system; (2.2) AnyBURL evaluates each rule in the rule set through the inference results in pre-training; the support of rule r indicates how many triples in a given knowledge graph the rule can match in the dam emergency response task, and the formula is as follows: in, Indicates the number of correct triples that satisfy the rule; (2.3) The confidence of the rule is evaluated and determined by calculating the ratio of the correct facts of the rule head node prediction. The confidence of rule r is expressed as follows: in, Indicates the number of all triples matched by the rule.
4. The sparse knowledge graph reasoning method based on fusion reward for reinforcement learning according to claim 1 is characterized in that: The specific steps of designing the fusion reward function in step (3) are as follows: (3.1) Design a rule reward function and evaluate the rule quality through rule support and confidence: First, design a rule reward function based on the rule set constructed by AnyBURL. The rule reward function evaluates the quality of the rule based on the rule support and confidence. Specifically, the support measures the matching of the rule in the dam emergency response system, and the confidence reflects the predictive ability of the rule. R r (Q)=∑ r∈Q support(r)×conf(r) Among them, Q is the rule set constructed by AnyBURL, R r (Q) is the definition of reward function corresponding to rule-based reward shaping; (3.2) Definition of terminal reward: The terminal reward is given only when the agent reaches the target entity correctly, and the value is 1; otherwise the reward is 0. The terminal reward ensures that the agent's reasoning can focus on the correct target entity, thereby completing the correct emergency response; the formula is as follows: Among them, s T Indicates the end state of the search process, e s It represents the starting entity of the search process, r q Indicates the relationship of the query, e t represents the target entity of the search process; (3.3) The rule reward is calculated through the support and confidence provided by the rule set, and finally the rule-based reward and hit reward are fused; the rule reward is calculated based on the support and confidence provided by the rule set; then, the rule reward and the terminal reward are weightedly fused to form the final fusion reward function to optimize the learning process of the intelligent agent; the final fusion reward function is: R total =λR r +(1-λ)R h Among them, R h is the hit reward function after reward shaping, R r is a rule-based reward function, and λ is a reward balancing factor that is set; (3.4) Application of fusion rewards: Through the guidance of rule rewards, the agent can explore potential relationship paths in the sparse dam emergency response knowledge graph to ensure that the emergency response path is found; the introduction of terminal rewards ensures that the agent's path reasoning can focus on the correct target entity.
5. The sparse knowledge graph reasoning method based on fusion reward for reinforcement learning according to claim 1 is characterized in that: The specific steps of the path reasoning based on reinforcement learning in step (4) are as follows: (4.1) Path reasoning strategy modeling: divided into state definition and action definition. State definition: the state of the agent in the dam emergency response knowledge graph consists of the current entity, query relationship and historical path; action definition: the action space consists of all connected relationships of the current entity; reward feedback: according to the fusion reward function, the feedback value of the agent's behavior is calculated, combined with the rule reward and the terminal reward, the reasoning process is optimized to ensure that the agent gradually makes correct emergency response decisions; (4.2) Strategy network design: Use the long short-term memory network to encode the agent's historical path, generate path features, and output the probability distribution of the current action space through the feedforward neural network to guide the agent to choose the optimal action; (4.3) Reasoning path output: The agent starts from the starting entity, gradually selects relationships and entities through multi-step reasoning, and finally reaches the target entity.
6. A sparse knowledge graph reasoning system based on fusion reward, characterized in that: It includes the following four modules: (1) Constructing a fusion embedding module: By fusing semantic association information with graph structure information, the sparse knowledge graph in the dam emergency response system is fully represented. This module specifically includes the following contents: Semantic association information extraction: Based on the TransE embedding method, the semantic association between entities and relations in the dam emergency response system is captured through vector offset. Graph structure information encoding: Based on the graph convolutional network, the context structure information in the dam emergency response system is extracted to help the intelligent agent understand the relationship between entities; fusion representation generation: Through the weighted fusion method, the output feature vectors of TransE and GCN are combined to construct a graph fusion embedding vector; (2) Rule construction module: The rule-based AnyBURL method is used to summarize the rule set from the sparse knowledge graph of the dam emergency response system. The Horn rule set is constructed through path sampling and inductive logic programming techniques to capture the potential relationship laws in the dam emergency response. The support and confidence of the obtained rules are calculated and evaluated to assess the quality of the rules and provide prior knowledge for the agent reasoning. (3) Reward shaping module: Using rule sets and rule-based reward shaping methods, we design a fusion reward function to alleviate the reward sparsity problem in reinforcement learning. This module specifically includes: Rule-based rewards: Calculate the quality of rules according to the support and confidence of rules, and inject rule rewards into the reinforcement learning process as prior knowledge; Fusion reward function: Fusion the rule rewards with the traditional terminal rewards according to the set weights to generate the final reward function and optimize the reasoning learning strategy of the intelligent agent; (4) Reinforcement learning reasoning module: A reinforcement learning-based reasoning strategy network is used for path reasoning, including: strategy network design: using a long short-term memory network to encode the agent's historical path and generate a path feature vector; action selection: using the strategy network to select the optimal action in the action space according to the current state, and updating the reasoning path until the target entity is reached; path generation: generating the final reasoning path through multi-hop reasoning, outputting the target entity and path results, and providing an explainable reasoning path for emergency response.
7. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the sparse knowledge graph reasoning method based on fusion reward as described in any one of claims 1 to 6 are implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the sparse knowledge graph reasoning method based on fusion reward as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Knowledge reasoning method and system based on agent dynamic path completion strategy
CN115526321A
Multi-view learning entity alignment method and system for dam emergency response knowledge base linkage
CN115982374A
Dam emergency event causal relationship identification method and system based on feature fusion
CN116738366A
Engineering emergency plan generation method based on knowledge graph
CN117077631A
Cited By
Construction method of double-path collaborative decision network for multi-agent collaborative path optimization
CN120598148A
Image semantic feature and pilot signal mapping method and device based on reinforcement learning
CN120823406A
Psychiatry pre-diagnosis triage knowledge reasoning method based on knowledge graph
CN121543713A
Bayesian network-based tin-based material knowledge graph query method and system and storage medium
CN122509315A