A dynamic knowledge graph driven SCADA policy generation method
Patent Information
- Application Number
- CN202611318129.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本申请提供了一种动态知识图谱驱动的SCADA策略生成方法,用于针对解决现有SCADA策略生成方式依赖固定知识库难以动态适配工况,无法兼顾检索精度、策略质量与推理效率的技术问题
上传SCADA策略生成任务,执行动态构型过程重建,得到自适应图谱;依据所述自适应图谱,通过定义基于增益项与惩罚项的多维奖励函数体系,以检索精度、策略质量与推理效率间的自主权衡为目标,确定检索-推理策略;在所述生成智能体中,生成智能体的动作空间与状态空间,执行循环推理决策,生成任务SCADA策略;在终端显示界面,对所述任务SCADA策略进行可视化。达到了实现自适应知识图谱构建与强化学习循环推理,自主权衡检索精度、策略质量与推理效率,提升SCADA调控策略智能生成能力的技术效果。
Smart Images

Figure CN122840193A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent decision-making technology in industrial automation, specifically to a dynamic knowledge graph-driven SCADA strategy generation method. Background Technology
[0002] Currently, most SCADA system control strategies rely on manual experience and generally use static, fixed knowledge bases to support strategy derivation. These knowledge bases cannot be dynamically updated according to on-site working conditions and tasks, resulting in poor adaptability. Traditional strategy reasoning mechanisms lack quantitative balancing mechanisms, making it difficult to simultaneously consider knowledge retrieval accuracy, generated strategy compliance quality, and reasoning efficiency. This can easily lead to problems such as insufficient knowledge retrieval, excessive reasoning iteration, and redundant output instructions. Consequently, it is difficult to autonomously achieve multi-objective balance and cannot automatically generate reliable SCADA control strategies for diverse industrial scenarios.
[0003] Existing SCADA strategy generation methods rely on fixed knowledge bases, making it difficult to dynamically adapt to working conditions and failing to balance retrieval accuracy, strategy quality, and inference efficiency. Summary of the Invention
[0004] This application provides a dynamic knowledge graph-driven SCADA strategy generation method to address the technical problem that existing SCADA strategy generation methods rely on fixed knowledge bases, making it difficult to dynamically adapt to working conditions and simultaneously achieve retrieval accuracy, strategy quality, and inference efficiency.
[0005] In view of the above problems, this application provides a dynamic knowledge graph-driven SCADA strategy generation method.
[0006] This application provides a dynamic knowledge graph-driven SCADA policy generation method, the method comprising: Upload the SCADA strategy generation task. Based on the graph policy network within the generated agent, perform a dynamic configuration process reconstruction with the graph configuration reward function as a constraint to obtain an adaptive graph. Based on the adaptive graph, determine the retrieval-inference strategy by defining a multi-dimensional reward function system based on gain and penalty terms, with the goal of autonomously balancing retrieval accuracy, policy quality, and inference efficiency. In the generated agent, using the adaptive graph as a reference, define the action space and state space of the generated agent based on the retrieval-inference strategy, perform cyclical inference decision-making, and generate the task SCADA strategy. Visualize the task SCADA strategy on the terminal display interface.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: Upload the SCADA policy generation task, perform dynamic configuration process reconstruction to obtain an adaptive knowledge graph; based on the adaptive knowledge graph, determine the retrieval-inference strategy by defining a multi-dimensional reward function system based on gain and penalty terms, aiming at an autonomous trade-off between retrieval accuracy, policy quality, and inference efficiency; in the generating agent, generate the agent's action space and state space, perform cyclic inference decision-making, and generate the task SCADA strategy; visualize the task SCADA strategy on the terminal display interface. This achieves the technical effect of realizing adaptive knowledge graph construction and reinforcement learning cyclic inference, autonomously balancing retrieval accuracy, policy quality, and inference efficiency, and improving the intelligent generation capability of SCADA control strategies. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic diagram of a dynamic knowledge graph-driven SCADA strategy generation method provided in an embodiment of this application.
[0010] Figure 2 This is a schematic diagram illustrating the process of generating a graph policy network within an intelligent entity in a dynamic knowledge graph-driven SCADA policy generation method provided in an embodiment of this application.
[0011] Figure 3 A bar chart comparing the key performance indicators of different solutions provided in the embodiments of this application. Detailed Implementation
[0012] This application provides a dynamic knowledge graph-driven SCADA strategy generation method to address the technical problem that existing SCADA strategy generation methods rely on fixed knowledge bases, making it difficult to dynamically adapt to working conditions and balance retrieval accuracy, strategy quality, and inference efficiency.
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0014] Examples, such as Figure 1As shown, this application provides a dynamic knowledge graph-driven SCADA policy generation method, the method comprising: Step S100: Upload the SCADA policy generation task, and based on the graph policy network in the generated intelligent body, perform a dynamic configuration process reconstruction with the graph configuration reward function as a constraint to obtain an adaptive graph.
[0015] Specifically, the system receives SCADA policy generation tasks uploaded from external sources and invokes a pre-built graph policy network within the generating agent to perform dynamic graph configuration reconstruction. This graph policy network is constructed based on a Markov decision process, with optimization objectives consisting of task quality indicators such as policy satisfaction, control accuracy, and physical constraint coverage. Its core constraint graph configuration reward function is a weighted fusion of a first reward component representing knowledge coverage sufficiency and based on the reciprocal of the average shortest reachable path length of knowledge entities, and a second reward component representing knowledge retrieval efficiency and based on the average graph traversal hop count of key knowledge entities. The dynamic configuration process uses the graph structure feature vector as the system state and adds new knowledge entity nodes and new entity relationship edges as optional configuration actions. After completing one round of policy execution, the immediate reward is calculated based on the actual running results. The state, actions, and reward trajectories are stored in the experience replay pool and the graph policy network is fine-tuned in reverse. After multiple rounds of iterative optimization, an adaptive graph adapted to the current SCADA task is output.
[0016] Step S200: Based on the adaptive graph, a retrieval-inference strategy is determined by defining a multidimensional reward function system based on gain and penalty terms, with the goal of autonomously balancing retrieval accuracy, strategy quality, and inference efficiency.
[0017] Specifically, the output adaptive graph serves as a shared interactive environment for retrieval and inference. A multi-dimensional reward function system, including gain and penalty terms, is constructed. The gain term consists of a first gain component representing the hit effect of the graph retrieval and a second gain component representing the degree to which the strategy conforms to the physical constraints of the field equipment. The penalty term consists of a first penalty component corresponding to the number of inference steps exceeding a preset threshold and a second penalty component counting the number of redundant and repeated control commands within the strategy. Using this multi-dimensional reward function as the evaluation criterion, multiple rounds of knowledge retrieval are carried out in the adaptive graph. The knowledge subgraph obtained in each round of retrieval is encoded into a graph embedding vector and concatenated with the real-time inference context vector to obtain a complete state feature. Based on this state feature, a dynamic balance and optimization of retrieval accuracy, policy compliance quality, and inference efficiency is achieved, ultimately determining a retrieval-inference strategy suitable for the current SCADA task.
[0018] Step S300: In the generated agent, based on the adaptive graph, the action space and state space of the generated agent are defined by the retrieval-reasoning strategy, and cyclic reasoning decision-making is performed to generate a task SCADA strategy.
[0019] Specifically, within the agent, the obtained adaptive graph serves as the knowledge benchmark. Combined with a defined retrieval-reasoning strategy, the agent's action space and state space are delineated. The action space includes thinking actions, stopping actions, query generation actions, graph retrieval actions, and strategy output actions. The state space consists of the current inference context text sequence, the retrieved knowledge subgraph structure, and the current strategy completeness score. A continuous cyclical inference decision-making process is then executed. Each round of inference takes the real-time inference context text sequence and the retrieved knowledge subgraph structure as input. First, it determines whether the required knowledge completeness threshold for strategy generation has been reached. If the threshold is met, the strategy output action is executed, directly terminating the inference loop and outputting the SCADA strategy adapted to the task. If the threshold is not reached, the query generation action is executed, generating a new graph query sub-target based on the existing inference context. A graph retrieval tool is then used to retrieve and supplement new knowledge subgraphs in the adaptive graph. The newly added knowledge subgraph is integrated into the inference context before entering the next round of inference iterations, until the knowledge completeness threshold is met, at which point the final SCADA strategy is output.
[0020] Step S400: Visualize the task SCADA strategy on the terminal display interface.
[0021] Specifically, the final generated SCADA strategy is transmitted to the front-end terminal display interface for visualization. The interface synchronously loads the adaptive graph structure on which the generation process is based, the complete retrieval-inference execution trajectory, the knowledge subgraph embedding features corresponding to each round of inference, the calculated values of each component of the multi-dimensional reward function, and the final formed control instruction scheme. The knowledge entities, relationships, inference iteration paths, and complete SCADA control strategy text are rendered in layers, intuitively presenting the logic of the entire strategy generation process and the final strategy content, making it convenient for operation and maintenance personnel to view, verify, and adjust the SCADA control strategy.
[0022] In one possible implementation, such as Figure 2 As shown, step S100 further includes: Step S110: Construct a graph configuration reward function adapted to the SCADA policy generation task, wherein the task quality index composed of policy satisfaction, control accuracy, and physical constraint coverage is used as the optimization objective, and the graph configuration reward function is obtained by weighted fusion of the first reward component and the second reward component.
[0023] Step S120: Model the dynamic configuration process of the knowledge graph as a Markov decision process.
[0024] Step S130: Based on the graph configuration reward function and Markov decision process, construct a graph policy network within the generating intelligent body, wherein the input of the graph policy network is the current graph structure features and the semantic description of the SCADA policy generation task, and the output is the probability distribution of the next graph configuration action.
[0025] Specifically, the overall optimization goal for task quality is set based on strategy satisfaction. Control precision Physical constraint coverage Common constraints guide the construction of a graph configuration reward function. First, define the first reward component. The formula is used to quantify the sufficiency of knowledge coverage in a knowledge map. , The average shortest reachable path length between all knowledge entities required to generate strategies within the graph is determined by the length of the path between entities. The higher the value, the more likely a second reward component will be defined. The formula used to quantify map retrieval efficiency is as follows: , This represents the average number of graph traversal hops in the key knowledge retrieval process; the fewer the traversal hops, the better. The higher the value, the better; configure the preset weighting coefficient. , ,satisfy The two components are weighted and fused to obtain the complete graph configuration reward function: During construction, all related knowledge entities in the current knowledge graph are first collected, and the average shortest path between entities is calculated. Statistical analysis of the average number of traversal jumps in multiple rounds of knowledge retrieval Substitute into the formula and solve separately , Then, by combining the preset weights of the process, a weighted calculation is performed to obtain the total reward value corresponding to the current graph configuration, which serves as a feedback constraint indicator for the graph configuration action in the Markov decision-making process.
[0026] A Markov Decision Process (MDP) model is constructed for the dynamic configuration process of a knowledge graph. The model first defines four core elements: a four-tuple (S, A, T, R). The state set S uses a global graph structure feature vector (GDV) obtained by concatenating the node and relation features of the current knowledge graph as a single state representation, completely recording the current distribution of entities and relationships in the graph. The action set A sets only two types of executable configuration actions: adding knowledge entity nodes and adding relation edges between entities, used for iterative adjustments to the graph structure. The state transition function T represents the mapping transformation rule that updates graph nodes and relations and synchronously refreshes the graph structure feature vector after executing any configuration action in the current state. The immediate reward function R directly adopts the pre-constructed graph configuration reward function. After each graph configuration action is completed and the graph state is updated, the average shortest reachable path length of graph knowledge entities is recalculated. Average number of traversal jumps for key knowledge retrieval Substitute the values into the reward function to calculate the immediate reward value corresponding to the action. After fully determining the four elements of state, action, state transition, and immediate reward, the MDP standardized modeling of the graph dynamic configuration process is completed, providing an iteratively computeable decision framework for reinforcement learning training of the graph policy network.
[0027] Based on the aforementioned graph configuration reward function and the already modeled Markov decision process, a graph policy network is constructed within the generative intelligent body. The construction process is as follows: First, the network input dimension and input content are determined. The current graph global structure feature vector and the SCADA policy generation task semantic encoding vector are concatenated to form the overall network input features. Second, the network backbone layer is built, using a multi-layer fully connected neural network as the feature extraction module. The concatenated fused features are subjected to nonlinear transformations layer by layer to mine the correlation features between the graph structure and task requirements. The network output layer is set with output channels that correspond one-to-one with the action set. The action set includes two types of configuration actions: adding knowledge entity nodes and adding relationship edges between entities. The output layer is connected via S... After normalization using the oftmax activation function, the network outputs the probability distributions corresponding to various next graph configuration actions. During the training phase, the states, actions, and immediate reward trajectories obtained through Markov decision process interaction sampling are used as training samples. The total reward value calculated by the graph configuration reward function is used as the basis for loss function optimization. The reinforcement learning gradient descent algorithm is used to iteratively update the network weights in reverse, continuously reducing the deviation between the policy output action probability and the optimal graph configuration action. After training convergence, a graph policy network that can output graph adjustment action probabilities is obtained. During the runtime phase, the network can directly output the probability distributions of each selectable graph configuration action by inputting real-time graph structure features and the current SCADA task semantic description, guiding subsequent graph adaptive reconstruction operations.
[0028] In one possible implementation, step S110 further includes: Step S111: The first reward component is quantified by the reciprocal of the average shortest reachable path length of the knowledge entities required for strategy generation in the graph, which represents the sufficiency of the knowledge coverage of the graph for the strategy generation task. The shorter the average shortest reachable path of the knowledge entities in the graph, the higher the value of the first reward component.
[0029] Step S112: The second reward component is quantified by the average number of graph traversal jumps required for key knowledge retrieval, which represents the knowledge retrieval efficiency of the graph structure. The fewer the average number of traversal jumps for knowledge retrieval, the higher the value of the second reward component.
[0030] Specifically, firstly, all target knowledge entities corresponding to this SCADA strategy generation task are identified. Target knowledge entities refer to graph node entities that describe SCADA equipment, control logic, and operational constraints. Then, a breadth-first search algorithm is used to solve for the shortest connected path length between any two target knowledge entities. The shortest connected path length is the number of edges contained in the path between the two entity nodes that traverses the fewest relational edges. Finally, the arithmetic mean of the shortest connected path lengths of all pairwise entity combinations is calculated to obtain the average shortest reachable path length. The first reward component is obtained by taking the reciprocal of the average shortest reachable path length. When the links between target knowledge entities within the graph are dense and the number of relation edges required for communication is reduced, the average shortest reachable path length is [not specified]. The value is smaller, and the first reward component is calculated by reciprocal. A higher value indicates better connectivity and more comprehensive knowledge coverage in the knowledge graph; conversely, a lower value indicates sparser connectivity links between entities and more jump steps. Too large, first reward amount The corresponding reduction serves as a positive reward basis for the optimization of the graph structure.
[0031] The key knowledge is defined as the graph entities, such as equipment parameters, control rules, and process constraints, necessary to complete the current SCADA strategy generation task. The graph traversal hop count is the number of relational edges traversed from the starting entity node to the target key knowledge entity node. For multiple different retrieval query tasks, graph retrieval traversal operations are performed separately. The traversal hop count from the starting node to the target key entity is counted for each query. The arithmetic mean of the traversal hop counts for all retrieval tasks is then calculated to obtain the average graph traversal hop count. In this embodiment, the formula for calculating the second reward component is set as follows: If the graph structure is compact and the relationships between entities are direct, the number of relational edges required to reach key knowledge is less, then the average number of graph traversal hops is lower. The second reward component, calculated by reciprocal, has a smaller value. A higher value indicates better graph retrieval efficiency; if the relationships between entities are loose and retrieval requires multiple hops, the average number of graph traversal hops will be higher. The value is too large, corresponding to the second reward portion. The value decreases, and this value is used as a positive feedback indicator to evaluate the performance of map retrieval and guide the adaptive adjustment of map structure.
[0032] In one possible implementation, step S110 further includes: The dynamic configuration process takes the graph structure feature vector of the current graph as the system state and adds new knowledge entity nodes and new relationship edges between entities as executable configuration actions.
[0033] Specifically, the dynamic configuration process is the complete workflow for iterative structural adjustments in the knowledge graph adaptation SCADA strategy generation task. This workflow corresponds to the state and action definitions in the Markov decision process. The system state refers to the observed representation of the input in a single iteration of the Markov decision process, using the graph structure feature vector of the current graph as the system state. The graph structure feature vector is a one-dimensional numerical vector obtained by concatenating and normalizing all knowledge entity node features, inter-entity relation edge features, and node connectivity topology features of the graph, which can completely digitally represent the overall topology and entity attribute information of the current graph. The executable configuration actions are the graph adjustment operations that the agent can choose to perform during the Markov decision process, including only two basic types of operations. The system employs two main methods: First, adding new knowledge entity nodes. These nodes represent business objects such as SCADA equipment, control parameters, process constraints, and control logic. Adding these nodes supplements the missing business objects required for the task. Second, adding inter-entity relationship edges. These edges connect two knowledge entity nodes, representing the logical association between the objects. Adding these edges improves the inter-entity association pathways. In each round of dynamic configuration iteration, the real-time graph is extracted to generate corresponding graph structure feature vectors, which are then used as inputs to the graph strategy network to determine the current system state. The network outputs the execution probability of each configuration action, selects the optimal action to update the graph nodes or relationship structure, and completes a single state-to-action interactive iteration.
[0034] In one possible implementation, step S110 further includes: After each round of SCADA strategy generation and execution tasks is completed, the graph configuration reward function is updated based on the actual execution results of the strategy, and the instantaneous reward value corresponding to the current graph configuration is calculated.
[0035] The instant reward value, graph state, and configuration action trajectory of the current round are synchronously stored in the experience replay pool, triggering the graph structure fine-tuning of the graph policy network to obtain an adaptive graph.
[0036] Specifically, after each complete round of SCADA strategy generation and on-site execution, the actual operational feedback data of this strategy is collected to update the graph configuration reward function and solve for the immediate reward value. The specific process is as follows: First, extract the three task quality indicators corresponding to the strategy execution results of this round, namely, strategy satisfaction. Control precision Physical constraint coverage The three indicators together serve as the optimization constraint benchmark for the graph configuration reward function; secondly, all knowledge entities required for the current graph task are read, and the average shortest reachable path length between entities is calculated using a breadth-first search algorithm. Substitute into the formula Solving for the first reward component of representing the sufficiency of knowledge coverage Then, calculate the average number of graph traversal jumps across multiple rounds of key knowledge retrieval. Substitute into the formula Solve for the second reward component that represents retrieval efficiency. ; Retrieve preset weighting coefficients , And satisfy the constraints Through weighted fusion formula The instantaneous reward value corresponding to the current map configuration is calculated. The immediate reward value is adjusted by the reverse constraint of the actual execution effect of the strategy in this round. If problems such as control deviation, constraint failure, or excessively long knowledge retrieval time occur during strategy execution, it will... , As the value increases, the immediate reward value decreases. Conversely, if the strategy is executed successfully and the graph retrieval is efficient, the immediate reward value will be higher. After calculation, the immediate reward value is used as a feedback signal for the Markov decision process and is used for subsequent weight iteration optimization of the graph policy network.
[0037] After completing the graph configuration action and calculating the immediate reward value for each round, the graph state (the graph structure feature vector corresponding to the current graph), the actual configuration action performed (adding knowledge entity nodes / adding relationship edges between entities), and the calculated immediate reward value are extracted. These three are packaged into a single interaction trajectory sample and stored in the experience replay pool. The experience replay pool continuously accumulates historical interaction samples from multiple rounds to eliminate temporal correlations in the training data. When the number of samples in the pool reaches a preset sampling threshold, a batch of samples is randomly selected and sent to the graph policy network for reinforcement learning training. During the training process, the immediate reward calculated by the graph configuration reward function is used as the optimization objective. A policy gradient loss function is constructed, and the weight parameters inside the graph policy network are corrected layer by layer through the backpropagation algorithm. The network is fine-tuned to optimize the output probability distribution of actions for each graph configuration, making it more inclined to output graph adjustment actions that can obtain higher rewards. After training, based on the updated graph policy network, the semantics of the current SCADA task and the graph structure features are input, the network outputs the probability of each configuration action and selects the optimal action to execute, and adds entity nodes or relation edges to the knowledge graph to complete the structural update. The complete process of "generating SCADA policy - calculating immediate reward - storing interaction trajectory - sampling, training and fine-tuning policy network - updating graph topology" is executed cyclically to continuously optimize the graph entity and relation structure until the immediate reward values converge and stabilize after multiple rounds and the graph topology no longer changes significantly, finally obtaining an adaptive graph that fits the current SCADA policy generation task.
[0038] In one possible implementation, step S200 further includes: Step S210: Using the adaptive graph as a shared interactive environment for retrieval and reasoning, construct a multi-dimensional reward function system that includes gain and penalty terms.
[0039] Step S220: Based on the multidimensional reward function, perform multiple rounds of retrieval in the adaptive graph, encode the accumulated knowledge subgraph into a graph embedding vector, and concatenate it with the current inference context vector to form state features.
[0040] Step S230: Determine the retrieval-reasoning strategy based on the state characteristics.
[0041] Specifically, the adaptive graph obtained through iterative optimization serves as the interactive carrier environment shared by knowledge retrieval and logical reasoning. The adaptive graph refers to a knowledge graph that, after multiple rounds of fine-tuning by the graph strategy network, adapts to the current SCADA strategy's generated task topology. When constructing the multi-dimensional reward function system, it is first divided into two main modules: a gain term for positive incentives and a penalty term for negative constraints. The gain term has a first gain component and a second gain component, and the penalty term has a first penalty component and a second penalty component. The first gain component represents the retrieval hit rate of intermediate queries in the adaptive graph's graph retrieval sample out of the total number of queries, used to measure the knowledge retrieval matching effect. The second gain component represents the retrieval hit rate of the reasoning output strategy scheme. The compliance rate of SCADA system equipment, process, and safety physical constraints is used to evaluate the compliance quality of the generated strategy. The first penalty component is the difference between the total number of retrieval and inference steps and the preset standard step threshold, used to limit excessive iteration in the inference process. The second penalty component is the statistical number of repeated and meaningless redundant control instructions within the strategy, used to constrain the simplicity of the strategy. After defining the four types of components, independent weight coefficients are set for each component. All gain and penalty components are merged by weighted summation to form a complete and quantifiable multidimensional reward function system. In subsequent multi-round retrieval and inference interactions, the comprehensive reward value output by this function is used to evaluate the merits of the current retrieval-inference strategy, providing quantitative feedback for strategy optimization.
[0042] Using a constructed multidimensional reward function as the criterion for evaluating the quality of retrieval actions, multiple rounds of knowledge retrieval operations are iteratively performed within an adaptive graph. Each round generates a unique graph query instruction based on the real-time inference context, matching corresponding SCADA equipment, control constraints, and process logic-related entity nodes and relationships within the adaptive graph. The matched graph fragments are aggregated round by round, accumulating and integrating to form a local knowledge subgraph covering the knowledge required for the current task. Subsequently, a graph neural network encoder is invoked to extract global topology and attribute features from this accumulated knowledge subgraph, fusing all node features and relational edge topology information within the subgraph, and compressing it to generate a fixed-dimensional graph embedding vector. This vector quantifies and carries the total amount of structured knowledge retrieved. Simultaneously, [further details are needed for accurate translation]. A pre-trained text encoder is used to semantically encode the continuously updated inference context text sequence to generate an inference context vector of equal dimension. This vector records the task progress, intermediate query logic, and generated control command information of the current inference stage. The graph embedding vector and the inference context vector are directly concatenated and fused along the feature dimension to obtain a unified state feature. This state feature contains two core types of information: the total amount of structured knowledge acquired and the real-time inference progress. This information can be used by the agent to make adaptive decisions. When knowledge is insufficient, the agent can choose to continue deep searching to supplement the graph information. When the knowledge completeness reaches the target, the inference can be terminated in time and the SCADA strategy can be output. This achieves a dynamic trade-off decision of deep exploration on demand and rapid stopping when the target is reached.
[0043] The state features obtained by splicing and fusing are input into the retrieval-inference strategy discriminant network to determine the retrieval-inference strategy. The state features contain both the graph embedding vector corresponding to the accumulated knowledge subgraph and the current inference context vector, fully carrying the graph retrieval knowledge and semantic information of the inference stage. The discriminant network is a pre-trained multi-layer fully connected network with the optimization objective of maximizing the comprehensive reward corresponding to the multi-dimensional reward function. After receiving the state features, the network performs nonlinear feature extraction layer by layer and outputs the estimated comprehensive reward value corresponding to different retrieval and inference combinations. It traverses all available retrieval and inference behaviors, compares the estimated reward values corresponding to each behavior, and selects the set of retrieval frequency, graph query generation rules, and inference iteration termination conditions with the highest comprehensive reward as the optimal behavior combination. This optimal combination is the retrieval-inference strategy adapted to the current SCADA task. This strategy clarifies the query generation logic, graph retrieval scope, knowledge supplementation timing, and knowledge completeness judgment criteria in each subsequent round of inference, guiding the generative agent to carry out cyclical inference decisions, balancing retrieval hit rate, policy compliance quality, and inference efficiency.
[0044] In one possible implementation, step S210 further includes: Step S211: The gain term includes a first gain component and a second gain component. The first gain component is the hit rate obtained by the generated intermediate query in the adaptive graph execution graph retrieval, and the second gain component is the degree to which the strategy scheme output in the current inference stage satisfies the physical constraints of the SCADA system.
[0045] Step S212: The penalty item includes a first penalty component and a second penalty component. The first penalty component is the step excess value corresponding to the actual number of reasoning steps exceeding the preset step threshold during the retrieval-reasoning process.
[0046] Step S213: The second penalty component is the number of redundant and repetitive control instructions within the policy scheme generated in the current inference phase.
[0047] Specifically, the gain term is a quantitative indicator that positively incentivizes retrieval and reasoning behaviors, comprising two independently computable indicators: a first gain component and a second gain component. The first gain component characterizes the graph retrieval matching effect, specifically referring to the proportion of queries that successfully match valid knowledge entities or relationships after graph retrieval operations are performed in the adaptive graph, out of the total number of intermediate queries. This is the retrieval hit rate. A higher value for this component indicates stronger retrieval effectiveness if intermediate queries can hit a large number of valid business knowledge entries within the graph. The second gain component characterizes the compliance quality of the reasoning output strategy. Specifically, this refers to the SCADA strategy scheme initially generated during the current inference phase. Each control instruction within the scheme is verified to ensure compliance with the physical constraints of the SCADA system, such as equipment operating limits, process operation specifications, and safety interlock rules. The ratio of the number of control instructions that satisfy all physical constraints to the total number of instructions in the strategy is used as the physical constraint satisfaction. The closer the strategy scheme is to the physical operating boundary of the field equipment, the higher this component's value. Two types of gain components provide positive rewards from the dimensions of retrieval effectiveness and strategy compliance, respectively. Subsequently, they will be combined with the penalty component to calculate a weighted comprehensive reward value, which is used to evaluate the merits of the current retrieval-inference strategy and guide iterative optimization of the strategy.
[0048] The penalty term includes a first penalty component and a second penalty component. The first penalty component is used to constrain the iteration time of the retrieval-inference process to prevent infinite iterations that would waste computing resources. A threshold for the number of inference steps is pre-set based on the real-time requirements of SCADA policy generation. This threshold is the reasonable maximum number of iterations required to complete knowledge retrieval and policy derivation normally. During the continuous execution of retrieval and inference, the actual number of inference steps is accumulated in real time. The excess step value is obtained by subtracting the preset step threshold from the actual number of inference steps, and this value is used as the first penalty component. When the actual number of inference steps is less than or equal to the preset step threshold, the excess step value is zero, and no negative constraint is introduced. When the actual number of inference steps exceeds the preset step threshold, the excess step value is greater than zero. This component participates in the multidimensional reward function calculation and reduces the comprehensive reward value. Relying on the negative feedback mechanism, the agent is guided to actively reduce the ineffective iteration process, ensuring the inference efficiency of the SCADA policy generation task.
[0049] The second penalty component is part of the penalty term within the multidimensional reward function system. It is used to constrain the simplicity of the reasoning-generated strategy and avoid invalid instructions in the output. After each round of reasoning iteration is completed and a preliminary SCADA strategy scheme is obtained, all control instructions contained in the strategy are analyzed one by one. Based on the execution object, control action, and triggering condition, similarity judgment is performed to identify redundant control instructions with semantic repetition, functional overlap, and no actual regulatory effect. The total number of such instructions is counted as the second penalty component. The more redundant instructions there are, the larger the value of the second penalty component. Including this component in the multidimensional reward function for weighted calculation will reduce the overall reward result. Through negative incentives, the agent is guided to simplify the strategy content, prompting the final output SCADA strategy to eliminate invalid and repetitive instructions, thereby improving the readability of the strategy and the efficiency of on-site execution.
[0050] In one possible implementation, step S300 further includes: Step S310: Using the retrieval-reasoning strategy as the execution strategy, define the action space and state space of the generated agent.
[0051] Step S320: In each round of reasoning, using the current reasoning context text sequence and the retrieved knowledge subgraph structure as input, evaluate whether the knowledge completeness threshold required for policy generation is met.
[0052] Step S330: If satisfied, execute the policy output action, terminate the inference loop, and output the task SCADA policy.
[0053] Specifically, a defined retrieval-inference strategy is used as the execution benchmark for generating the agent, and the action space and state space of the agent are standardized. The state space contains all legal states, and each state is represented by a state feature formed by concatenating the fusion graph embedding vector and the inference context vector. Each set of state features corresponds to the current knowledge scale and inference progress within the adaptive graph. The action space contains all optional operations that the agent can execute, specifically divided into two basic actions: continuously initiating a new round of graph retrieval and stopping retrieval and outputting the final SCADA strategy. During the strategy execution cycle, the agent selects actions to be executed within the action space based on the input real-time state features, evaluates the expected benefits corresponding to different actions based on the multi-dimensional reward function, and thus completes adaptive decision-making, realizing the operational logic of continuing retrieval when knowledge is insufficient and terminating inference and outputting the strategy when information is sufficient.
[0054] In each round of inference, the current inference context text sequence and the retrieved knowledge subgraph structure are used as evaluation inputs to quantitatively determine whether the existing information meets the knowledge completeness threshold required for policy generation. First, the set of target knowledge entities required for the SCADA policy generation task, as defined by the inference context, is extracted. Simultaneously, extract the set of entities that have been successfully acquired in the current knowledge subgraph. Define knowledge coverage , This represents the number of entities within the set; simultaneously, it represents the total number of relationships that need to be established between the target knowledge entities. And the number of valid relationships already present in the current knowledge subgraph. Define association completeness Configure weight coefficients , And satisfy Construct a formula for calculating knowledge completeness. The calculated knowledge completeness S is compared with a pre-set knowledge completeness threshold. To make a comparison, if If sufficient information is available, the retrieval and reasoning process can be terminated and a SCADA strategy can be generated. If the information is found to be missing, a new round of knowledge retrieval needs to be carried out in the adaptive graph to supplement the missing entities and relationships.
[0055] When the assessment determines that the current knowledge completeness is greater than or equal to the preset knowledge completeness threshold, the agent selects a strategy within the action space to output an action and execute it, immediately terminating the continuous iterative retrieval and inference loop. The agent integrates the current inference context information and all valid knowledge, such as equipment parameters, control logic, and safety constraints, contained in the accumulated retrieved knowledge subgraph. It then organizes this knowledge in a structured manner according to the standard instruction format of the SCADA system, automatically generates a complete SCADA strategy that adapts to the on-site control requirements, and outputs it outward. The output SCADA strategy carries timing information, controlled objects, execution conditions, and operating parameters, and can be directly sent to the SCADA control system to complete equipment control. The current retrieval and inference process then ends.
[0056] In one possible implementation, step S300 further includes: Step S340: If not satisfied, execute the query generation action.
[0057] Step S350: Generate a new graph query sub-target based on the current inference context text sequence, call the graph retrieval tool to perform retrieval in the adaptive graph to obtain a new knowledge sub-graph, integrate it into the current inference context text sequence, and then proceed to the next iteration.
[0058] Specifically, when the assessment determines that the current knowledge completeness is less than the preset knowledge completeness threshold, the agent performs a query generation action. The agent identifies the missing target knowledge entities and their relationships based on the current inference context, and constructs a targeted structured query statement by combining the entity type and relationship definition rules of the adaptive graph. The generated query statement is then input into the adaptive graph to conduct a new round of knowledge retrieval, obtain the missing entity and relationship information, update the knowledge subgraph structure and inference context, and then enter the next round of inference loop to conduct another knowledge completeness assessment and continuously supplement the business knowledge required for strategy generation.
[0059] Analyze the current inference context text sequence to locate the missing target knowledge entities and relationships in the existing knowledge subgraph, and generate new graph query sub-targets based on task semantics; employ a missing entity matching loss function. Quantify the knowledge gap, where, This represents the complete set of target knowledge entities required for strategy generation. This represents the set of knowledge entities that have been acquired in the current knowledge subgraph. This indicates the number of entities within the set; a higher loss value signifies more missing knowledge, which guides the targeted construction of query sub-targets. A graph retrieval tool is invoked to perform queries within the adaptive graph, resulting in newly added local knowledge subgraphs, and the acquired entity set is updated synchronously. With the set of relationships The structured semantic information corresponding to the newly added knowledge subgraph is fused into the current inference context text sequence, and the knowledge completeness calculation formula is updated. ,in All target knowledge entities required for the task. This refers to all the relationships that should exist between the target entities. For the currently acquired entity relationships, , Preset weighting coefficients and satisfying The influence weights of entity completeness and relation completeness on the completeness are adjusted respectively; after the information is updated, the next round of reasoning loop is entered, and the knowledge completeness judgment is carried out again.
[0060] In one possible implementation, step S310 further includes: Step S311: The action space includes thinking actions, stopping actions, query generation actions, graph retrieval actions, and strategy output actions.
[0061] Step S312: The state space contains the current inference context text sequence, the retrieved knowledge subgraph structure, and the completeness score of the current policy scheme.
[0062] Specifically, the action space includes thinking actions, stopping actions, query generation actions, graph retrieval actions, and strategy output actions. Among them, the thinking action is used to analyze the reasoning context, identify the current knowledge gaps, and conduct logical deductions to provide a basis for subsequent decisions. The stopping action is used to directly terminate the current task process and abandon the generation of SCADA strategies. The query generation action is used to construct graph query sub-targets based on the identified knowledge gap elements. The graph retrieval action is used to call retrieval tools, perform queries within the adaptive graph, and obtain newly added knowledge sub-graphs. The strategy output action refers to integrating all information after the knowledge completeness target is met, structurally generating, and issuing SCADA control strategies. In each iteration, the agent selects one of the above five types of actions to execute based on the real-time state characteristics, driving the continuous operation of the retrieval and reasoning process.
[0063] The state space is the set of all lawful decision states of the agent, uniformly composed of three core information types: the current inference context text sequence, the retrieved knowledge subgraph structure, and the current strategy completeness score. It comprehensively represents the real-time decision state during the retrieval and inference iteration process. The current inference context text sequence records the task objective, reasoned logic, historical query content, and intermediate decision information for each iteration, providing global semantic temporal features of the inference process. The retrieved knowledge subgraph structure is a local topological structure accumulated from multiple rounds of graph retrieval, containing effective knowledge entities, inter-entity relationships, and topological attributes, providing structured knowledge features to support strategy inference. The current strategy completeness score is a quantitative indicator calculated based on entity coverage and relationship matching rate, used to numerically represent the knowledge completeness of the current semi-finished strategy. These three types of information complement each other, uniquely determining the agent's current decision state from three dimensions: semantic inference progress, structured knowledge reserves, and strategy formation quality. This serves as a unified input for action selection, reward calculation, and network training, supporting the agent in completing adaptive retrieval, iterative inference, and strategy output decisions.
[0064] A field test was conducted on a SCADA load control scenario in a power distribution substation. Multiple control schemes were set up, and key indicators such as knowledge retrieval hit rate, policy physical constraint satisfaction, average inference steps, number of redundant instructions, and policy generation time were statistically analyzed. The test results are shown in Table 1. Table 1 Comparison of SCADA strategy generation performance under different schemes Comparison Option 1 Static knowledge graph 72.3 75.1 18.6 4.2 682 Comparison Option 2 Simple Atlas Update 79.5 80.4 15.3 3.1 564 Embodiments of the present invention Adaptive Dynamic Knowledge Graph 92.6 94.8 8.2 0.7 316 As shown in Table 1, Comparison Scheme 1, which uses a static knowledge graph, has the lowest performance in both metrics, while Comparison Scheme 2 only shows a slight improvement. This embodiment relies on dynamic knowledge graph reconstruction to form an adaptive knowledge graph and optimizes retrieval accuracy and strategy quality through a multi-dimensional reward function, significantly improving retrieval hit rate and constraint satisfaction. The results are plotted based on the experimental data in Table 1. Figure 3 The bar chart shown is composed of... Figure 3 It can be intuitively seen that, compared to Scheme 1, Scheme 2 are limited by static knowledge graphs or simple graph updates, resulting in low retrieval hit rates and constraint satisfaction. This invention obtains an adaptive graph through dynamic graph configuration reconstruction and achieves coordinated optimization of retrieval accuracy and policy quality by relying on a multi-dimensional reward function containing gain and penalty terms. It has higher retrieval hit rates and constraint satisfaction, verifying that this method has better overall performance in the scenario of automatic SCADA policy generation.
[0065] It should be noted that the order of the embodiments described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. Specific embodiments of this specification have been described above. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0066] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0067] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A dynamic knowledge graph-driven SCADA strategy generation method, characterized in that, The method includes: Upload the SCADA policy generation task, and based on the graph policy network in the generated intelligent body, perform a dynamic configuration process reconstruction with the graph configuration reward function as a constraint to obtain an adaptive graph. Based on the adaptive graph, a retrieval-inference strategy is determined by defining a multidimensional reward function system based on gain and penalty terms, with the goal of autonomously balancing retrieval accuracy, strategy quality, and inference efficiency. In the generated agent, the action space and state space of the generated agent are defined based on the adaptive graph and the retrieval-reasoning strategy, and cyclic reasoning decision-making is performed to generate the task SCADA strategy. The task SCADA strategy is visualized on the terminal display interface.
2. The dynamic knowledge graph-driven SCADA strategy generation method as described in claim 1, characterized in that, Generate a graph policy network within the intelligent entity, including: A graph configuration reward function adapted to SCADA policy generation tasks is constructed, wherein the task quality index composed of policy satisfaction, control accuracy, and physical constraint coverage is used as the optimization objective, and the graph configuration reward function is obtained by weighted fusion of the first reward component and the second reward component; The dynamic configuration process of knowledge graphs is modeled as a Markov decision process; Based on the graph configuration reward function and Markov decision process, a graph policy network is constructed within the generating intelligent body. The input of the graph policy network is the current graph structure features and the semantic description of the SCADA policy generation task, and the output is the probability distribution of the next graph configuration action.
3. The dynamic knowledge graph-driven SCADA strategy generation method as described in claim 2, characterized in that, The first reward component is quantified by the reciprocal of the average shortest reachable path length of the knowledge entities required for strategy generation in the graph, which represents the sufficiency of the graph's knowledge coverage of the strategy generation task. The shorter the average shortest reachable path of knowledge entities in the graph, the higher the value of the first reward component. The second reward component is quantified by the average number of graph traversal jumps required for key knowledge retrieval, representing the knowledge retrieval efficiency of the graph structure. The fewer the average number of traversal jumps for knowledge retrieval, the higher the value of the second reward component.
4. The dynamic knowledge graph-driven SCADA strategy generation method as described in claim 3, characterized in that, The dynamic configuration process takes the graph structure feature vector of the current graph as the system state and adds new knowledge entity nodes and new relationship edges between entities as executable configuration actions.
5. The dynamic knowledge graph-driven SCADA strategy generation method as described in claim 4, characterized in that, After each round of SCADA strategy generation and execution task is completed, the graph configuration reward function is updated based on the actual execution result of the strategy, and the instantaneous reward value corresponding to the current graph configuration is calculated. The instant reward value, graph state, and configuration action trajectory of the current round are synchronously stored in the experience replay pool, triggering the graph structure fine-tuning of the graph policy network to obtain an adaptive graph.
6. The dynamic knowledge graph-driven SCADA strategy generation method as described in claim 1, characterized in that, Determine the retrieval-reasoning strategy, including: Using the adaptive graph as a shared interactive environment for retrieval and reasoning, a multi-dimensional reward function system containing gain and penalty terms is constructed; Based on the multidimensional reward function, multiple rounds of retrieval are performed in the adaptive graph, and the accumulated knowledge subgraphs are encoded into graph embedding vectors, which are then concatenated with the current inference context vector to form state features; Based on the state characteristics, a retrieval-reasoning strategy is determined.
7. The dynamic knowledge graph-driven SCADA strategy generation method as described in claim 6, characterized in that, The gain term includes a first gain component and a second gain component. The first gain component is the hit rate obtained by the generated intermediate query in the adaptive graph execution graph retrieval, and the second gain component is the degree to which the strategy scheme output in the current inference stage satisfies the physical constraints of the SCADA system. The penalty item includes a first penalty component and a second penalty component. The first penalty component is the number of steps that exceed a preset step threshold during the retrieval-reasoning process. The second penalty component is the number of redundant and repetitive control instructions within the policy scheme generated in the current inference phase.
8. The dynamic knowledge graph-driven SCADA strategy generation method as described in claim 1, characterized in that, Perform cyclical reasoning decisions to generate task SCADA strategies, including: Using the aforementioned retrieval-reasoning strategy as the execution strategy, the action space and state space of the generated agent are defined; In each round of reasoning, the current reasoning context text sequence and the retrieved knowledge subgraph structure are used as inputs to evaluate whether the knowledge completeness threshold required for policy generation is met. If the conditions are met, the policy output action is executed, the inference loop is terminated, and the task SCADA policy is output.
9. The dynamic knowledge graph-driven SCADA strategy generation method as described in claim 8, characterized in that, If the conditions are not met, then execute the query generation action; A new graph query sub-target is generated based on the current inference context text sequence. The graph retrieval tool is called to perform a retrieval in the adaptive graph to obtain a new knowledge sub-graph. After being integrated into the current inference context text sequence, it enters the next iteration.
10. The dynamic knowledge graph-driven SCADA strategy generation method as described in claim 8, characterized in that, The action space includes thinking actions, stopping actions, query generation actions, graph retrieval actions, and strategy output actions; The state space contains the current inference context text sequence, the retrieved knowledge subgraph structure, and the completeness score of the current policy scheme.