Energy storage power station operation and maintenance plan dynamic generation method and device
Patent Information
- Application Number
- CN202611011849.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明提供了一种储能电站运维计划动态生成方法及装置,以解决无法在多告警并发、资源受限的场景下自动生成全局优化的运维计划的问题
将实时告警数据中与同一根因节点在设备拓扑图中距离不超过一跳的告警聚合作为联合运维任务,并将实时告警数据中无法关联至任一根因节点的告警作为独立运维任务,以构建运维任务列表。
Smart Images

Figure CN122840667A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy storage power station operation and maintenance technology, specifically to a method and apparatus for dynamically generating energy storage power station operation and maintenance plans. Background Technology
[0002] With the large-scale development of the energy storage industry, a single 100-megawatt energy storage power station typically contains hundreds of battery clusters, thousands of battery modules and tens of thousands of cells. The number of equipment nodes is huge, the operating conditions are complex and variable, and the stability of its operation is highly dependent on effective operation and maintenance management.
[0003] Existing solutions achieve task scheduling management by analyzing the execution priority of various operation and maintenance tasks and assessing the load status. When equipment anomalies are detected, work orders are automatically generated, and intelligent work orders are dispatched based on personnel skills and geographical location. Some solutions also introduce alarm correlation analysis technology, which uses the Apriori association rule mining algorithm to perform correlation analysis on alarm data and output alarm correlation graphs to help operation and maintenance personnel quickly locate the root cause.
[0004] However, existing solutions typically split multiple alarms into multiple independent work orders for processing in alarm correlation scenarios, resulting in dispersed operation and maintenance resources and low processing efficiency. The alarm correlation analysis results also mostly remain at the information display level and fail to be transformed into intelligent generation and scheduling decisions for operation and maintenance tasks. In addition, in actual operation and maintenance, resource constraints such as personnel skill division and spare parts inventory are handled passively after work orders are generated, rather than being actively incorporated into the optimization in the decision-making process. This makes it impossible to proactively optimize the handling sequence based on historical experience to reduce future risks. Summary of the Invention
[0005] This invention provides a method and apparatus for dynamically generating operation and maintenance plans for energy storage power stations, in order to solve the problem of not being able to automatically generate globally optimized operation and maintenance plans in scenarios with multiple concurrent alarms and limited resources.
[0006] In a first aspect, the present invention provides a method for dynamically generating operation and maintenance plans for energy storage power stations, including: Collect real-time alarm data and real-time resource information of the energy storage power station, and construct the equipment topology diagram of the energy storage power station; Graph neural networks are used to locate root causes in device topology and real-time alarm data to obtain a list of maintenance tasks. Construct a state vector based on real-time resource information and the list of operation and maintenance tasks; The state vector is input into the pre-trained policy network to determine the scheduling action, and an operation and maintenance plan is generated based on the scheduling action.
[0007] Beneficial effects: It realizes the automatic identification of alarm root causes and the automatic aggregation of related alarms, avoiding the dispersion of operation and maintenance resources. At the same time, it improves the scientificity and adaptability of operation and maintenance plans by dynamically optimizing scheduling decisions under resource constraints through deep reinforcement learning, laying the foundation for subsequent long-term self-learning optimization.
[0008] In one alternative implementation, constructing the device topology diagram of the energy storage power station includes: Using the equipment in the energy storage power station as nodes and the connection relationships between the equipment in the energy storage power station as edges, construct the equipment topology graph of the energy storage power station. The equipment includes battery clusters, converter units, battery management system controllers, environmental control units, and fire detectors. The equipment connection relationships include electrical connection edges, communication dependency edges, and spatial proximity edges.
[0009] Beneficial effects: Improved accuracy and reliability of alarm root cause localization. When the number of nodes is small, a fully connected graph structure can be used as an alternative, enhancing the adaptability of the technical solution.
[0010] In one optional implementation, a graph neural network is used to locate root causes in the device topology and real-time alarm data to obtain a list of maintenance tasks, including: Construct a feature vector for each node, which includes device type code, operating status code, alarm count, health score, and alarm level code; Based on the nodes in the device topology graph and the feature vectors corresponding to each node, a device feature matrix is constructed. The device feature matrix and the adjacency matrix of the device feature matrix are subjected to two-layer graph convolution to obtain the root cause probability and non-root cause probability of each node. Input the root cause probability and non-root cause probability of each node into the softmax function to obtain the alarm root cause probability of each node. Nodes whose alarm root cause probability is greater than a preset probability threshold are identified as root cause nodes. Alarms in real-time alarm data that are no more than one hop away from the same root node in the device topology graph are aggregated as joint operation and maintenance tasks, while alarms in real-time alarm data that cannot be associated with any root node are treated as independent operation and maintenance tasks, in order to construct an operation and maintenance task list.
[0011] Beneficial effects: It enables causal root cause localization of alarms and automatic aggregation of related alarms. The use of two-layer graph convolution allows each node to aggregate information from neighboring nodes within two hops, which perfectly matches the physical characteristic of energy storage power stations where fault propagation does not exceed two hops, thus avoiding misjudgments caused by excessive information diffusion.
[0012] In one optional implementation, a state vector is constructed based on real-time resource information and an operation and maintenance task list, including: Based on the information of each operation and maintenance task in the operation and maintenance task list, a task queue vector is constructed; based on real-time resource information, a resource status vector is constructed; and based on the health information of the equipment in the energy storage power station, an equipment health vector is constructed. The task queue vector, resource status vector, and device health vector are concatenated to obtain the status vector.
[0013] Beneficial effects: The three-dimensional splicing structure of the state vector enables the deep reinforcement learning model to comprehensively consider the urgency of the task, the availability of resources, and the long-term health trend of the equipment when making decisions, avoiding the limitation of passively matching resources only after the work order is generated in the existing technology.
[0014] In one alternative implementation, the method further includes: Based on the alarm type of each alarm in the real-time alarm data, determine the risk function corresponding to each alarm; The risk value corresponding to each alarm is calculated based on the risk function, which is used to construct the reward function in the training process of the policy network. The risk functions include linear growth type, exponential growth type, S-shaped type and step type.
[0015] Beneficial effects: This technique enables the model to prioritize tasks with rapidly increasing risks during decision-making, achieving adaptive priority under dynamic risk evolution.
[0016] In one alternative implementation, the state vector is input into a pre-trained policy network to determine the scheduling action, including: The state vector is input into the hidden layer of the policy network for function activation, and normalized using the softmax function of the output layer of the policy network to obtain the probability distribution of selecting each scheduling action under the state vector. The scheduling actions include assigning maintenance tasks to designated personnel, rearranging the task queue vector, merging multiple joint maintenance tasks, delaying the processing of low-priority maintenance tasks, and requesting the allocation of spare parts. Based on the probability distribution, determine the scheduling action at the current moment.
[0017] Beneficial effects: The policy network is a three-layer fully connected neural network that maps the state vector to the selection probability of each action through forward propagation. This enables it to automatically learn the optimal task execution order and resource matching scheme under the current resource state. When multiple emergency alarms occur concurrently and resources are limited, it can intelligently arrange the processing order, shorten the average fault recovery time, and reduce the idle rate of operation and maintenance personnel.
[0018] In one alternative implementation, the method further includes: If there are alarms in the real-time alarm data that meet the preset emergency conditions, a mandatory work order will be generated. The preset emergency conditions include the alarm level being urgent and the alarm type being exponentially increasing, the battery temperature exceeding the preset protection threshold, and the fire alarm being issued by the fire protection system. If no alarms that meet the preset emergency conditions are found in the real-time alarm data, the operation and maintenance plan corresponding to the scheduling action will be pushed to the terminal. In response to the terminal's confirmation or adjustment instructions, the confirmed or adjusted operation and maintenance plan will be used as the final execution plan.
[0019] Beneficial effects: It addresses the special requirements of the energy storage industry for high safety, while the manual confirmation mode reduces the barrier for operation and maintenance personnel to accept AI decision-making, which is conducive to the promotion and application of technical solutions in actual power plants.
[0020] In one alternative implementation, the method further includes: The mandatory work order and the maintenance plan confirmed by the terminal are converted into maintenance work orders. The maintenance work order includes the work order number, task type, task description, response time limit, assigned personnel information and required spare parts list. Push maintenance work orders to the corresponding terminals and record the time of acceptance and completion of the maintenance work orders.
[0021] Beneficial effects: It achieves a seamless connection from decision-making to execution, ensuring that operation and maintenance plans can be quickly and accurately transformed into executable operation instructions, thereby improving the closed-loop efficiency of operation and maintenance management.
[0022] In one alternative implementation, the method further includes: Record the actual execution data of the scheduling action. The actual execution data includes the actual completion time of the task, the alarm information newly generated during the task execution, and the recurrence of similar alarms within a preset time period after the task is completed. The actual execution data within the preset period is added to the training set, and the policy network is updated using the near-end policy optimization algorithm to obtain the updated policy network. The updated policy network and the original policy network are tested on historical data. If the overall performance score of the updated policy network is higher than that of the original policy network, the updated policy network replaces the original policy network.
[0023] Beneficial effects: It overcomes the limitation of existing technologies in actively learning forward-looking strategies from historical experience. Furthermore, by employing a proximal policy optimization algorithm to prune the objective function, the model update process remains stable, ensuring the reliability of engineering deployment.
[0024] Secondly, the present invention provides a device for dynamically generating operation and maintenance plans for energy storage power stations, comprising: The data acquisition module is used to collect real-time alarm data and real-time resource information of the energy storage power station, and to construct the equipment topology diagram of the energy storage power station. The positioning module is used to locate root cause nodes by using graph neural networks to analyze the device topology and real-time alarm data, and obtain a list of maintenance tasks. The building module is used to construct state vectors based on real-time resource information and the list of operation and maintenance tasks; The determination module inputs the state vector into the pre-trained policy network, determines the scheduling action, and generates an operation and maintenance plan based on the scheduling action.
[0025] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform a method for dynamically generating an energy storage power station operation and maintenance plan as described in the first aspect or any corresponding embodiment.
[0026] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions, which are used to cause a computer to execute a method for dynamically generating an energy storage power station operation and maintenance plan according to the first aspect or any corresponding embodiment described above.
[0027] Fifthly, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute a method for dynamically generating an energy storage power station operation and maintenance plan according to the first aspect above or any corresponding embodiment. Attached Figure Description
[0028] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0029] Figure 1 This is a flowchart illustrating a method for dynamically generating an operation and maintenance plan for an energy storage power station according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating a method for constructing a device topology diagram according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating a method for constructing an operation and maintenance task list according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating a method for obtaining a state vector according to an embodiment of the present invention; Figure 5This is a flowchart illustrating a method for constructing a reward function according to an embodiment of the present invention; Figure 6 This is a flowchart illustrating a method for determining a scheduling action at the current moment according to an embodiment of the present invention; Figure 7 This is an architecture diagram of a method for dynamically generating operation and maintenance plans for energy storage power stations according to an embodiment of the present invention; Figure 8 This is a flowchart illustrating an alarm aggregation method using a graph neural network according to an embodiment of the present invention. Figure 9 This is a structural block diagram of a dynamic generation device for operation and maintenance plans of an energy storage power station according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0031] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0032] In the context of alarm association scenarios in energy storage power stations, multiple alarms are split into multiple independent work orders for separate processing, resulting in dispersed operation and maintenance resources. The alarm association analysis results are only displayed at the information level and fail to be transformed into intelligent generation and scheduling decisions for operation and maintenance tasks. Resource constraints in actual operation and maintenance are handled by passive matching after work order generation rather than being actively incorporated into the optimization in the decision-making process. Furthermore, it is impossible to proactively optimize the handling sequence based on historical experience to reduce future risks.
[0033] This invention provides a structured graph data foundation for alarm root cause localization by collecting real-time alarm data and resource information from energy storage power stations and constructing an equipment topology map of the power stations. It utilizes a graph neural network to locate root cause nodes in the equipment topology map and real-time alarm data, obtaining a list of maintenance tasks. Alarms associated with the same root cause node are aggregated into joint maintenance tasks, while alarms not associated with a root cause node are treated as independent maintenance tasks. This allows for causal aggregation and task integration of multiple alarms during root cause node localization, fundamentally preventing maintenance resources from being scattered across redundant alarm processing. A state vector is constructed based on real-time resource information and the maintenance task list, enabling multi-dimensional resource constraints such as personnel skills and spare parts inventory to be proactively incorporated into scheduling optimization during the decision-making stage, rather than passively matched after work orders are generated. By inputting the state vector into a pre-trained policy network, scheduling actions are determined, and maintenance plans are generated based on these actions. This allows for the learning of long-term optimization strategies from historical data to drive current scheduling decisions, achieving intelligent aggregation of maintenance tasks under causal relationships of multiple alarms and dynamic scheduling decisions under resource constraints without the need for manually preset rules or fixed priorities.
[0034] According to an embodiment of the present invention, a method for dynamically generating operation and maintenance plans for energy storage power stations is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0035] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which can be used in the aforementioned electronic devices. Figure 1 This is a flowchart of a method for dynamically generating an operation and maintenance plan for an energy storage power station according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Collect real-time alarm data and real-time resource information of the energy storage power station, and construct the equipment topology diagram of the energy storage power station.
[0036] Specifically, real-time alarm data refers to abnormal status signals generated by each subsystem of the energy storage power station during operation. Each alarm data entry must include at least the alarm device identifier, alarm type, alarm level, and alarm occurrence time. Real-time resource information refers to the available maintenance manpower and material resources at the current moment, including the skill tags, busy / idle status of each on-duty maintenance personnel, and spare parts inventory list. The device topology graph is an undirected or directed graph constructed using the devices in the energy storage power station as nodes and the connections between devices as edges. , where the node set Includes battery clusters, converter units, battery management system controllers, environmental control units, and fire detectors, side collections This includes electrical connection edges to reflect the electrical wiring relationships between devices, communication dependency edges to reflect the communication protocols and data flow relationships between devices, and spatial proximity edges to reflect the physical proximity relationships between devices.
[0037] Real-time alarm data and resource information of the energy storage power station are collected in real time with a configurable sampling period of 1 to 10 seconds. The collected data is then deduplicated and timestamped before being stored in a real-time database. At the same time, a device topology diagram is constructed based on the device connection relationships of the energy storage power station.
[0038] Step S102: Use graph neural networks to locate root cause nodes in the device topology diagram and real-time alarm data to obtain a list of operation and maintenance tasks.
[0039] Specifically, a graph neural network (GNN) is a neural network model that performs message passing and feature learning on graph-structured data. Root cause localization refers to using a GNN to calculate the probability that each device node in the device topology graph is the root cause of an alarm. and the probability is greater than the preset probability threshold. The node is identified as the root cause device that triggered the current alarm. The operation and maintenance task list refers to the collection of operation and maintenance jobs to be executed generated from the root cause node location results. It includes joint operation and maintenance tasks formed by multiple alarms caused by the same root cause node, as well as independent operation and maintenance tasks formed by alarms that cannot be associated with any root cause node.
[0040] Based on device topology diagram Using real-time alarm data as input, a graph neural network is used for message passing to calculate the data for each device node in the device topology graph. Probability as the root cause of an alarm .Will > The node is identified as the root cause node. For all unprocessed real-time alarm data at the current moment, if the distance between the located device and the same root cause node in the device topology graph is no more than one hop, these alarms are aggregated into a joint operation and maintenance task. The joint operation and maintenance task includes the root cause device identifier, the list of associated alarms, the task priority, the estimated processing time, the required skills, and the response time limit; alarms that cannot be associated with any root cause node constitute independent operation and maintenance tasks. All joint operation and maintenance tasks and independent operation and maintenance tasks together constitute the operation and maintenance task list.
[0041] Step S103: Construct a state vector based on real-time resource information and the list of operation and maintenance tasks.
[0042] Specifically, the state vector refers to a multi-dimensional vector obtained by numerically encoding the task feature information in the operation and maintenance task list and the resource status in the real-time resource information, which is used as the input to the policy network.
[0043] The task feature information of each operation and maintenance task in the operation and maintenance task list is integrated with the resource status information in the real-time resource information, and then constructed into a status vector after numerical encoding. .
[0044] Step S104: Input the state vector into the pre-trained policy network, determine the scheduling action, and generate an operation and maintenance plan based on the scheduling action.
[0045] Specifically, the policy network refers to a pre-trained neural network model used to output scheduling decisions based on the current state vector. A scheduling action refers to the optimal operational instruction selected from a predefined action space for the current moment. An operational plan refers to an executable operational scheme generated based on the scheduling actions.
[0046] The state vector The input is a pre-trained policy network, which, after forward propagation, outputs the probability distribution of each action in the action space. The scheduling action at the current time is then determined based on the probability distribution. According to the scheduling action Generate the corresponding operation and maintenance plan and issue it for execution.
[0047] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations. It realizes automatic identification of alarm root causes and automatic aggregation of related alarms, avoiding the dispersion of operation and maintenance resources. At the same time, it improves the scientificity and adaptability of operation and maintenance plans by dynamically optimizing scheduling decisions under resource constraints through deep reinforcement learning, laying the foundation for subsequent long-term self-learning optimization.
[0048] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which can be used for the aforementioned electronic devices, etc. Figure 2 This is a flowchart illustrating a method for constructing a device topology diagram according to an embodiment of the present invention, as shown below. Figure 2 As shown, the process includes the following steps: Step S201: Using the equipment in the energy storage power station as nodes and the connection relationships of the equipment in the energy storage power station as edges, construct the equipment topology graph of the energy storage power station. The equipment includes battery clusters, converter units, battery management system controllers, environmental control units, and fire detectors. The equipment connection relationships include electrical connection edges, communication dependency edges, and spatial proximity edges.
[0049] Specifically, in an energy storage power station, the electrical connection edges in the equipment connection relationships refer to edges established based on the electrical wiring relationships in the equipment ledger information, used to reflect the electrical connections between equipment. Communication dependency edges refer to edges established based on the communication protocols and data flow directions in the configuration file, used to reflect the communication dependencies between equipment. Spatial proximity edges refer to edges established based on the physical location coordinates in the equipment ledger information, used to reflect the spatial proximity relationship between equipment; a spatial proximity edge is established when the physical distance between two equipment nodes is less than a preset distance threshold.
[0050] Based on the energy storage power station's documentation, determine the nodes and their attribute information. Then, establish electrical connection edges between each device node according to the electrical wiring relationships. Establish communication dependency edges between each device node according to the communication protocols and data flow directions in the configuration file. Establish spatial proximity edges between each device node according to the physical location coordinates in the device ledger information. For small power stations with fewer than 50 nodes, this explicit topology construction step can be omitted, and the default fully connected graph structure can be used subsequently.
[0051] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which improves the accuracy and reliability of alarm root cause localization. When the number of nodes is small, a fully connected graph structure can be used as an alternative, enhancing the adaptability of the technical solution.
[0052] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which can be used for the aforementioned electronic devices, etc. Figure 3 This is a flowchart illustrating a method for constructing an operation and maintenance task list according to an embodiment of the present invention, as shown below. Figure 3 As shown, the process includes the following steps: Step S301: Construct a feature vector for each node. The feature vector includes device type code, operating status code, alarm count, health score, and alarm level code.
[0053] Specifically, a feature vector refers to a numerical descriptive vector constructed for each device node in the device topology graph. This is used to characterize the status information of the device node in multiple dimensions, including device type code to characterize the type of device, operating status code to characterize the current working status of the device, alarm count to quantify the recent abnormal activity level of the device, health score to characterize the degree of health degradation of the device, and alarm level code to characterize the severity level of the alarm currently issued by the device.
[0054] Device topology diagram Each device node in Constructing feature vectors The feature vector includes device type code, operating status code, alarm count, health score, and alarm level code.
[0055] Step S302: Construct a device feature matrix based on the nodes in the device topology graph and the feature vector corresponding to each node. Perform a two-layer graph convolution on the device feature matrix and the adjacency matrix of the device feature matrix to obtain the root cause probability and non-root cause probability of each node.
[0056] Specifically, the device feature matrix refers to the matrix obtained by stacking the feature vectors of all device nodes in the device topology graph row by row. ,in This represents the total number of device nodes. The dimension is denoted by , where each row corresponds to a feature vector of a device node. The adjacency matrix is a matrix used to represent the connection relationships between device nodes in a device topology graph. Its dimension is |V|×|V|. Two-layer graph convolution refers to performing message passing and feature transformation on the input features in two steps. The first layer graph convolution maps the input features to 64-dimensional intermediate features, and the second layer graph convolution maps the intermediate features to 2-dimensional output features. The 2-dimensional output features correspond to the root cause feature value and non-root cause feature value of each device node, respectively.
[0057] Device topology diagram The feature vectors corresponding to all nodes in Stack the rows to construct the device feature matrix. A two-layer graph convolutional network is used for message passing, and its propagation rule is shown in formula (1): (1) in, For the first Layer node feature matrix For the first Layer node feature matrix, rows correspond to device nodes, columns correspond to feature dimensions; , It is an adjacency matrix. It is the identity matrix; for The degree matrix, diagonal elements ; For the first Layer-trainable weight matrix; The ReLU activation function is used. The first layer of graphical convolution converts the input features... Mapped to The second layer of graph convolution will Mapped to . Each device node in the system corresponds to two feature values, which serve as the feature value for the root cause of the alarm and the feature value for the non-root cause of the alarm.
[0058] In an optional implementation, the graph convolutional network in step S302 can be replaced by a graph attention network. The graph attention network assigns different weight coefficients to different neighbor nodes through an attention mechanism, enabling more accurate root cause localization results in scenarios with complex device topology and significant differences in edge importance. Alternatively, the graph convolutional network can be replaced by a graph isomorphic network, which has stronger expressive power and is suitable for scenarios requiring fine differentiation of different graph structures. All of the above alternatives take the device topology graph as input, output the probability distribution of root cause nodes, and aggregate related alarms into joint operation and maintenance tasks.
[0059] Step S303: Input the root cause probability and non-root cause probability of each node into the softmax function to obtain the alarm root cause probability of each node.
[0060] Specifically, the root cause feature value refers to the output of the second-layer graph convolution. Each device node corresponds to the first of two output values. This is used to characterize the original response strength of the node as the root cause of the alarm. Non-root cause eigenvalues refer to the output of the second-layer graph convolution. Each device node corresponds to the second of two output values. This is used to characterize the original response strength of a node that is not the root cause of an alarm. The softmax function is a preset normalization function used to convert root cause and non-root cause feature values into a normalized probability distribution. The alarm root cause probability refers to the probability value obtained after normalization by the softmax function that the device node is the root cause of the alarm. Its value ranges from 0 to 1, and the sum of the root cause probability and non-root cause probability of alarms for the same node is 1.
[0061] Output of the second layer graph convolution Apply the softmax function to obtain the result for each device node. Root cause probability of alarms The specific calculation formula is shown in formula (2): (2) in, For nodes Non-root cause eigenvalues For nodes The root cause eigenvalues are denoted by exp, which represents the natural exponential function, i.e., an exponential function with the mathematical constant e as its base. The softmax function normalizes the root cause eigenvalues and non-root cause eigenvalues of each node into a probability distribution, such that the sum of the root cause probability and the non-root cause probability of the same node is 1.
[0062] Step S304: Nodes with alarm root cause probabilities greater than a preset probability threshold are identified as root cause nodes.
[0063] Specifically, preset probability threshold It is a pre-defined probability threshold used to determine whether a device node is the root cause of an alarm. A root cause node is a device node in the device topology diagram that is identified as the source of the current alarm.
[0064] Probability of root cause of alarm device nodes If a device node is identified as a root cause node, and the probability of the alarm root cause is not greater than a preset probability threshold, then it is determined not to be a root cause node.
[0065] Step S305: Aggregate alarms in the real-time alarm data that are no more than one hop away from the same root node in the device topology graph as joint operation and maintenance tasks, and treat alarms in the real-time alarm data that cannot be associated with any root node as independent operation and maintenance tasks, so as to construct an operation and maintenance task list.
[0066] Specifically, distance in the device topology graph refers to the number of edges in the shortest path between two device nodes. When two device nodes are directly adjacent, the distance is one hop. A distance of no more than one hop means the two device nodes are the same node or directly adjacent. A joint maintenance task refers to a single maintenance work unit formed by aggregating all alarms that are no more than one hop away from the same root node in the device topology graph. An independent maintenance task refers to a maintenance work unit consisting solely of alarms in the real-time alarm data that cannot be associated with any root node.
[0067] For all unprocessed real-time alarms at the current moment, the corresponding location device node for each alarm is obtained. If the distance between the location device node and the same root cause node in the device topology graph is no more than one hop, these alarms are aggregated into a joint operation and maintenance task. The joint operation and maintenance task includes the root cause device identifier, a list of associated alarms, a task priority (the highest level among the associated alarms), an estimated processing time, required skills, and a response time limit. If the location device node of an alarm in the real-time alarm data cannot be associated with any root cause node, or if the root cause probability of the alarm is not greater than a preset probability threshold, then the alarm is constituted as an independent operation and maintenance task. All joint operation and maintenance tasks and independent operation and maintenance tasks together constitute the operation and maintenance task list.
[0068] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, realizing the causal root cause localization of alarms and the automatic aggregation of related alarms. A two-layer graph convolution is used, enabling each node to aggregate information from neighboring nodes within two hops. This perfectly matches the physical characteristic of energy storage power stations where fault propagation does not exceed two hops, avoiding misjudgments caused by excessive information diffusion.
[0069] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which can be used for the aforementioned electronic devices, etc. Figure 4 This is a flowchart illustrating a method for obtaining a state vector according to an embodiment of the present invention, as shown below. Figure 4 As shown, the process includes the following steps: Step S401: Construct a task queue vector based on the information of each operation and maintenance task in the operation and maintenance task list, construct a resource status vector based on real-time resource information, and construct a device health vector based on the health information of the devices in the energy storage power station.
[0070] Specifically, the information for each operation and maintenance task includes the task priority code and the current risk value. The system includes risk growth type coding, estimated processing time, required skill codes, and remaining response time. Real-time resource information includes the number of currently available personnel for each skill type and the inventory of critical spare parts. Health information for equipment in the energy storage power station includes the average health status of all critical equipment and the station-wide alarm frequency over the past 24 hours.
[0071] Extract information for each operation and maintenance task from the operation and maintenance task list, including task priority code and current risk value. The risk growth type, estimated processing time, required skill code, and remaining response time are numerically encoded and constructed into a task queue vector. The number of currently available personnel and the inventory of critical spare parts for each skill type are extracted from real-time resource information and numerically encoded to construct a resource status vector. The average health status of all critical equipment and the station-wide alarm frequency over the past 24 hours are extracted from the health information of equipment in the energy storage power station and numerically encoded to construct an equipment health vector.
[0072] Step S402: Concatenate the task queue vector, resource status vector, and device health vector to obtain the status vector.
[0073] Specifically, concatenation refers to linking the task queue vector, resource state vector, and device health vector sequentially, end to end, to merge them into a single vector of a higher dimension. (State vector) It is a multi-dimensional vector obtained after the concatenation operation, which serves as the current state description for the subsequent policy network input.
[0074] The task queue vector, resource state vector, and device health vector are concatenated sequentially, end to end, to form a single multidimensional vector, which is the state vector. .
[0075] In an optional implementation, the state vector in step S402 above can omit the device health information portion corresponding to the device health vector. For energy storage power stations with limited data acquisition capabilities or scarce storage resources, the device health vector can be omitted, and the state vector can be constructed solely by concatenating the task queue vector and the resource state vector. Although this alternative will slightly reduce the model's ability to perceive long-term device degradation trends, it can still achieve task scheduling optimization under resource constraints and obtain basic decision-making results.
[0076] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations. By using a three-dimensional splicing structure of state vectors, the deep reinforcement learning model can comprehensively consider the urgency of the task, resource availability, and long-term health trends of the equipment when making decisions, thus avoiding the limitation of passively matching resources only after the work order is generated in the prior art.
[0077] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which can be used for the aforementioned electronic devices, etc. Figure 5 This is a flowchart illustrating a method for constructing a reward function according to an embodiment of the present invention, as shown below. Figure 5 As shown, the process includes the following steps: Step S501: Determine the risk function corresponding to each alarm based on the alarm type of each alarm in the real-time alarm data.
[0078] Specifically, alarm type refers to the category label attached to each real-time alarm data entry, used to identify the type of fault to which the alarm belongs, such as abnormal temperature, voltage exceeding limits, communication interruption, fire alarm, etc. The risk function is a mathematical model used to describe the evolution of alarm risk over time. That is, the first The risk value of each alarm changes over time from the moment it occurs. The risk function is a changing function, and different alarm types correspond to different risk evolution patterns. Risk functions include linear growth, exponential growth, S-shaped, and step-shaped functions.
[0079] Based on the alarm type of each alarm in the real-time alarm data, determine the risk function corresponding to each alarm from the preset mapping relationship between alarm types and risk functions. Risk functions include: linear growth type It is applicable to general abnormalities, among which Reflects the average rate at which the corresponding alarm risk develops. , This represents a typical cycle from the emergence of a general anomaly to its development into a major risk, lasting from several days to tens of days; exponential growth type. It is suitable for abnormal temperature and early signs of thermal runaway, among which Reflecting the initial risk value, The rate of risk growth is reflected by fitting historical failure data. It is the base of the natural logarithm; S-shaped growth type This is applicable to equipment aging, where k is the aging rate. This refers to the inflection point of the aging curve, that is, the point at which the battery enters the accelerated aging stage. It is the base of the natural logarithm. The upper limit of the S-shaped growth type, that is, the maximum safety risk value when the equipment ages to its limit; step type It is suitable for threshold-protected anomalies such as overvoltage and undervoltage, among which To determine the risk level corresponding to the alarm. The system automatically matches the corresponding function type and parameters based on the alarm type, including: .
[0080] In an optional implementation, the four risk functions in step S501 above can be replaced with a simplified risk evolution model. For energy storage power stations with relatively simple alarm types and relatively simple risk evolution patterns, only a linear growth model can be used. This approach treats all alarm risks as increasing linearly over time, thus reducing model complexity and parameter tuning difficulty. Alternatively, for scenarios primarily involving temperature-related alarms, an exponential growth model can be used. This highlights the urgency of early signs of thermal runaway. While the simplified scheme described above reduces accuracy, it still achieves the basic goal of prioritizing tasks with rapidly increasing risks.
[0081] Step S502: Calculate the risk value corresponding to each alarm based on the risk function. The risk function is used to construct the reward function in the training process of the policy network. The risk functions include linear growth type, exponential growth type, S-type type and step type.
[0082] Specifically, the risk value refers to the quantified value of the risk accumulated from the time the alarm occurred to the current time, expressed as a risk function. At the present moment Substituting the values, we obtain the result. The reward function is a scoring function used to evaluate the quality of scheduling decisions during the training of the policy network. The risk value is incorporated into the reward function in the form of a cumulative negative reward term, so that scheduling actions that prioritize handling alarms with rapidly increasing risk during training can obtain higher cumulative rewards.
[0083] Based on the risk function corresponding to each alarm , the current moment Substitute into the risk function to calculate the current risk value corresponding to each alarm. The risk function is used to construct the reward function during the training process of the policy network. The reward function includes a negative reward term for accumulated risk, which makes the policy network tend to prioritize the operation and maintenance tasks corresponding to alarms with rapidly increasing risk during the training process.
[0084] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which enables the model to prioritize tasks with rapidly increasing risks when making decisions, thus achieving priority adaptation under dynamic risk evolution.
[0085] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which can be used for the aforementioned electronic devices, etc. Figure 6 This is a flowchart illustrating a method for determining a scheduling action at the current moment according to an embodiment of the present invention, as shown below. Figure 6 As shown, the process includes the following steps: Step S601: The state vector is input into the hidden layer of the policy network for function activation, and normalized using the softmax function of the output layer of the policy network to obtain the probability distribution of selecting each scheduling action under the state vector. The scheduling actions include assigning maintenance tasks to designated personnel, rearranging the task queue vector, merging multiple joint maintenance tasks, delaying the processing of low-priority maintenance tasks, and requesting the allocation of spare parts.
[0086] Specifically, the policy network is a three-layer fully connected neural network pre-trained using a proximal policy optimization algorithm. Its input layer dimension equals the dimension of the state vector, each hidden layer contains 128 neurons, and the activation function is the hyperbolic tangent function. The output layer dimension equals the action space size, and the activation function is the softmax function. Function activation refers to the process of sequentially passing the state vector through the first hidden layer of the policy network, performing a linear transformation and activation with the hyperbolic tangent function to obtain the first hidden feature vector, and then passing it through the second hidden layer, performing the same linear transformation and activation to obtain the second hidden feature vector. The action space refers to the set of all possible scheduling actions, including assigning maintenance tasks to designated personnel. Rearrange the task queue vector Merging multiple joint operation and maintenance tasks Delay processing of low-priority operation and maintenance tasks And request the allocation of spare parts The probability distribution refers to the probability vector formed by the probability values of each scheduled action being selected in the action space after normalization by the softmax function, and the sum of the probabilities of all actions is 1.
[0087] The policy network is responsible for generating the probability distribution of actions based on the current state and determining the scheduling actions. The value network is used to estimate the state value, providing an evaluation benchmark for updating the policy network. The near-end policy optimization algorithm updates the policy network through the clip function, and its optimization objective function uses the clip function to restrict the probability ratio to an interval. To avoid overly large step sizes in a single update, which could lead to training instability, an experience replay buffer is maintained during the update process, and the learning rate is set to one-tenth of the initial value for fine-tuning.
[0088] The state vector Input pre-trained policy network The state vector is calculated via forward propagation through the policy network. First, a linear transformation is performed through the first hidden layer. The first hidden feature vector is obtained by activation with the hyperbolic tangent function. The first hidden feature vector is then subjected to a linear transformation and hyperbolic tangent activation function in the second hidden layer to obtain the second hidden feature vector. Finally, the output layer performs a linear transformation and normalizes the result using a preset softmax function, as shown in formula (3): (3) Specifically, Indicates based on parameters The policy network outputs the state vector in the current state. Select Action The probability distribution, the action here It refers to a specific scheduling action, such as discharging at a 1C rate to the cutoff voltage; The softmax activation function is used to convert the network output into a probability distribution, ensuring that the sum of the probabilities of all actions is 1. is the hyperbolic tangent function, which is the activation function of the hidden layer; , These are the trainable weight matrices for the first and second hidden layers of the policy network, respectively. This is the trainable weight matrix for the output layer of the policy network; , These are the bias terms for the first and second hidden layers of the policy network, respectively. This is the bias term for the output layer of the policy network.
[0089] Obtain the state vector The probability distribution of selecting each scheduling action The softmax function ensures that the sum of the probabilities of all scheduled actions is 1.
[0090] Step S602: Determine the scheduling action for the current moment based on the probability distribution.
[0091] Specifically, the scheduling action at the current moment refers to a specific operation and maintenance instruction selected from the action space at the current decision moment.
[0092] According to probability distribution The scheduling action at the current moment is determined from the action space using a preset sampling strategy. Determine the scheduling action at the current moment. Then, the scheduling action is executed to drive the operation and maintenance of the energy storage power station.
[0093] In an alternative implementation, the scheme in step S602 where the scheduling actions are entirely output by the deep reinforcement learning model can be replaced by a decision architecture that combines a rule engine and deep reinforcement learning. Specifically, a confidence evaluation module is set up. When the confidence of the action output by the policy network is lower than a preset threshold, the system automatically switches to the decision engine based on expert rules to generate an operation and maintenance plan; when the confidence is higher than the preset threshold, the output of the policy network is used. This alternative solution balances intelligence and robustness, and can rely on the rule engine to ensure basic decision-making capabilities when the model encounters abnormal scenarios that have not appeared in the training set.
[0094] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations. The strategy network is a three-layer fully connected neural network. Through forward propagation, the state vector is mapped to the selection probability of each action, which enables the system to automatically learn the optimal task execution order and resource matching scheme under the current resource status. When multiple emergency alarms occur concurrently and resources are limited, the system can intelligently arrange the processing order, shorten the average fault recovery time, and reduce the idle rate of operation and maintenance personnel.
[0095] Specifically, the method for dynamically generating an energy storage power station operation and maintenance plan provided in this embodiment of the invention further includes: Step S701: If there are alarms in the real-time alarm data that meet the preset emergency conditions, a mandatory work order is generated. The preset emergency conditions include an alarm level of emergency and an alarm type of exponential growth, battery temperature exceeding a preset protection threshold, and fire alarm issued by the fire protection system.
[0096] Specifically, preset emergency conditions refer to a set of pre-defined alarm judgment rules used to trigger the generation of mandatory work orders. Preset emergency conditions include alarm level being urgent and alarm type being exponentially increasing, battery temperature exceeding a preset protection threshold, and fire alarm issued by the fire protection system. Mandatory work orders refer to emergency maintenance work orders generated without policy network recommendation.
[0097] Safety verification is performed on the scheduling actions. If any alarm in the real-time alarm data meets any of the preset emergency conditions, i.e., the alarm level is emergency and the alarm type is exponential growth, the battery temperature exceeds the preset protection threshold, or the fire protection system issues a fire alarm, a mandatory work order is generated directly, ignoring the scheduling actions output by the deep reinforcement learning decision module.
[0098] Step S702: If there are no alarms in the real-time alarm data that meet the preset emergency conditions, the operation and maintenance plan corresponding to the scheduling action is pushed to the terminal. In response to the terminal's confirmation or adjustment instruction, the confirmed or adjusted operation and maintenance plan is used as the final execution plan.
[0099] Specifically, "terminal" refers to the terminal device used by the operations manager, including web terminals or mobile terminals. "Operations plan" refers to the operations recommendation plan generated based on scheduling actions, displayed in a structured list format.
[0100] If no alarms meeting the preset emergency conditions are found in the real-time alarm data, the corresponding operation and maintenance plan will be pushed to the operation and maintenance supervisor's terminal in the form of a structured list for confirmation or adjustment. In response to the confirmation or adjustment command returned by the terminal, the confirmed or adjusted operation and maintenance plan will be used as the final execution plan.
[0101] In an optional implementation, the manual confirmation step in step S702 can be omitted. For energy storage power stations aiming for unattended or minimally staffed operation, provided that thorough verification and strict safety fallback conditions are configured, the scheduling actions output by the deep reinforcement learning decision module can be directly issued for execution. In this alternative solution, the safety fallback module plays a more critical role; all alarms meeting preset emergency conditions are still required to generate work orders, while other scenarios are decided automatically by the model. This alternative solution further enhances the level of operation and maintenance automation and is suitable for scenarios with extremely high requirements for response timeliness and where manual intervention is costly.
[0102] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which addresses the special requirements of the energy storage industry for high security. The manual confirmation mode reduces the barrier for operation and maintenance personnel to accept AI decision-making, which is conducive to the promotion and application of the technical solution in actual power stations.
[0103] In some optional implementations, the step S701 above is followed by: Step S7011: Convert the mandatory work order and the maintenance plan confirmed by the terminal into a maintenance work order. The maintenance work order includes the work order number, task type, task description, response time limit, assigned personnel information, and a list of required spare parts.
[0104] Specifically, an operation and maintenance (O&M) work order refers to an executable O&M work document converted from a mandatory work order or an O&M plan confirmed by the terminal. The work order number is a unique identifier for the O&M work order. The task type refers to the type of O&M work, including equipment repair, equipment inspection, and spare parts replacement. The task description is a written description of the specific content of the O&M work. The response time limit refers to the preset time limit for completing the O&M work. The assigned personnel information refers to the name and contact information of the personnel assigned to perform the O&M work. The required spare parts list refers to a list of the names and quantities of spare parts required to perform the O&M work.
[0105] Mandatory work orders or maintenance plans confirmed by the terminal are converted into maintenance work orders. Maintenance work orders include work order number, task type, task description, response time limit, assigned personnel information, and a list of required spare parts.
[0106] Step S7012: Push the maintenance work order to the corresponding terminal and record the acceptance time and completion result of the maintenance work order.
[0107] Specifically, the corresponding terminal refers to the handheld terminal used by the operations and maintenance personnel assigned to execute the maintenance work order. The order acceptance time refers to the time when the operations and maintenance personnel receive the maintenance work order through their terminal. The completion result refers to the processing result reported by the operations and maintenance personnel after completing the maintenance work order.
[0108] The maintenance work order is pushed to the terminal corresponding to the assigned personnel through the maintenance system, and the time when the personnel accept the order is recorded. After the personnel complete the maintenance work order, the completion result reported by them is recorded.
[0109] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which achieves seamless connection from decision-making to execution, ensuring that operation and maintenance plans can be quickly and accurately transformed into executable operation instructions, thereby improving the closed-loop efficiency of operation and maintenance management.
[0110] Specifically, the method for dynamically generating an energy storage power station operation and maintenance plan provided in this embodiment of the invention further includes: Step S801: Record the actual execution data of the scheduling action. The actual execution data includes the actual completion time of the task, newly generated alarm information during task execution, and the recurrence of similar alarms within a preset time period after the task is completed.
[0111] Specifically, the actual execution data of the scheduling action refers to the various recorded data generated during the actual execution of the scheduling action. The actual completion time of the task refers to the actual time taken from the start of the operation and maintenance task to its actual completion. The alarm information newly generated during the task execution refers to the alarm records newly generated by the energy storage power station during the execution of the operation and maintenance task. The recurrence of similar alarms within a preset time period after the task is completed refers to whether the same equipment or the same type of alarm occurs again within a preset time period after the task is completed. The preset time period can be 24 hours.
[0112] Record the actual execution data after the scheduling action is executed. The actual execution data includes the actual completion time of the task, the alarm information newly generated during the task execution, and whether the same type of alarm recurs within 24 hours after the task is completed.
[0113] Step S802: Add the actual execution data within the preset period to the training set, and update the policy network using the near-end policy optimization algorithm to obtain the updated policy network.
[0114] Specifically, the preset cycle refers to the pre-defined model update cycle, which can be once a week. The training set refers to the set of historical interaction data used to train and update the policy network. The proximal policy optimization algorithm is a policy gradient algorithm widely used in the field of reinforcement learning. It aims to solve the problems of unstable training and low sample efficiency in traditional policy gradient methods. Its optimization objective function is shown in formula (4): (4) in, Refers to time step Expectations of experience For probability ratios, For the estimation of the advantage function, The function will use the importance sampling ratio Limited to the range Inside, To prune hyperparameters, a value of 0.2 is often used. Its purpose is to limit the difference between the old and new strategies to avoid training instability or policy collapse caused by excessively large step sizes in a single update, thereby allowing multiple parameter updates while ensuring convergence.
[0115] The actual execution data from the past seven days can be added to the training set once a week, and the policy network can be fine-tuned and updated using the near-end policy optimization algorithm. During fine-tuning, the original experience replay buffer is retained, and the learning rate is set to one-tenth of the initial value to obtain the updated policy network.
[0116] In an alternative implementation, the scheme using a single policy network for decision-making in step S802 can be replaced by a centralized training, distributed execution architecture. For scenarios involving multiple geographically dispersed energy storage power stations, a global policy network is centrally trained in the cloud using historical data from all power stations. Then, copies of the global policy network are deployed to the edge of each power station, with each power station independently performing inference based on its own state. Execution data from each power station is periodically transmitted back to the cloud for incremental updates to the global policy network. This alternative reduces the computing power requirements of individual power stations while allowing smaller power stations to benefit from the operational experience accumulated by larger power stations.
[0117] Specifically, a GPU cluster is deployed on the cloud server to handle offline training and weekly incremental updates of the graph neural network and policy network; a lightweight inference engine is deployed on the edge industrial gateway to load the latest model parameters and perform forward inference when a new alarm is received or when it is triggered every 5 minutes, outputting a recommended operation and maintenance plan with an inference latency of less than or equal to 100ms.
[0118] The safety fallback module runs before decision output, determining if any of the following alarms exist: the alarm level is urgent and the risk growth type is exponential, the equipment temperature exceeds a preset protection threshold, or the fire alarm system issues a fire alarm. If any of these exist, a mandatory work order is directly generated without strategy network recommendation. If none exist, the strategy network recommendation result is pushed to the operations manager's terminal in the form of an operations and maintenance suggestion plan, which is then manually confirmed and executed. The actual execution results of each decision are recorded, including task completion time, the actual handler, and subsequent alarm changes. New data is added to the training set weekly to fine-tune the graph neural network and strategy network, enabling system strategy self-learning and continuous optimization. This implementation method, through a collaborative architecture of centralized cloud training and distributed edge inference, balances the computational resource requirements of model training with the real-time requirements of on-site decision-making, while ensuring decision security in extreme emergency situations through a safety fallback mechanism.
[0119] In step S803, the updated policy network and the policy network are tested on historical data. If the overall performance score of the updated policy network is higher than that of the policy network, the updated policy network replaces the policy network.
[0120] Specifically, historical data refers to historical replay datasets used for performance comparison tests. The overall performance score is a comprehensive quantitative indicator used to evaluate the decision-making quality of the policy network.
[0121] The updated policy network is compared with the original policy network on historical replay data. If the overall performance score of the updated policy network on the historical replay data is higher than that of the original policy network, the updated policy network is deployed to the edge to replace the original policy network.
[0122] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, overcoming the shortcomings of existing technologies that cannot proactively learn forward-looking strategies from historical experience. Furthermore, by employing a near-end strategy optimization algorithm with a pruning objective function, the model update process remains stable, ensuring the reliability of engineering deployment.
[0123] This embodiment also provides a dynamic generation device for energy storage power station operation and maintenance plans. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0124] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which can be used for the aforementioned electronic devices, etc. Figure 7 This is an architecture diagram of a method for dynamically generating operation and maintenance plans for energy storage power stations according to an embodiment of the present invention.
[0125] This embodiment provides a method for dynamically generating operation and maintenance plans for energy storage power stations, which can be used for the aforementioned electronic devices, etc. Figure 8 This is a flowchart illustrating an alarm aggregation method using a graph neural network according to an embodiment of the present invention.
[0126] For example, such as Figure 7 and Figure 8 As shown in the embodiment of the present invention, the method for dynamically generating operation and maintenance plans for energy storage power stations includes the following functional modules: a data acquisition and preprocessing module, an equipment topology diagram construction module, a graph neural network alarm aggregation module, a Markov decision process state management module, a deep reinforcement learning decision module, a safety fallback and manual confirmation module, an operation and maintenance task generation and dispatch module, and a closed-loop feedback and model update module.
[0127] The output of the data acquisition and preprocessing module is connected to the input of the equipment topology map construction module and the Markov decision process state management module, respectively. The output of the equipment topology map construction module is connected to the input of the graph neural network alarm aggregation module. The output of the graph neural network alarm aggregation module is connected to the input of the Markov decision process state management module. The output of the Markov decision process state management module is connected to the input of the deep reinforcement learning decision module. The output of the deep reinforcement learning decision module is connected to the input of the safety fallback and manual confirmation module. The output of the safety fallback and manual confirmation module is connected to the input of the operation and maintenance task generation and dispatch module. The output of the operation and maintenance task generation and dispatch module is connected to the input of the closed-loop feedback and model update module, and the output of the closed-loop feedback and model update module is fed back to the update port of the deep reinforcement learning decision module.
[0128] The workflow of implementing the dynamic generation method for energy storage power station operation and maintenance plans provided by this invention through the above functional modules is as follows: The data acquisition and preprocessing module is used to collect operating parameters, real-time alarm data, and real-time resource information of each subsystem of the energy storage power station in real time at a configurable sampling period, and to clean and format the collected data. The device topology graph construction module is used to construct a topology graph structure with devices in the energy storage power station as nodes and device connections as edges, based on the device connection relationships of the energy storage power station. The graph neural network alarm aggregation module is used to calculate the root cause probability of each node through a graph convolutional network, taking the device topology graph and real-time alarm data as input. When the root cause probability of the alarm exceeds a preset probability threshold... The system generates joint operation and maintenance tasks in real time. The Markov Decision Process (MDP) state management module integrates the operation and maintenance task list, resource status from real-time resource information, and equipment health baseline into a state vector. The deep reinforcement learning (DLM) decision module loads the pre-trained policy network and outputs operation and maintenance scheduling actions based on the current state vector. The safety fallback and manual confirmation module verifies the safety of the scheduling actions output by the DLM decision module and, when preset emergency conditions are met (e.g., emergency alarms, exponential risks, over-temperature, fire hazards), forcibly generates an emergency operation and maintenance work order; otherwise, it submits the recommended plan for manual confirmation. The operation and maintenance task generation and dispatch module converts the confirmed scheduling actions into executable operation and maintenance work orders and dispatches them to designated personnel. The closed-loop feedback and model update module records the decision execution results, such as task execution and alarm changes, and adds the actual execution data within a preset period to the training set to update the graph neural network and reinforcement learning model, achieving policy self-learning and continuous optimization.
[0129] This embodiment provides a device for dynamically generating operation and maintenance plans for energy storage power stations, such as... Figure 9 As shown, it includes: The data acquisition module 901 is used to collect real-time alarm data and real-time resource information of the energy storage power station, and to construct the equipment topology diagram of the energy storage power station. The positioning module 902 is used to locate the root cause node by using a graph neural network to analyze the device topology and real-time alarm data, and obtain a list of maintenance tasks. Module 903 is used to construct a state vector based on real-time resource information and the list of operation and maintenance tasks; The determination module 904 inputs the state vector into the pre-trained policy network, determines the scheduling action, and generates an operation and maintenance plan based on the scheduling action.
[0130] In some alternative implementations, the acquisition module 901 includes: The device topology graph construction unit is used to construct the device topology graph of the energy storage power station with the devices in the energy storage power station as nodes and the device connection relationships in the energy storage power station as edges. The devices include battery clusters, converter units, battery management system controllers, environmental control units and fire detectors. The device connection relationships include electrical connection edges, communication dependency edges and spatial proximity edges.
[0131] In some alternative implementations, the positioning module 902 includes: The feature vector construction unit is used to construct the feature vector corresponding to each node. The feature vector includes device type code, operating status code, alarm count, health score and alarm level code.
[0132] The device feature matrix construction unit is used to construct the device feature matrix based on the nodes in the device topology graph and the feature vector corresponding to each node. The device feature matrix and the adjacency matrix of the device feature matrix are subjected to two-layer graph convolution to obtain the root cause probability and non-root cause probability of each node.
[0133] The alarm root cause probability calculation unit is used to input the root cause probability and non-root cause probability of each node into the softmax function to obtain the alarm root cause probability of each node.
[0134] The root cause node determination unit is used to determine the nodes whose alarm root cause probability is greater than a preset probability threshold as root cause nodes.
[0135] The operation and maintenance task list construction unit is used to aggregate alarms in real-time alarm data that are no more than one hop away from the same root node in the device topology graph as joint operation and maintenance tasks, and to treat alarms in real-time alarm data that cannot be associated with any root node as independent operation and maintenance tasks, so as to construct the operation and maintenance task list.
[0136] In some alternative implementations, the construction module 903 includes: The task queue vector construction unit is used to construct a task queue vector based on the information of each operation and maintenance task in the operation and maintenance task list, construct a resource status vector based on real-time resource information, and construct a device health vector based on the health information of the equipment in the energy storage power station.
[0137] The vector concatenation unit is used to concatenate the task queue vector, resource status vector, and device health vector to obtain a status vector.
[0138] In some alternative embodiments, the apparatus further includes: The risk function determination unit is used to determine the risk function corresponding to each alarm based on the alarm type of each alarm in the real-time alarm data.
[0139] The risk value calculation unit is used to calculate the risk value corresponding to each alarm based on the risk function. The risk function is used to construct the reward function in the training process of the policy network. The risk functions include linear growth type, exponential growth type, S-shaped type and step type.
[0140] In some alternative implementations, the determining module 904 includes: The probability distribution determination unit is used to input the state vector into the hidden layer of the policy network for function activation, and normalize it using the softmax function of the output layer in the policy network to obtain the probability distribution of selecting each scheduling action under the state vector. The scheduling actions include assigning maintenance tasks to designated personnel, rearranging the task queue vector, merging multiple joint maintenance tasks, delaying the processing of low-priority maintenance tasks, and requesting the allocation of spare parts.
[0141] The scheduling action determination unit is used to determine the scheduling action at the current moment based on the probability distribution.
[0142] In some alternative embodiments, the apparatus further includes: The forced work order generation unit is used to generate a forced work order if there is an alarm in the real-time alarm data that meets the preset emergency conditions. The preset emergency conditions include the alarm level being urgent and the alarm type being exponentially increasing, the battery temperature exceeding the preset protection threshold, and the fire alarm being issued by the fire protection system.
[0143] The operation and maintenance plan push unit is used to push the operation and maintenance plan corresponding to the scheduling action to the terminal if there is no alarm in the real-time alarm data that meets the preset emergency conditions. In response to the terminal's confirmation or adjustment instructions, the confirmed or adjusted operation and maintenance plan is used as the final execution plan.
[0144] In some alternative embodiments, the apparatus further includes: The maintenance work order conversion unit is used to convert mandatory work orders and maintenance plans confirmed by the terminal into maintenance work orders. Maintenance work orders include work order number, task type, task description, response time limit, assigned personnel information, and a list of required spare parts.
[0145] The maintenance work order push unit is used to push maintenance work orders to the corresponding terminals and record the acceptance time and completion result of the maintenance work orders.
[0146] In some alternative embodiments, the apparatus further includes: The actual execution data recording unit is used to record the actual execution data of the scheduling action. The actual execution data includes the actual completion time of the task, newly generated alarm information during task execution, and the recurrence of similar alarms within a preset time period after the task is completed.
[0147] The policy network update unit is used to add the actual execution data within a preset period to the training set and update the policy network using the near-end policy optimization algorithm to obtain the updated policy network.
[0148] The policy network replacement unit is used to perform performance tests on historical data between the updated policy network and the original policy network. If the overall performance score of the updated policy network is higher than that of the original policy network, the updated policy network will replace the original policy network.
[0149] The energy storage power station operation and maintenance plan dynamic generation device provided in this embodiment of the invention can execute the energy storage power station operation and maintenance plan dynamic generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments above, and will not be repeated here.
[0150] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0151] The following is a detailed reference. Figure 10 This diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from memory 1008 into random access memory (RAM) 1003. The RAM 1003 also stores various programs and data required for the operation of the electronic device. The processor 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0152] Typically, the following devices can be connected to the I / O interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; memory devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to exchange data via wireless or wired communication with other devices. Although Figure 10 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0153] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 1009, or installed from a memory 1008, or installed from a ROM 1002. When the computer program is executed by the processor 1001, it performs the functions defined in the method for dynamically generating an energy storage power station operation and maintenance plan according to an embodiment of the present invention.
[0154] Figure 10 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0155] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, it implements the dynamic generation method for energy storage power station operation and maintenance plans shown in the above embodiments.
[0156] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0157] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for dynamically generating operation and maintenance plans for energy storage power stations, characterized in that, The method includes: Collect real-time alarm data and real-time resource information of the energy storage power station, and construct the equipment topology diagram of the energy storage power station; A graph neural network is used to locate root cause nodes in the device topology and real-time alarm data to obtain a list of maintenance tasks. A state vector is constructed based on real-time resource information and the aforementioned operation and maintenance task list; The state vector is input into a pre-trained policy network to determine scheduling actions, and an operation and maintenance plan is generated based on the scheduling actions.
2. The method according to claim 1, characterized in that, Constructing the equipment topology diagram of the energy storage power station includes: Using the devices in the energy storage power station as nodes and the device connection relationships in the energy storage power station as edges, a device topology graph of the energy storage power station is constructed. The devices include battery clusters, converter units, battery management system controllers, environmental control units, and fire detectors. The device connection relationships include electrical connection edges, communication dependency edges, and spatial proximity edges.
3. The method according to claim 1, characterized in that, By using a graph neural network to locate root causes in the device topology and real-time alarm data, a list of maintenance tasks is obtained, including: Construct a feature vector for each node, the feature vector including device type code, operating status code, alarm count, health score and alarm level code; Based on the nodes in the device topology graph and the feature vector corresponding to each node, a device feature matrix is constructed. The device feature matrix and the adjacency matrix of the device feature matrix are subjected to two-layer graph convolution to obtain the root cause probability and non-root cause probability of each node. The root cause probability and non-root cause probability of each node are input into the softmax function to obtain the alarm root cause probability of each node. Nodes whose alarm root cause probability is greater than a preset probability threshold are identified as root cause nodes. Alarms in the real-time alarm data that are no more than one hop away from the same root cause node in the device topology graph are aggregated as joint operation and maintenance tasks, and alarms in the real-time alarm data that cannot be associated with any of the root cause nodes are treated as independent operation and maintenance tasks, so as to construct an operation and maintenance task list.
4. The method according to claim 1, characterized in that, Based on real-time resource information and the aforementioned operation and maintenance task list, a state vector is constructed, including: Based on the information of each maintenance task in the maintenance task list, a task queue vector is constructed; based on the real-time resource information, a resource status vector is constructed; and based on the health information of the equipment in the energy storage power station, an equipment health vector is constructed. The task queue vector, the resource status vector, and the device health vector are concatenated to obtain the status vector.
5. The method according to claim 1, characterized in that, The method further includes: Based on the alarm type of each alarm in the real-time alarm data, determine the risk function corresponding to each alarm; The risk value corresponding to each alarm is calculated based on the risk function, which is used to construct the reward function in the training process of the policy network. The risk function includes linear growth type, exponential growth type, S-shaped type and step type.
6. The method according to claim 1, characterized in that, The state vector is input into a pre-trained policy network to determine the scheduling action, including: The state vector is input into the hidden layer of the policy network for function activation, and normalized using the softmax function of the output layer of the policy network to obtain the probability distribution of selecting each scheduling action under the state vector. The scheduling actions include assigning the operation and maintenance task to a designated person, rearranging the task queue vector, merging multiple joint operation and maintenance tasks, delaying the processing of low-priority operation and maintenance tasks, and requesting the allocation of spare parts. Based on the probability distribution, determine the scheduling action at the current moment.
7. The method according to claim 1, characterized in that, The method further includes: If there is an alarm in the real-time alarm data that meets the preset emergency conditions, a mandatory work order is generated. The preset emergency conditions include an alarm level of emergency and an alarm type of exponential growth, a battery temperature exceeding a preset protection threshold, and a fire alarm issued by the fire protection system. If there are no alarms in the real-time alarm data that meet the preset emergency conditions, the operation and maintenance plan corresponding to the scheduling action will be pushed to the terminal. In response to the confirmation or adjustment instruction from the terminal, the confirmed or adjusted operation and maintenance plan will be used as the final execution plan.
8. The method according to claim 7, characterized in that, The method further includes: The mandatory work order and the maintenance plan confirmed by the terminal are converted into a maintenance work order. The maintenance work order includes a work order number, task type, task description, response time limit, assigned personnel information, and a list of required spare parts. The maintenance work order is pushed to the corresponding terminal, and the time of acceptance and completion of the maintenance work order are recorded.
9. The method according to claim 1, characterized in that, The method further includes: Record the actual execution data of the scheduling action, including the actual task completion time, newly generated alarm information during task execution, and the recurrence of similar alarms within a preset time period after the task is completed; The actual execution data within a preset period is added to the training set, and the policy network is updated using the near-end policy optimization algorithm to obtain the updated policy network. The updated policy network and the original policy network are tested on historical data. If the overall performance score of the updated policy network is higher than that of the original policy network, the updated policy network replaces the original policy network.
10. A device for dynamically generating operation and maintenance plans for energy storage power stations, characterized in that, The device includes: The data acquisition module is used to collect real-time alarm data and real-time resource information of the energy storage power station, and to construct the equipment topology diagram of the energy storage power station. The positioning module is used to locate the root cause node of the device topology map and the real-time alarm data using a graph neural network to obtain a list of operation and maintenance tasks. The construction module is used to construct a state vector based on real-time resource information and the operation and maintenance task list; The determination module inputs the state vector into the pre-trained policy network, determines the scheduling action, and generates an operation and maintenance plan based on the scheduling action.