Multi-layer distributed micro-grid control system and method based on edge cloud collaborative lightweight reinforcement learning

By adopting a multi-layer distributed control system with edge-cloud collaborative lightweight reinforcement learning in the microgrid, and using the lightweight DQN model for real-time decision-making and scheduling, the shortcomings of the existing microgrid control systems in real-time, reliability and economics are solved, and the effective management of high-permeability renewable energy access is achieved.

CN120073869APending Publication Date: 2025-05-30YANTAI DEV ZONE DELIAN SOFTWARE CO LTD
View PDF 0 Cites 32 Cited by

Patent Information

Application Number
CN202510399305.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing microgrid control systems have shortcomings in real-time, reliability and economics, and are difficult to meet the requirements of high permeability renewable energy access.

Method used

A multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning is adopted. Through the collaborative work of cloud management units, edge node units and microgrid subgroup units, the lightweight DQN model is used for real-time decision-making and scheduling.

Benefits of technology

It significantly improves the real-time response capability and system robustness of the microgrid, reduces the computing and communication pressure of the central control layer, and enhances the scalability and self-healing ability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120073869A_ABST
    Figure CN120073869A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of micro-grid control, in particular to a multi-layer distributed micro-grid control system and method based on edge cloud collaborative lightweight reinforcement learning. The method comprises the following steps: a cloud management unit trains a DQN model by using historical environment data, generates a lightweight DQN model through pruning and quantification, issues the lightweight DQN model to an edge node unit, and formulates day-level and weekly-level optimization scheduling strategies based on a function layering + layering cooperation strategy; the edge node unit monitors electrical quantity and environmental parameters in real time through a sensor, deploys a lightweight DQN model to monitor abnormal conditions and trigger local early warning and emergency strategies, and makes a fast decision by using the lightweight DQN model; the micro-grid subgroup units collect data of the edge node units. According to the invention, through close cooperation of the edge node layer, the micro-grid subgroup layer and the cloud management layer, a distributed cooperative control mode of edge node real-time decision, subgroup cooperative management and cloud global optimization is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of microgrid control, and more specifically, to a multi-layer distributed microgrid control system and method based on edge-cloud collaborative lightweight reinforcement learning. Background Art

[0002] With the rapid development of renewable energy, the application scale of distributed energy and energy storage technologies in the distribution network is continuously expanding, and the traditional centralized power grid dispatching mode is facing huge challenges. As an advanced concept that can coordinate and manage distributed energy, energy storage systems and loads, the microgrid has gradually become an important direction for the development of the power system. However, there are still many deficiencies in the existing technology for the distributed control of microgrids, making it difficult to meet the requirements of real-time performance, reliability and economy under the access of high-penetration renewable energy.

[0003] The existing microgrid control system is gradually evolving from a centralized or hierarchical centralized mode to a distributed and intelligent direction. The emergence of edge computing technology and lightweight reinforcement learning algorithms provides new ideas for improving the real-time response ability of microgrids, reducing communication and computing pressure. However, there is still a lack of systematic and mature solutions for key issues such as the division of distributed control levels, the lightweight deployment of reinforcement learning, and the handling of large-scale abnormal events and self-healing mechanisms. Therefore, a multi-layer distributed microgrid control system and method based on edge-cloud collaborative lightweight reinforcement learning are provided. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-layer distributed microgrid control system and method based on edge-cloud collaborative lightweight reinforcement learning, so as to solve the problems of insufficient real-time performance and reliability in microgrid management, poor scalability of the centralized architecture, the randomness problems brought by the grid connection of renewable energy, and the resource allocation problems in large-scale scenarios proposed in the above background art.

[0005] To achieve the above purpose, on the one hand, the present invention aims to provide a multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning, including: A cloud management unit, which trains a DQN model using historical environmental data, generates a lightweight DQN model through pruning and quantization, sends the lightweight DQN model to the edge node unit, and formulates daily and weekly optimized scheduling strategies based on a function layer + hierarchical collaboration strategy; An edge node unit, which monitors electrical quantities and environmental parameters in real time through sensors, deploys a lightweight DQN model to monitor abnormal situations and trigger local early warning and emergency strategies, and makes quick decisions using the lightweight DQN model; The microgrid subgroup unit collects data from the edge node units, forms a regional energy balance analysis, comprehensively schedules power generation, energy storage, and loads within a local area, introduces a confidence factor and the total power of energy exchange between subgroups for optimization during the scheduling process, and introduces a multi-agent collaborative self-healing mechanism for fault isolation and self-healing.

[0006] As a further improvement of this technical solution, the cloud management unit includes a global optimization and strategy generation module, a model training and compression module, and an anomaly linkage and fault tolerance management module; Among them, the global optimization and strategy generation module stores and processes data of subgroups and the external environment, and formulates daily and weekly optimization scheduling strategies based on a functional layer + hierarchical collaboration strategy; The model training and compression module uses historical environment data to train the DQN model, generates a lightweight model through pruning and quantization, and periodically sends the pruned and quantized lightweight DQN model to the edge node units; The anomaly linkage and fault tolerance management module analyzes network-wide anomaly events, coordinates the subgroup layer to initiate a cross-domain emergency plan, periodically optimizes the microgrid operation strategy, and sends emergency instructions to the microgrid subgroup unit.

[0007] As a further improvement of this technical solution, the model training and compression module uses historical environment data to train the DQN model, generates a lightweight model through pruning and quantization, including the following steps: S1.1. Construct the state vector of reinforcement learning : ; Among them, represents the output power of the distributed generation side; represents the power consumption of the load side; represents the state of charge of the energy storage unit; represents the renewable energy surplus; represents the voltage; represents the current; represents the device temperature; represents the time; Define the set of actions : ; Among them, represents the set of all possible actions; represents a specific action option; represents the different types of actions; represents the total number of actions in the action set; S1.2. Use historical environmental data in the cloud to perform offline training on the Deep Q-Network (DQN) model, design a reward function, and iteratively update the Q-value function; S1.3. Determine the part of the weight matrix of the DQN with a contribution degree less than the threshold a to the Q-value output through redundancy analysis, and apply pruning technology to remove the part with a contribution degree less than the threshold a; S1.4. Perform small-scale fine-tuning on the pruned DQN model, and convert the pruned floating-point weights into low-precision fixed-point numbers for quantization processing; S1.5. Use sparse matrix storage technology to further compress the file size of the pruned and quantized lightweight DQN model; S1.6. Transmit the optimized lightweight DQN model to the edge node unit through an encrypted communication protocol and the MQTT lightweight protocol.

[0008] As a further improvement of this technical solution, in S1.1, the reward function is: ; Among them, represents the immediate reward; represents the operating cost; represents the current moment the power balance deviation of power generation, energy storage, and load; represents the fault penalty term; represents the adjustable weight coefficient of the operating cost; represents the adjustable weight coefficient of the power balance deviation; represents the adjustable weight coefficient of the fault penalty term; The Q-value function is: ; Among them, represents the parameter vector of the main network; is the target network parameter; represents the learning rate, is the discount factor.

[0009] As a further improvement of this technical solution, formulating daily and weekly optimization scheduling strategies based on the function stratification + stratification cooperation strategy includes the following steps: S1.7. Obtain the historical operation data of the microgrid from the subgroup controller and collect meteorological data; S1.8. Define the local multi-objective optimization function: ; Among them, represents the comprehensive objective function at the subgroup level; represents the power generation subsystem; represents the energy storage subsystem; Represents the load subsystem; Represents the operating cost of the subsystem; Represents the deviation degree of the energy utilization rate of the subsystem; Represents the risk index of the subsystem; Represents the weight coefficient of the operating cost; Represents the weight coefficient of the deviation degree of the energy utilization rate; Represents the weight coefficient of the risk index; Represents the index of the subsystem; Define the system constraints: Power balance: ; Energy storage state of charge (SOC) limit: ; The charge and discharge power of the energy storage does not exceed the rated value: ; Wherein, Represents the actual output power of the power generation subsystem at time t; Represents the charge and discharge power of the energy storage subsystem; Represents the power consumed by the load; Represents the interactive power with the external power grid; Represents the state of charge of the energy storage system at time ; Represents the rated power constraint of the energy storage charge and discharge; S1.9. Select the linear programming method to construct the optimization model; S1.10. Based on historical data and meteorological data, use the deep learning LSTM model to predict the renewable energy output and load demand in the future period; S1.11. Use particle swarm optimization to solve the above-established optimization model to find the optimal scheduling strategy; S1.12. Generate a specific scheduling plan according to the solution results.

[0010] As a further improvement of the technical solution, the edge node unit includes a data acquisition and preprocessing module, a lightweight reinforcement learning decision module, and a local anomaly detection and emergency module; Wherein, the data acquisition and preprocessing module monitors electrical parameters through sensors, and at the same time uses edge nodes to monitor environmental parameters and equipment health status, and performs data preprocessing at the edge nodes; The lightweight reinforcement learning decision-making module integrates the collected and monitored data to form the state input of reinforcement learning, deploys the pruned and quantized lightweight DQN model, and infers the optimal action in real time based on the local state; The local anomaly detection and emergency module calculates the anomaly index through multi-source data fusion, triggers local warnings using a lightweight anomaly detection mechanism based on threshold judgment, executes local emergency strategies, gives priority to ensuring local safety, and synchronously reports fault information to the cloud management unit.

[0011] As a further improvement of this technical solution, the lightweight reinforcement learning decision-making module integrates the collected and monitored data to form the state input of reinforcement learning, deploys the pruned and quantized lightweight DQN model, and infers the optimal action in real time based on the local state, including the following steps: S2.1. Deploy the pruned and quantized lightweight DQN model on the edge node to perform fast value calculations within the second and minute-level scheduling cycles; S2.2. When receiving the state vector at the current moment , use the lightweight DQN model for forward inference to calculate the Q values of each action : ; Among them, represents the Q value; represents using the lightweight DQN model to calculate the Q values of all actions in a given state; S2.3. Select the action with the highest Q value as the optimal action at the current moment : ; S2.4. Execute the optimal action ; S2.5. Periodically upload the execution data and the execution effect to the cloud management unit, and send the execution instruction to the microgrid subgroup unit.

[0012] As a further improvement of this technical solution, the microgrid subgroup unit collects the data of the edge nodes, forms an analysis of regional energy balance, conducts comprehensive scheduling of power generation, energy storage, and loads in the local area, and introduces a multi-agent collaborative self-healing mechanism for fault isolation and self-healing, including the following steps: S3.1. Summarize the data of all edge nodes in the subgroup and calculate the energy balance situation at each time point in the region: ; Among them, represents the power generation subsystem at time The actual output power; Indicates the charge and discharge power of the energy storage subsystem; Indicates the power consumed by the load; Indicates the interactive power with the external power grid; S3.2. Establish the objective function and constraint conditions, use the linear programming method for optimization, and introduce a confidence factor into the objective function And the total power of energy exchange between subgroups For optimization; S3.3. When a fault is detected, start the multi-agent collaborative self-healing mechanism. The subgroup controller analyzes the abnormal signals of the edge node units and performs fault isolation, compensates for the faults through the remaining edge nodes in the edge node units, and synchronously uploads the fault control status to the cloud management unit.

[0013] As a further improvement of this technical solution, in the above S3.2, the objective function is: ; Wherein, Represents the emergency decision of the th edge node at time , Is the fault risk function of the edge node In the emergency mode, Is the energy imbalance of node ; Represents the weight coefficient of the fault risk; Represents the weight coefficient of the energy imbalance; Represents the index of the edge node; Represents the total number of edge nodes; Aiming at the problem that the objective function cannot respond to load sensitivity and environmental fluctuations, generate dynamic weights through the spatio-temporal attention mechanism , and considering the reliability of model prediction, introduce a confidence factor Into the objective function for optimization: ; Aiming at the problem of energy interaction cost between subgroups, introduce the total power of energy exchange between subgroups Into the objective function for further optimization: ; ; Wherein, Represents the adjustment coefficient that controls the influence degree of energy interaction between subgroups on the total objective; Represents subgroup To Transmission efficiency matrix; Represents subgroup To the actual power transmitted; and represents the index of the subgroup.

[0014] On the other hand, the present invention provides a multi-layer distributed microgrid control method based on edge-cloud collaborative lightweight reinforcement learning for a multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning described in any one of the above, including the following steps:

[0015] S4.1. Use historical environmental data to train a DQN model, generate a lightweight DQN model through pruning and quantization, send the lightweight DQN model to the edge nodes, and formulate daily and weekly optimal scheduling strategies based on a function-layered + hierarchical collaborative strategy; S4.2. Real-time monitor electrical quantities and environmental parameters through sensors, deploy the lightweight DQN model to monitor abnormal conditions and trigger local early warning and emergency strategies, and use the lightweight DQN model to make quick decisions; S4.3. Collect data from the edge nodes, form a regional energy balance analysis, comprehensively schedule power generation, energy storage, and loads within a local area, and introduce a multi-agent collaborative self-healing mechanism for fault isolation and self-healing.

[0016] Compared with the prior art, the beneficial effects of the present invention:

[0017] In a multi-layer distributed microgrid control system and method based on edge-cloud collaborative lightweight reinforcement learning, through the close cooperation of the edge node layer, the microgrid subgroup layer, and the cloud management layer, not only a distributed collaborative control mode of "edge node real-time decision-making + subgroup collaborative management + cloud global optimization" is realized, but also the computing and communication pressure of the central control layer is significantly reduced, and the robustness and scalability of the system under the condition of high-penetration renewable energy grid connection are enhanced. The edge nodes can quickly respond to local changes or abnormal events using lightweight reinforcement learning algorithms, the microgrid subgroup layer is responsible for comprehensive scheduling and fault self-healing within a local area, while the cloud management center focuses on large-scale data mining and long-term planning, thereby realizing the efficient management and optimal scheduling of the microgrid at multiple time scales and multiple spatial levels. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is the overall flow block diagram of the present invention; The meanings of the various labels in the figure are as follows: 1. Cloud management unit; 11. Global optimization and policy generation module; 12. Model training and compression module; 13. Abnormal linkage and fault tolerance management module; 2. Edge node unit; 21. Data acquisition and preprocessing module; 22. Lightweight reinforcement learning decision-making module; 23. Local abnormal detection and emergency module; 3. Microgrid subgroup unit. Detailed implementation manners

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0020] Embodiment 1: Please refer to Figure 1 As shown, a multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning is provided, including: The cloud management unit 1 trains a DQN model using historical environment data, generates a lightweight DQN model through pruning and quantization, sends the lightweight DQN model to the edge node unit 2, and formulates daily and weekly optimization scheduling strategies based on the function layer + layer collaborative strategy; In this embodiment, the cloud management unit 1 includes a global optimization and policy generation module 11, a model training and compression module 12, and an abnormal linkage and fault tolerance management module 13; Among them, the global optimization and policy generation module 11 stores and processes data of subgroups and the external environment, and formulates daily and weekly optimization scheduling strategies (such as economic optimization and carbon emission optimization) based on the function layer + layer collaborative strategy; The model training and compression module 12 trains a DQN model using historical environment data, generates a lightweight model through pruning (removing redundant weights) and quantization (converting floating-point to fixed-point), and periodically sends the pruned and quantized lightweight DQN model to the edge node unit 2 to ensure continuous optimization of the strategy; Among them, the historical environmental data includes: electrical parameters: output power of the distributed generation side, power consumption of the load side, charge and discharge power of the energy storage, state of charge of the energy storage, voltage, current; environmental parameters: real-time monitoring data such as wind speed, sunlight intensity, air temperature, etc. that affect renewable energy generation. Equipment status data: equipment health status information such as equipment temperature, fault codes, etc. Operating cost data: economic indicators such as power generation cost, energy storage depreciation cost, electricity price cost, etc. Meteorological data: historical meteorological records used to predict renewable energy output (such as photovoltaic, wind power) and load demand. Fault and anomaly records: historical fault types, occurrence times, emergency measures and treatment results. Time series data: time series data of historical power generation and load demand, used to train the LSTM prediction model. Multi-objective optimization related indicators: comprehensive optimization parameters at the subgroup level such as energy utilization deviation degree, risk indicators; The anomaly linkage and fault tolerance management module 13 analyzes the abnormal events of the whole network (such as extreme weather, multi-region faults), coordinates the subgroup layer to start cross-domain emergency plans (such as calling standby power supplies, load transfer), periodically optimizes the microgrid operation strategy, and issues emergency instructions to the microgrid subgroup unit 3. When large-scale power fluctuations or extreme accidents (such as typhoons, earthquakes, etc.) cause multiple subgroups to face risks at the same time, start cross-regional emergency plans, and conduct emergency dispatching by integrating the energy storage capabilities and load elastic resources of each subgroup. If the cloud communication is blocked or temporarily fails, the microgrid subgroup unit 3 and the edge node unit 2 can still respond autonomously according to local rules and reinforcement learning strategies; after the communication is restored, the cloud will correct the global state of the system and conduct unified dispatching.

[0021] Among them, pruning can remove unimportant weight connections and reduce the number of multiplication and addition operations required; quantization converts floating-point numbers into fixed-point number representations, reducing the bit width required for each calculation. The combination of the two greatly reduces the amount of calculation required during inference. The size of the lightweight model is smaller, occupying less storage space and runtime memory, which is especially important for edge devices; due to the reduction in the amount of calculation and data transmission volume, the lightweight model can achieve a faster response time on the same hardware, which is crucial for application scenarios that require quick decision-making (such as autonomous driving, robot control, etc.); a more efficient model allows more tasks or services to be run simultaneously on the same device, improving the overall throughput of the system; The model training and compression module 12 trains the DQN model using historical environmental data and generates a lightweight model through pruning (removing redundant weights) and quantization (floating-point to fixed-point), including the following steps: S1.1. Construct the state vector of reinforcement learning : ; Among them, represents the output power of the distributed generation side; Represents the power consumption of the load end; Represents the state of charge of the energy storage unit (between 0 and 1); Represents the remaining renewable energy; Represents the voltage; Represents the current; Represents the device temperature; Represents the time; Defines the set of actions (Scheduling operations that edge nodes can perform at time , including setting the charging and discharging power of the energy storage, on / off strategies for controllable loads, etc.): ; Among them, Represents the set of all possible actions; Represents a specific action option; Represents different types of actions; Represents the total number of actions in the action set; S1.2. Offline train the deep Q-network DQN model using historical or simulation environment data in the cloud, design the reward function and iteratively update the Q-value function; Furthermore, the reward function is: ; Among them, Represents the immediate reward, which is a quantitative feedback on the current state of the system and the actions taken; Represents the operating cost or electricity price cost; Represents this moment The power balance deviation among power generation, energy storage, and load; Represents the device reliability or fault penalty term (such as overcharging or over-discharging of the energy storage, over-temperature operation of the device, etc.); Represents the adjustable weight coefficient of the operating cost; Represents the adjustable weight coefficient of the power balance deviation; Represents the adjustable weight coefficient of the fault penalty term; The Q-value function is: ; Among them, Represents the parameter vector of the main network; Is the parameter of the target network; Represents the learning rate, Is the discount factor.

[0022] S1.3. Determine the part in the weight matrix of the DQN that contributes less than the threshold α to the Q-value output through redundancy analysis, and apply pruning technology to remove the part with a contribution less than the threshold α to reduce the model volume. Pruning simplifies the model structure by removing weights with low contribution to model prediction, achieving the purpose of reducing the model volume; Further, to quantitatively describe the pruning process, a mask matrix is introduced , where represents retaining the th weight, while represents pruning this weight. Define the objective function: where, is the sampled validation data set, represents element-wise multiplication, represents the number of 1s in M. By restricting the number of retained weights not to exceed K, the Q-value error can be controlled at a low level while reducing the model volume by solving the above formula; S1.4. Perform a small-scale fine-tuning on the pruned DQN model to ensure that its performance does not drop significantly due to parameter reduction. Convert the floating-point weights after pruning to low-precision fixed-point numbers for quantization processing (INT8). Quantization simplifies the representation by reducing the numerical precision of the weights, thereby reducing the model size and accelerating the calculation for efficient operation on resource-constrained edge devices; Further, the quantized weights are represented as: ; where, represents the quantization step (generated by histogram statistics). At this time, simple integer operations can be used for inference on the edge side, significantly reducing the operation and storage costs; S1.5. Use sparse matrix storage technology to further compress the file size of the pruned and quantized lightweight DQN model. Sparse matrix storage reduces the storage requirements of the model by only retaining non-zero weights and their position information, achieving further compression; S1.6. Transmit the optimized lightweight DQN model to the edge node unit 2 through an encrypted communication protocol (SSL / TLS) and the MQTT lightweight protocol.

[0023] In this embodiment, to adapt to the resource limitations of edge devices, a lightweight protocol such as MQTT is adopted to achieve asynchronous message publishing and subscribing on the edge side; the cloud supports data interaction with the subgroup controller in ways such as RESTful or WebSocket. When the network is congested or interrupted, the edge node and subgroup layer can enter the "offline autonomous" mode and temporarily use local reinforcement learning strategies for decision-making; after the network is restored, key operation data or fault logs are batch-transmitted back to the cloud.

[0024] To reduce the risk of data leakage, two-way authentication is performed between the edge nodes and the cloud management center to prevent illegal devices from accessing the system; different levels of users (such as dispatchers, maintenance personnel) are assigned different access permissions. And sensitive data (such as device control instructions, fault information) is encrypted during transmission using SSL / TLS secure connections or lightweight encryption algorithms to reduce the risk of data leakage. A network security monitoring module is deployed in the cloud or sub-group controller, combined with traffic analysis and AI recognition, to detect malicious attacks or abnormal traffic in a timely manner, and isolate and warn against potential threats.

[0025] Through the above data collection and communication mechanism, on the premise of ensuring information security and reducing communication burden, efficient data interaction between "edge - sub-group - cloud" is achieved, enabling real-time control and global optimization strategies to be smoothly transmitted through the hierarchical network. At the same time, when local network or cloud connections are abnormal, the edge nodes and sub-groups can also operate autonomously based on local data and pre-deployed lightweight reinforcement learning models, providing a solid guarantee for the stability and scalability of the microgrid system.

[0026] Among them, daily and weekly optimal scheduling strategies are formulated based on the function hierarchical + hierarchical cooperation strategy, including the following steps: Based on the adaptive hierarchical control strategy of "function division + hierarchical autonomy + multi-layer cooperation", it aims to fully explore the operating characteristics of each functional module of the microgrid, and through the organic cooperation between multiple levels, achieve the overall optimal operating goal. The function hierarchical + hierarchical cooperation strategy, through the idea of "adaptive hierarchy", the upper layer can dynamically adjust equal weight coefficients, providing new scheduling guidance or policy focus for the lower-layer optimization model, so that the entire microgrid can flexibly switch the focus in different operating stages (such as peak shaving and valley filling, emergency support, economic optimization, etc.); The microgrid is mainly divided into the following subsystems according to its functions (the boundaries of each subsystem can be flexibly adjusted according to the actual microgrid application scenario), and hierarchical autonomy and coordination are formed on this basis: Generation subsystem: It includes wind power, photovoltaic, and other distributed generation units, mainly responsible for providing renewable energy supply. This type of subsystem has output volatility and uncertainty, and local prediction correction and power scheduling are required; Energy storage subsystem: It consists of various energy storage devices (such as battery energy storage, flywheel energy storage, supercapacitor, etc.), which are used for peak shaving, valley filling, suppressing fluctuations, or emergency power supply. This type of subsystem has physical constraints such as state of charge (SOC), and factors such as life attenuation and energy utilization efficiency need to be considered during scheduling; Load subsystem: It is composed of controllable loads (industrial and commercial loads, charging piles, shiftable loads, etc.) and some uncontrollable loads, providing electrical energy or other forms of energy services for the user side. The energy demand of this type of subsystem has time-varying characteristics and peak-valley differences, and flexible scheduling needs to be combined with price signals or demand response strategies; Scheduling and management subsystem: It is responsible for issuing scheduling instructions to each functional subsystem and comprehensive data processing, and at the same time maintaining communication with the cloud management center to receive long-term optimization suggestions or model updates. This type of subsystem focuses on global information integration and autonomous decision-making within the region, and can directly issue emergency instructions in local scenarios to cooperate with the rapid execution on the edge side; Based on the above "function division", this embodiment further proposes a strategy design of "hierarchical autonomy + cross-layer coordination", enabling the microgrid to form several levels from top to bottom. Each level executes control and optimization within its own function range, and at the same time exchanges necessary information or instructions with adjacent levels: Edge autonomy layer: Each generation, energy storage, and load subsystem operates independently at this level, executing lightweight reinforcement learning or rule control to achieve rapid adjustment of second / minute-level local fluctuations. This level focuses on "real-time performance". Once an abnormal event (such as power mutation, equipment failure, etc.) is detected, an emergency strategy can be made in a very short time; Regional coordination layer: Several subsystems with similar functions or adjacent geographical locations are divided into a coordination domain (which can be called a subgroup), and the scheduling and management subsystem or subgroup controller conducts unified coordination and data aggregation for it. This level focuses on "local optimization", and improves the overall local benefit by balancing the energy flow and sharing resources among subsystems. For example, the energy storage provides assistance to the load system in period A and supports the energy regulation of the generation system in period B; Cloud global layer: The cloud management center conducts comprehensive and long-term strategy analysis and reinforcement learning model training for the entire network, forms daily / weekly optimization plans, and issues them to the lower layer in the form of scheduling instructions or lightweight model updates. This level focuses on "cross-time and space optimization", and uses large-scale historical data mining and prediction algorithms to achieve iterative updates of global scheduling strategies, taking into account economy, environmental protection, and system safety and stability; S1.7. Obtain the historical operation data of the microgrid from the subgroup controller, including power generation, energy storage status, load demand, etc., and collect meteorological data, which is used for the historical meteorological records to predict the output of renewable energy (such as photovoltaic and wind power) and load demand; S1.8. Define the local multi-objective optimization function: ; Among them, represents the comprehensive objective function at the subgroup level; represents the power generation subsystem; represents the energy storage subsystem; represents the load subsystem; represents the th subsystem's operating cost (which can include power generation cost, energy storage depreciation, electricity price cost, etc.); represents the th subsystem's energy utilization rate deviation (such as the energy storage SOC not being in the reasonable range, photovoltaic curtailment rate, etc.); represents the th subsystem's reliability or risk index (which can quantify the impact of frequent start-stop on equipment life, failure probability, etc.); represents the weight coefficient of the operating cost; represents the weight coefficient of the energy utilization rate deviation; represents the weight coefficient of the risk index; represents the index of the subsystem; Define the system constraint conditions (the basic constraints that need to be satisfied for energy balance and safe operation among subsystems): Power balance: ; Energy storage state of charge (SOC) limit: ; The charge and discharge power of the energy storage does not exceed the rated value: ; Among them, represents the actual output power of the power generation subsystem at time t; represents the charge and discharge power of the energy storage subsystem; represents the power consumed by the load; represents the interactive power with the external power grid; represents the state of charge of the energy storage system at time ; represents the rated power constraint of the energy storage charge and discharge; S1.9. Select the linear programming method to construct the optimization model. Linear programming provides an efficient and analytically solvable method for complex optimization problems by solving the optimal value of a linear objective function subject to linear constraint conditions; S1.10. Based on historical data and meteorological data, use the deep learning LSTM model to predict the renewable energy output and load demand in the future period. The LSTM model can achieve accurate prediction of future renewable energy output and load demand by analyzing the time series characteristics and long-term dependencies in historical data. By introducing memory units and gating mechanisms (input gate, forget gate, output gate), the LSTM model can selectively remember or forget historical information, thus performing well in time series prediction tasks. Specifically, by analyzing the time series characteristics and long-term dependencies in historical data, the LSTM model can accurately predict the power generation of renewable energy and load demand. The inputs of the model include historical power generation, load demand data, and meteorological data, and the output is the predicted value for a future period. The training process of the LSTM model uses historical operation data and meteorological data as the training set, and minimizes the prediction error through an optimization algorithm. The trained LSTM model is used to generate the prediction results of renewable energy output and load demand, providing data support for the optimal scheduling of the microgrid; S1.11. Use particle swarm optimization to solve the above-established optimization model and find the optimal scheduling strategy. Particle swarm optimization searches for the optimal or near-optimal scheduling strategy in the solution space through swarm intelligence and iterative search; S1.12. Generate a specific scheduling plan according to the solution results, covering aspects such as power generation, energy storage, and load management.

[0027] The edge node unit 2 is deployed at the bottom layer of the microgrid. It uses sensors to monitor electrical quantities such as voltage and current in real time, as well as environmental parameters. A lightweight DQN model is deployed to monitor abnormal situations and trigger local early warnings and emergency strategies. The lightweight DQN model makes rapid decisions on energy storage charging and discharging power, load switching, etc., adapting to real-time changes at the second / minute level; In this embodiment, the edge node unit 2 includes a data acquisition and preprocessing module 21, a lightweight reinforcement learning decision-making module 22, and a local anomaly detection and emergency module 23; Among them, the data acquisition and preprocessing module 21 monitors electrical parameters such as voltage, current, power, energy storage state (SOC), and device temperature through sensors. At the same time, it uses edge nodes to monitor environmental parameters (wind speed, sunlight intensity, temperature, etc.) and the health status of devices, and performs data preprocessing at the edge node. According to pre-configured rules (including threshold filtering, denoising processing, etc.), the collected data is initially cleaned or compressed to reduce communication bandwidth occupancy. If data missing, mutation, or suspicious abnormal values occur, they are marked at the edge node to provide support for subsequent anomaly detection or fault diagnosis; The lightweight reinforcement learning decision-making module 22 integrates the collected and monitored data to form the state input for reinforcement learning, deploys the pruned and quantized lightweight DQN model, and infers the optimal actions (energy storage charge and discharge power setting, controllable load switching strategy) in real time based on local states (such as power generation, load demand, energy storage SOC). On the premise of ensuring normal operation, the edge node integrates the recently monitored data (load demand, energy storage SOC, wind and solar power output, etc.) to form the state input for reinforcement learning. The reinforcement learning model after lightweight processing (pruning, quantization, etc.) is inferred on the edge node to output the optimal or approximately optimal control actions (such as energy storage charge and discharge power, adjustable load switching time, power generation scheduling, etc.) in real time, achieving high-frequency decision-making within a second-level or minute-level cycle. In addition, if the edge node has a certain amount of computing power redundancy, short-cycle online learning or parameter fine-tuning can be performed locally to improve the adaptability of the strategy to random fluctuations; and the relevant training data or experience samples are reported asynchronously to the upper layer for retraining or overall evaluation of the cloud model. The local anomaly detection and emergency module 23 calculates the anomaly index through multi-source data fusion (voltage, current, temperature, etc.) (using the anomaly index to quantify the deviation degree between the current state and the normal working mode), triggers local warnings using a lightweight anomaly detection mechanism based on threshold judgment, and executes local emergency strategies (such as cutting off the faulty line, enabling energy storage emergency power supply), giving priority to ensuring local safety and synchronously reporting the fault information to the cloud management unit 1.

[0028] In this embodiment, in a microgrid with a high proportion of renewable energy access, equipment failures, line anomalies, and extreme working conditions (such as sudden load increase, external power grid impact, etc.) will have a serious impact on the system operation. In order to ensure that the microgrid can still quickly recover and maintain basic stability when anomalies occur, the present invention introduces an anomaly detection and fast self-healing strategy for multi-source data on the edge side and sub-group layer to form a multi-level fault prevention and control system. Its core ideas include: real-time identification of abnormal events, local emergency response, and multi-agent collaborative self-healing, striving to limit the impact of faults to the minimum range and quickly restore the system to normal operation. A lightweight anomaly detection mechanism based on threshold judgment is deployed in the edge node to scan the data in real time. If abnormal situations such as voltage, current, and equipment temperature that deviate significantly from the normal range are detected, local warnings are immediately triggered.

[0029] If the anomaly reaches a certain level (such as equipment failure or line short circuit), the edge node executes the pre-deployed local emergency strategy (such as cutting off the faulty line, entering the energy storage emergency power supply mode) to relieve the impact brought by local faults at the fastest speed. At the same time, the fault information or abnormal data is reported to the microgrid sub-group layer or the cloud to obtain broader linkage scheduling support. The basic measurement parameters include: voltage, current, active / reactive power, frequency, energy storage SOC (state of charge), equipment temperature, fault codes, etc.; when it is detected that the key measurement exceeds the normal operation threshold, a preliminary alarm is immediately triggered.

[0030] Let , representing the moment at which the multi-dimensional state vector collected by the edge node covers electrical quantities, equipment status, environmental information, etc.; In this embodiment, an anomaly index is used to quantify the deviation degree between the current state and the normal working mode. Let be the predicted or reference value in the "normal mode". The edge node is defined as: ; where represents the actual measured value of the th dimension in the state vector, represents the reference value, is the weight index of each index, is the total number of detection dimensions.

[0031] When exceeds a certain set threshold , the edge node determines that a potential anomaly has occurred and triggers a local emergency response or further inspection. This threshold can be adaptively adjusted according to historical data statistics and operation risk levels to achieve precise detection for diverse devices and environments.

[0032] If a serious anomaly is detected locally, such as a line short circuit, voltage dip, or energy storage overheating, etc., it will immediately report to the subgroup controller and perform preliminary isolation or protection actions by itself. The subgroup layer aggregates the alarm information of each edge node, judges the anomaly range, fault type, and emergency level. If it is found that multiple nodes are affected or cross-regional collaborative processing is required, a larger-scale linkage self-healing measure is initiated.

[0033] In this embodiment, the process of the localization emergency strategy is as follows: First, fault location and isolation are carried out. When the edge node detects a line fault or equipment anomaly and confirms it, it immediately cuts off the faulty branch or isolates the abnormal equipment from the main circuit to prevent the fault from spreading to other nodes. If a sudden drop in power generation side output or a sharp increase in load side power is detected, to avoid large fluctuations in voltage / frequency, the system can automatically activate the energy storage unit for emergency power supply or call for a backup power source (such as a diesel generator) to temporarily support the local load. In a fault or abnormal state, the present invention allows the edge node to temporarily make concessions to the original peak shaving and valley filling or economic objectives to improve safety and emergency priority. That is, some parameters of the original decision (such as the upper and lower limits of energy storage SOC, load transfer period) are dynamically relaxed to the safe range in the emergency mode; When a severe fault is detected, without waiting for an order from the upper layer, the edge node can directly perform the most basic open-circuit or load-shedding operations to ensure that local faults do not spread. If the abnormal influence range is large, the subgroup controller can uniformly call regional energy storage or standby power supply, and short-term dispatch the surplus energy of adjacent subgroups when necessary. The cloud management center records and analyzes the network-wide anomalies afterwards, such as equipment damage statistics, fault occurrence frequency, etc., to provide data support for long-term planning and model retraining.

[0034] Deploy an offline fallback mode. When the communication network is interrupted or the cloud management center fails, the subgroups and edge nodes can maintain basic self-healing and scheduling functions under local rules and reinforcement learning strategies to avoid system paralysis. Set conservative safety thresholds for key operating parameters (such as voltage, frequency, energy storage SOC). Once exceeded, the offline fallback mode is preferentially entered to limit part of the load or power generation to ensure the security of core power supply.

[0035] Among them, the lightweight reinforcement learning decision-making module 22 integrates the collected and monitored data to form the state input of reinforcement learning, deploys the pruned and quantized lightweight DQN model, and real-time infers the optimal actions based on local states (such as power generation, load demand, energy storage SOC), including the following steps: S2.1. Deploy the pruned and quantized lightweight DQN model on the edge node. This model is trained in the cloud and its volume and computational complexity are reduced through a series of optimization processes (such as pruning and quantization), making it suitable for running on resource-constrained edge devices for fast value calculation within the scheduling cycle of seconds and minutes; S2.2. When receiving the state vector at the current moment, use the lightweight DQN model for forward inference to calculate the Q-values of each action : ; Among them, represents the Q-value, which represents the expected cumulative reward that can be obtained by following a certain policy starting from the current moment after taking a specific action in a given state; represents using the lightweight DQN model to calculate the Q-values of all actions in a given state; S2.3. Select the action with the highest Q-value as the optimal action at the current moment : ; S2.4. Execute the optimal action , which involves operations such as adjusting the charge and discharge power of the energy storage system and changing the working mode of controllable loads, so as to achieve the purpose of peak shaving and valley filling, balancing supply and demand, and improving system stability; S2.5. Periodically upload the execution data and the execution effects to the cloud management unit 1. The execution effects include the change of energy storage SOC, grid operation indicators, fault handling results, etc., and send execution instructions to the microgrid subgroup unit 3.

[0036] The microgrid subgroup unit 3 collects the data of the edge node unit 2, forms a regional energy balance analysis, comprehensively schedules the power generation, energy storage and load in the local area, and introduces a multi-agent collaborative self-healing mechanism for fault isolation and self-healing; In this embodiment, the microgrid subgroup unit 3 collects the data of the edge nodes, forms a regional energy balance analysis, comprehensively schedules the power generation, energy storage and load in the local area, and introduces a multi-agent collaborative self-healing mechanism for fault isolation and self-healing, including the following steps: In this embodiment, several adjacent or functionally similar edge nodes are divided into a subgroup, and the subgroup controller or the dispatching center summarizes the data of the nodes under its jurisdiction to form a subgroup-level energy balance analysis (such as the matching situation of power generation, energy storage, and load). The subgroup controller can, according to the status and requirements of each node, uniformly coordinate the energy optimization scheduling or fault self-healing strategy within a local range. For example, in an emergency, it can allocate nodes with remaining energy storage capacity to compensate for the power supply of adjacent faulty nodes. At the subgroup level, the reinforcement learning decisions of each edge node are constrained and corrected to avoid the overall benefit decline caused by "selfish optimization". If greater range collaboration is required, the subgroup-level operation information and requirements can be reported to the cloud.

[0037] S3.1. Summarize the data of all edge nodes in the subgroup and calculate the energy balance situation at each time point in the area: ; Among them, represents the actual output power of the power generation subsystem at time ; represents the charge and discharge power of the energy storage subsystem; represents the power consumed by the load; represents the interaction power with the external power grid; S3.2. Establish the objective function and constraint conditions (including power balance constraint (ensuring local energy supply and demand balance), energy storage SOC constraint (ensuring the energy storage within a safe range), maximum power constraint (limiting the maximum charge and discharge capacity), load scheduling constraint (partial load adjustable)), and use the linear programming (LP) method for optimization to optimize the regional-level scheduling to ensure power balance and reduce costs; Furthermore, the objective function is: ; Among them, represents the The emergency decision-making of an edge node at time (such as energy storage discharge instruction, load shedding plan, etc.), is the fault risk or linkage loss function of the edge node in the emergency mode, is the energy imbalance degree or load transfer cost of node ; represents the weight coefficient of the fault risk; represents the weight coefficient of the energy imbalance degree, which is used to balance between ensuring the fault isolation efficiency and minimizing the power deviation; represents the index of the edge node; represents the total number of edge nodes; The inability to respond to the load sensitivity problem means that the power demands of industrial loads and residential loads in the microgrid show time-varying characteristics (such as differences between morning and evening peaks); the load mutation causes a chain reaction in the energy storage SOC and the inverter output, and the fixed weight cannot reflect the different sensitivities of loads to system stability at different times; the inability to respond to the environmental fluctuation problem means that the random fluctuations of photovoltaic / wind power (such as the sudden power drop caused by cloud occlusion) directly affect the power balance, but the economic weight in the traditional objective function is a fixed value and cannot dynamically adjust the optimization direction according to the fluctuation amplitude; the load demands (such as the start and stop of industrial equipment and the peak of residential electricity consumption) and renewable energy generation (such as photovoltaic and wind power) in the microgrid have significant spatio-temporal fluctuations. The traditional objective function with fixed weights is difficult to capture this dynamic change and may lead to scheduling deviations. Through the spatio-temporal attention mechanism, it can dynamically adjust the weight of the energy imbalance degree For example, increase during periods with high load sensitivity or severe environmental fluctuations, and give priority to balancing local energy; Aiming at the problem that the objective function cannot respond to load sensitivity and environmental fluctuations, generate dynamic weights through the spatio-temporal attention mechanism, and consider the reliability of model prediction, introduce a confidence factor into the objective function for optimization, Quantify the decision credibility based on the standard deviation of Q values to achieve the weight switching between the physical model and the data-driven method; ; There is an efficiency loss in the energy transmission between subgroups, which is manifested as the transmission efficiency matrix When the lines between subgroups are aged or the distance is far, the value of is low, resulting in a decrease in the utilization rate of the actual transmission power Introduce the objective function, and the system will automatically balance the transmission efficiency during optimization, preferentially select the interaction path with high efficiency, and reduce the overall energy loss; introduce the total power of energy exchange between subgroups and the adjustment coefficient , which can dynamically balance the costs of internal scheduling and external interaction. When the transmission efficiency between subgroups is low, increase to suppress frequent interactions and reduce energy loss; frequent cross-subgroup energy interactions may increase the system's dependence on the transmission network. For example, when a subgroup fails to transmit due to a fault, the subgroups that overly rely on interactions may face an energy shortage. After optimization, it can promote the internal energy self-balance of the subgroups, reduce the dependence on external transmission, and thus improve the power supply reliability of the local area; Regarding the problem of the energy interaction cost between subgroups, introduce the total power of energy exchange between subgroups into the objective function for optimization, model the energy transmission loss through the transmission efficiency matrix between subgroups to promote global optimization: ; ; ; wherein, represents the adjustment coefficient that controls the influence degree of the energy interaction between subgroups on the total objective; represents the subgroup to transmission efficiency matrix; represents the actual power transmitted from subgroup to ; and represent the indices of the subgroups; represents the device efficiency; represents the communication delay; represents the time decay constant; S3.3. When a fault is detected (such as energy storage overload, device anomaly, line short circuit), start the multi-agent collaborative self-healing mechanism. The subgroup controller analyzes the abnormal signals of the edge node unit 2 and performs fault isolation (disconnect the power connection of the fault node and isolate the fault line), compensates for the fault through the remaining edge nodes in the edge node unit 2, and synchronously uploads the fault control status to the cloud management unit 1 (let other energy storage devices increase discharge to make up for the power supply gap; let some controllable loads reduce the power demand) to ensure the stability of the local microgrid. This optimization can be quickly solved in the subgroup controller or the cloud, and coordinate the emergency actions of each node (such as energy storage scheduling, controllable load reduction, standby power grid connection time, etc.) through the downlink instruction to maximize the overall self-healing efficiency.

[0038] Embodiment 2: The difference between Embodiment 2 and Embodiment 1 of the present invention is that this embodiment introduces a lightweight reinforcement learning method used in a multi-layer distributed microgrid control system based on edge-cloud collaboration lightweight reinforcement learning.

[0039] A multi-layer distributed microgrid control method based on edge-cloud collaboration lightweight reinforcement learning, for a multi-layer distributed microgrid control system based on edge-cloud collaboration lightweight reinforcement learning in any one of the above, includes the following steps: S4.1. Train a DQN model using historical environment data, generate a lightweight DQN model through pruning and quantization, send the lightweight DQN model to the edge nodes, and formulate daily and weekly optimal scheduling strategies based on a function layer + hierarchical collaboration strategy; S4.2. Real-time monitor electrical quantities and environmental parameters through sensors, deploy the lightweight DQN model to monitor abnormal situations and trigger local early warning and emergency strategies, and make quick decisions using the lightweight DQN model; S4.3. Collect data from the edge nodes, form a regional energy balance analysis, comprehensively schedule power generation, energy storage, and loads in a local area, and introduce a multi-agent collaborative self-healing mechanism for fault isolation and self-healing.

[0040] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. A multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning, characterized in that: include: A cloud management unit (1), wherein the cloud management unit (1) uses historical environment data to train a DQN model, generates a lightweight DQN model through pruning and quantization, sends the lightweight DQN model to an edge node unit (2), and formulates a daily and weekly optimization scheduling strategy based on a functional layering + layered coordination strategy; An edge node unit (2), wherein the edge node unit (2) monitors electrical quantities and environmental parameters in real time through sensors, deploys a lightweight DQN model to monitor abnormal conditions and trigger local early warning and emergency strategies, and uses the lightweight DQN model to make rapid decisions; A microgrid subgroup unit (3) collects data from the edge node unit (2) to form a regional energy balance analysis, performs comprehensive scheduling of power generation, energy storage and load in a local area, introduces a confidence factor and the total power of energy exchange between subgroups to optimize the scheduling process, and introduces a multi-agent collaborative self-healing mechanism to perform fault isolation and self-healing.

2. The multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning according to claim 1 is characterized in that: The cloud management unit (1) comprises a global optimization and strategy generation module (11), a model training and compression module (12) and an abnormal linkage and fault tolerance management module (13); The global optimization and strategy generation module (11) stores and processes data of subgroups and external environments, and formulates daily and weekly optimization scheduling strategies based on functional stratification + stratification coordination strategies; The model training and compression module (12) trains the DQN model using historical environment data, generates a lightweight model through pruning and quantization, and periodically sends the pruned and quantized lightweight DQN model to the edge node unit (2); The abnormal linkage and fault-tolerance management module (13) analyzes abnormal events in the entire network, coordinates the sub-group layer to initiate cross-domain emergency plans, periodically optimizes the microgrid operation strategy, and issues emergency instructions to the microgrid sub-group unit (3).

3. The multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning according to claim 2 is characterized in that: The model training and compression module (12) uses historical environment data to train the DQN model and generates a lightweight model through pruning and quantization, including the following steps: S1.

1. Constructing the state vector of reinforcement learning : ; in, Indicates the output power of distributed generation side; Indicates the electrical power consumed by the load end; Indicates the charge state of the energy storage unit; represents the renewable energy surplus; Indicates voltage; Indicates current; Indicates the device temperature; Indicates time; Defining actions A collection of: ; in, Represents the set of all possible actions; Indicates specific action options; Indicates different kinds of actions; Represents the total number of actions in the action set; S1.

2. Use historical environment data to train the deep Q network DQN model offline in the cloud, design the reward function and iteratively update the Q value function; S1.3, determine the part of the DQN weight matrix whose contribution to the Q value output is less than the threshold a through redundancy analysis, and apply pruning technology to remove the part whose contribution is less than the threshold a; S1.

4. Perform small-scale fine-tuning on the pruned DQN model and convert the pruned floating-point weights into low-precision fixed-point numbers for quantization. S1.

5. Use sparse matrix storage technology to further compress the file size of the lightweight DQN model after pruning and quantization; S1.

6. Transmit the optimized lightweight DQN model to the edge node unit (2) through the encrypted communication protocol and the MQTT lightweight protocol.

4. The multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning according to claim 3 is characterized in that: In S1.1, the reward function is: ; in, Indicates immediate reward; Indicates the operating cost; Indicates this time Power balance deviation between generation, storage and load; represents the fault penalty term; An adjustable weight factor representing the operating cost; An adjustable weight factor representing the power balance deviation; represents the adjustable weight coefficient of the fault penalty term; The Q value function is: ; in, represents the parameter vector of the main network; is the target network parameter; represents the learning rate, is the discount factor.

5. The multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning according to claim 4 is characterized in that: The formulation of daily and weekly optimization scheduling strategies based on functional stratification + stratification coordination strategy includes the following steps: S1.7, obtain the historical operation data of the microgrid from the subgroup controller and collect meteorological data; S1.

8. Define the local multi-objective optimization function: ; in, represents the comprehensive objective function at the subgroup level; Indicates the electronic system; represents the energy storage subsystem; represents the load subsystem; Indicates Subsystem operating costs; Indicates Deviation of energy utilization of subsystems; Indicates Risk indicators of subsystems; Indicates the weight coefficient of operating cost; The weight coefficient representing the deviation of energy utilization rate; Indicates the weight coefficient of the risk indicator; Indicates the index of the subsystem; Define system constraints: Power Balance: ; Energy storage state of charge (SOC) limit: ; The energy storage charging and discharging power does not exceed the rated value: ; in, represents the actual output power of the power generation subsystem at time t; Indicates the charging and discharging power of the energy storage subsystem; Indicates the power consumed by the load; Represents the interactive power with the external power grid; Indicates that the energy storage system is at time The state of charge; Indicates the energy storage charging and discharging rated power constraint; S1.

9. Select the linear programming method to construct the optimization model; S1.

10. Based on historical data and meteorological data, use the deep learning LSTM model to predict renewable energy output and load demand in the future period; S1.

11. Use particle swarm optimization to solve the optimization model established above and find the optimal scheduling strategy; S1.

12. Generate a specific scheduling plan based on the solution results.

6. The multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning according to claim 5 is characterized in that: The edge node unit (2) comprises a data collection and preprocessing module (21), a lightweight reinforcement learning decision module (22) and a local anomaly detection and emergency response module (23); The data acquisition and preprocessing module (21) monitors electrical parameters through sensors, and uses edge nodes to monitor environmental parameters and equipment health status, and performs data preprocessing at the edge nodes; The lightweight reinforcement learning decision module (22) integrates the collected and monitored data to form a state input for reinforcement learning, deploys a pruned and quantized lightweight DQN model, and infers the optimal action in real time based on the local state; The local anomaly detection and emergency module (23) calculates an anomaly index by fusing multi-source data, uses a lightweight anomaly detection mechanism based on threshold judgment to trigger a local warning, executes a localized emergency strategy, prioritizes local safety, and simultaneously reports fault information to the cloud management unit (1).

7. The multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning according to claim 6 is characterized by: The lightweight reinforcement learning decision module (22) integrates the collected and monitored data to form the state input of reinforcement learning, deploys the pruned and quantized lightweight DQN model, and infers the optimal action in real time based on the local state, including the following steps: S2.

1. Deploy a pruned and quantized lightweight DQN model on the edge node to perform fast scheduling within seconds or minutes. Value calculation; S2.

2. When receiving the current state vector When , the lightweight DQN model is used for forward reasoning to calculate each action Q value: ; in, represents the Q value; Indicates the use of a lightweight DQN model to calculate the Q value of all actions in a given state; S2.

3. Select the action with the highest Q value as the optimal action at the current moment according to the Q value : ; S2.

4. Execute the best action ; S2.

5. Execute data The execution results are periodically uploaded to the cloud management unit (1), and execution instructions are sent to the microgrid subgroup units (3).

8. The multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning according to claim 7 is characterized in that: The microgrid subgroup unit (3) collects data from edge nodes, forms a regional energy balance analysis, performs comprehensive scheduling of power generation, energy storage and load in the local area, and introduces a multi-agent collaborative self-healing mechanism for fault isolation and self-healing, including the following steps: S3.

1. Summarize the data of all edge nodes in the subgroup and calculate the energy balance at each time point in the region: ; in, Indicates the electronic system at time The actual output power; Indicates the charging and discharging power of the energy storage subsystem; Indicates the power consumed by the load; Represents the interactive power with the external power grid; S3.

2. Establish the objective function and constraints, use linear programming method for optimization, and introduce the confidence factor in the objective function Total power of energy exchange between subgroups Optimize S3.

3. When a fault is detected, the multi-agent collaborative self-healing mechanism is activated, the sub-group controller analyzes the abnormal signal of the edge node unit (2) and isolates the fault, performs fault compensation through the remaining edge nodes in the edge node unit (2), and synchronously uploads the fault control status to the cloud management unit (1).

9. The multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning according to claim 8 is characterized in that: In S3.2, the objective function is: ; in, Indicates The edge nodes at time emergency decision making, For edge nodes Failure risk function in emergency mode, For Node The energy imbalance The weight coefficient representing the failure risk; The weight coefficient representing the energy imbalance; Indicates the index of the edge node; Indicates the total number of edge nodes; To solve the problem that the objective function cannot respond to load sensitivity and environmental fluctuations, dynamic weights are generated through the spatiotemporal attention mechanism. , and considering the model prediction reliability, the confidence factor is introduced into the objective function To optimize: ; In order to solve the energy interaction cost problem between subgroups, the total power of energy exchange between subgroups is introduced into the objective function. Further optimization: ; ; in, The adjustment coefficient represents the influence of energy interaction between control subgroups on the overall target; Represents subgroup arrive The transmission efficiency matrix of Represents subgroup Towards The actual power transmitted; and Indicates the index of the subgroup.

10. A multi-layer distributed microgrid control method based on edge-cloud collaborative lightweight reinforcement learning, used in a multi-layer distributed microgrid control system based on edge-cloud collaborative lightweight reinforcement learning as claimed in any one of claims 1 to 9, characterized in that: The steps include: S4.

1. Use historical environment data to train the DQN model, generate a lightweight DQN model through pruning and quantization, send the lightweight DQN model to the edge node, and formulate daily and weekly optimization scheduling strategies based on functional stratification + stratified collaboration strategy; S4.

2. Real-time monitoring of electrical quantities and environmental parameters through sensors, deployment of lightweight DQN models to monitor abnormal conditions and trigger local early warning and emergency strategies, and use lightweight DQN models to make quick decisions; S4.

3. Collect data from edge nodes to form a regional energy balance analysis, conduct comprehensive scheduling of power generation, energy storage and loads in the local area, and introduce a multi-agent collaborative self-healing mechanism for fault isolation and self-healing.

Citation Information

Cited By

  • Mobile energy storage vehicle energy management method based on intelligent algorithm

    CN120262410A

  • Cloud edge collaborative robot cluster simulation training and optimization system

    CN120524844A

  • Microgrid operation management method, device and equipment and readable storage medium

    CN120541495A

  • Image encryption system and method in smart power grid environment

    CN120547282A

  • Tunnel modular prefabricated cabin power supply and distribution intelligent substation self-adaptive regulation and control system based on edge calculation

    CN120657962A