Multi-time scale coordination control device and method for novel power system

Through multi-time-scale coordinated control devices, combined with deep reinforcement learning and transfer learning mechanisms, the problems of infeasible control strategies and insufficient data processing capabilities in new power systems have been solved, efficient and safe multi-objective optimization and complex scenario simulation have been achieved, and the stability and economy of the power system have been significantly improved.

CN120691608APending Publication Date: 2025-09-23STATE GRID GRID GANSU ELECTRIC POWER CO QINGYANG POWER SUPPLY CO
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510997695.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-19
Publication Date
2025-09-23

Smart Images

  • Figure CN120691608A_ABST
    Figure CN120691608A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-time-scale coordination control device and method for a novel electric power system, and relates to the technical field of electric power automation. Aiming at the problems of difficulty in multi-source data collaboration, complexity in control strategy optimization and insufficient scene adaptability in a novel power system, the device realizes flexible deployment and mobile operation through a high-strength aluminum alloy frame integrated modular functional assembly; the environment monitoring module is used for high-frequency collection of multi-parameter environment data, and equipment operation safety is guaranteed; a plurality of communication protocols are supported by means of a multi-time scale control module, and distributed energy equipment is seamlessly connected; based on three core units of a data processing center module, a scene simulation unit and a cooperative control center module, second-to-hour-level electric parameter data generation, complex scene simulation verification and intelligent generation, execution and optimization of a control strategy are realized; and finally, under the complex working condition of high-proportion new energy access, the control precision, the operation stability and the comprehensive benefits of the power system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electric power automation technology, and more particularly to a multi-time scale coordinated control device and method for a novel electric power system. Background Art

[0002] With the large-scale integration of renewable energy and the widespread use of distributed energy devices, new power systems are characterized by a high proportion of intermittent power sources, complex topologies, and diverse load characteristics. Traditional power system control methods face numerous challenges. On the one hand, the volatility and randomness of renewable energy generation exacerbate the difficulty of balancing system power, significantly increasing the response speed and precision requirements for real-time frequency regulation and voltage stability control. On the other hand, the decentralized and diverse nature of distributed energy resources further complicates the coordinated control of multiple timescales. Traditional single-timescale control strategies struggle to meet the coordinated demands of second-level rapid regulation, minute-level rolling optimization, and hour-level planned scheduling.

[0003] At the same time, power system operation is subject to multiple physical constraints, such as node voltage amplitude ranges, phase angle difference limits, and thermal power unit ramp rate constraints. Traditional control algorithms struggle to effectively integrate these constraints, leading to strategy infeasibility and system operational risks. Furthermore, power system operation scenarios are diverse and varied, with a wide variety of fault types. Existing control devices often lack the ability to simulate and verify complex scenarios, making it difficult to ensure the reliability and robustness of control strategies.

[0004] In terms of data processing and communication, the multi-source, heterogeneous data generated by new power systems is characterized by high-frequency acquisition and diverse protocols. Traditional devices have limited data acquisition, transmission, and processing capabilities, and cannot meet real-time and high-efficiency requirements. Furthermore, when evaluating and optimizing control strategies, traditional methods struggle to comprehensively consider the multi-objective optimization requirements of economy, stability, and low carbon efficiency. They lack dynamic adjustment mechanisms and are unable to adapt to the complex operating environment of new power systems. Therefore, there is an urgent need to develop a new device with multi-timescale coordinated control capabilities, the ability to effectively integrate physical constraints, and the ability to achieve multi-objective optimization, in order to improve the operating efficiency and reliability of new power systems. Summary of the Invention

[0005] In view of the deficiencies of the existing technology, the present invention discloses a multi-time-scale coordinated control method and device for a new type of power system.

[0006] The present invention adopts the following technical solutions: A multi-time-scale coordinated control device for a new power system, comprising: The main frame of the device is a high-strength aluminum alloy rectangular parallelepiped structure, with hydraulic lifting support legs and heavy-duty directional wheels set at the bottom of the main frame; an environmental monitoring module is set at the top of the main frame; The environmental monitoring module integrates temperature, humidity, dust, vibration and light intensity sensors to achieve high-frequency acquisition of multiple parameters and transmit data via RS485 and Modbus-RTU protocols; A composite heat dissipation structure is provided on the side of the main frame, and the composite heat dissipation structure includes a variable frequency heat dissipation fan and a heat dissipation structure air inlet; The front of the main frame is equipped with a modular functional panel, equipped with a multi-timescale control module and a system interaction module. The multi-timescale control module has the functions of real-time control, rolling optimization and scheduling, supports the communication protocol transmission functions of IEC61850, Modbus-TCP, OPC UA and MQTT, and adapts to the access requirements of distributed energy DER equipment. The system interaction module is equipped with a touch screen, which displays the system operation status and control strategy through a three-dimensional visual interface. The remote collaboration interface enables real-time monitoring and collaborative control by remote terminals such as mobile phones and tablets. The main framework includes a data processing hub module, a scene simulation unit and a collaborative control hub module; The data processing hub module generates electrical parameter data at different time scales from seconds to hours based on the real-time operation data and historical data of the new power system, including simulations of ≥20 complex operation scenarios. It also has data playback and comparison capabilities, supports historical scenario reproduction, and is used for strategy debugging and fault tracing. The scenario simulation unit simulates operating scenarios such as grid frequency fluctuations, relay protection actions, energy storage charging and discharging strategy conflicts, renewable energy output fluctuations, load mutations, and equipment failures to test the effectiveness of multi-timescale coordinated control strategies. The communication link uses SSL / TLS encryption to prevent data tampering and malicious attacks. The collaborative control center module is responsible for the formulation, execution and optimization of multi-time scale control strategies, with real-time control strategy generation, rolling optimization scheme adjustment, and planned scheduling arrangements; the collaborative control center module includes a strategy generation unit, an execution monitoring unit, a data acquisition unit, an effect evaluation unit, and a strategy optimization unit; the strategy generation unit is based on the new power system topology and operating characteristics, and adopts the deep reinforcement learning DRL algorithm that integrates physical constraints to construct a multi-time scale control strategy template, including a transfer learning mechanism, and uses historical scenario training data to pre-initialize strategy network parameters; the execution monitoring unit monitors the execution of the control strategy in real time; the data acquisition unit A distributed sensor network collects system operation data; the effect evaluation unit uses a hierarchical analysis method to perform a multi-dimensional evaluation of the control strategy effect, combines it with grey correlation analysis to locate performance shortcomings, and generates optimization suggestions; the strategy optimization unit constructs a multi-objective optimization model based on a non-dominated sorting genetic algorithm II, and simultaneously optimizes the three objective functions of economy, stability, and low carbon; the output end of the strategy generation unit is connected to the input end of the execution monitoring unit, the output end of the data acquisition unit is connected to the input end of the effect evaluation unit, and the strategy generation unit and the strategy optimization unit are interconnected; the output end of the effect evaluation unit is connected to the input end of the execution monitoring unit; The multi-time scale control module is interconnected with the collaborative control central module; the data processing central module is interconnected with the collaborative control central module; the scene simulation unit is interconnected with the collaborative control central module; and the system interaction module is interconnected with the collaborative control central module.

[0007] As a further technical solution of the present invention, the environmental monitoring module includes multi-parameter sensors for temperature and humidity, dust concentration, vibration amplitude and light intensity, and realizes data transmission with a frequency of more than 10Hz through multi-channel synchronous acquisition and RS485 and Modbus-RTU protocols; when the environment is abnormal, it triggers an audible and visual alarm and synchronously starts heat dissipation or power reduction protection.

[0008] As a further technical solution of the present invention, the strategy generation unit includes a DRL algorithm module integrating physical constraints, a transfer learning mechanism module, a multi-time scale strategy template generation module, and a strategy evaluation and screening module; The strategy generation unit includes a DRL algorithm module integrating physical constraints, a transfer learning mechanism module, a multi-time scale strategy template generation module, and a strategy evaluation and screening module; The DRL algorithm module integrating physical constraints adopts the deep deterministic policy gradient (DDPG) framework, in which the Actor network consists of six fully connected layers with the number of neurons in the order of 256-128-64-32-16-action space dimensions, and the Critic network adopts a dual-Q structure. The physical constraints are embedded in the reward function through the augmented Lagrange multiplier method, including: voltage amplitude constraint, phase angle difference constraint, thermal power unit ramp rate constraint, and photovoltaic curtailment rate constraint. The node voltage amplitude constraint dynamically adjusts the initial value to 10 and updates the Lagrange multiplier every 500 steps to ensure that the voltage operates within the range of 0.95-1.05. The phase angle difference constraint limits the phase angle difference to within 30 degrees, and a penalty mechanism is triggered when it exceeds the range. The thermal power unit ramp rate constraint is that the ramp rate does not exceed 5% of the rated power per minute, and the minimum start and stop time is 4 hours and 2 hours respectively. The photovoltaic curtailment rate constraint is that the photovoltaic power station curtailment rate does not exceed 5%, and the wind farm power generation utilization rate is not less than 95%. The implementation method of the transfer learning mechanism module includes scenario similarity calculation and parameter migration; the scenario similarity calculation is based on the system topology structure of the node number and line connection relationship, the load peak-valley difference, and the operating characteristic parameters of the proportion of new energy to calculate the cosine similarity, and migration is triggered when the similarity is ≥70%; the layer-by-layer migration strategy fully reuses the parameters of the underlying feature extraction layer, and inherits the parameters of the upper decision layer at a ratio of 0.7; the adaptive learning rate is the initial learning rate, set to 0.0005, and uses a cosine annealing scheduler with a decay of 0.98 per training cycle; The implementation process of the multi-timescale policy template generation module includes: second-level policy and hour-level policy. The second-level policy is for frequency fluctuation scenarios, using an improved gated recurrent unit (GRU) network combined with an attention mechanism to generate a second-level policy; the prediction time domain is 30 seconds and the control period is 100 milliseconds; the hour-level policy is for load peak and valley scheduling, using a graph convolutional network (GCN) combined with a spatiotemporal transformer architecture to construct an hour-level policy, with a prediction time domain of 24 hours and a control period of 15 minutes. The strategy evaluation and screening module generates 10 candidate strategy templates for each training cycle. Through Monte Carlo Tree Search (MCTS) evaluation, the top three with the highest average cumulative rewards are selected to enter the verification phase. Templates with a verification pass rate of ≥90% are confirmed as valid strategy templates. The output segment of the DRL algorithm module integrating physical constraints is connected to the input end of the transfer learning mechanism module, the output end of the transfer learning mechanism module is connected to the input end of the multi-time scale policy template generation module, and the output end of the multi-time scale policy template generation module is connected to the input end of the policy evaluation and screening module.

[0009] As a further technical solution of the present invention, the distributed sensor network of the data acquisition unit includes: optical fiber current sensors deployed on transmission lines with a measurement accuracy of 0.1 and a sampling frequency ≥ 200 Hz; a meteorological sensor group installed at a new energy station to collect wind speed, light intensity, and temperature data in real time; a smart meter set on the load side that supports three-phase four-wire measurement; all sensors transmit data through a wireless sensor network based on the IEEE802.15.4 protocol, with a transmission delay of ≤ 30 ms.

[0010] As a further technical solution of the present invention, the effect evaluation unit includes a hierarchical analysis model, a grey relational analysis model, performance short board positioning rules and optimization suggestion generation; The AHP model is an evaluation framework that constructs a three-level indicator system. The first-level indicators include safety (weighted between 0.35 and 0.55), economy (weighted between 0.25 and 0.45), and environmental protection (weighted between 0.15 and 0.3). Second-level indicators include frequency stability (weighted between 0.4 and 0.6 for safety), voltage compliance (weighted between 0.3 and 0.5), operating cost (weighted between 0.5 and 0.7 for economy), network loss rate (weighted between 0.2 and 0.4), and carbon emission intensity (weighted between 0.5 and 0.7 for environmental protection). Third-level indicators are further refined to include equipment-level parameters such as transformer load factor and line transmission efficiency. The grey correlation analysis uses an improved Deng's correlation algorithm to locate performance shortcomings by calculating the correlation between each evaluation object and the ideal solution. When the correlation is less than 0.6, a deep analysis process is triggered to identify the specific indicators that cause the shortcomings. The performance shortcoming identification rule is as follows: if the frequency stability correlation is less than 0.5 and the frequency deviation standard deviation is greater than 0.2Hz, it is identified as insufficient frequency regulation capability; if the economic correlation is less than 0.5 and the proportion of purchased electricity costs is greater than 70%, it is identified as an unreasonable energy structure; if the environmental protection correlation is less than 0.5 and the carbon emission intensity is greater than 0.5kg / kWh, it is identified as a need to optimize the low-carbon scheduling strategy; The optimization suggestion generation is to generate optimization suggestions for specific improvement measures based on the short board positioning results.

[0011] As a further technical solution of the present invention, in the strategy optimization unit, the working method of the multi-objective optimization model constructed based on the non-dominated sorting genetic algorithm II is: setting the population size to 150, the maximum number of iterations to 80, the crossover probability to 0.8, and the mutation probability to 0.02; the economic objective function aims to minimize the system operating cost, including the electricity purchase cost, equipment operation and maintenance cost, and energy storage charging and discharging loss cost; the stability objective function aims to minimize the sum of squared frequency deviations and the voltage offset, of which the frequency deviation weight accounts for 0.4 and the voltage offset weight accounts for 0.6; the low-carbon objective function aims to minimize the total carbon emissions of the system; through congestion distance calculation and fast non-dominated sorting mechanism, a Pareto front solution set is generated, and when the system new energy penetration rate exceeds 60%, the weight of the low-carbon objective function is automatically increased to above 0.5 to adapt to the optimization needs under the scenario of high proportion of new energy access.

[0012] As a further technical solution of the present invention, S1, real-time acquisition and preprocessing of multi-source heterogeneous data The environmental monitoring module on top of the device's main frame uses temperature, humidity, dust, vibration, and light intensity sensors to collect equipment operating environment data at a frequency of more than 10Hz. At the same time, the data acquisition unit drives the distributed sensor network to obtain power system operating data. Environmental data is transmitted to the data processing hub module via RS485 and Modbus-RTU protocols. Power system data is uploaded through a multi-time scale control module that supports IEC61850, Modbus-TCP, OPCUA, and MQTT protocols. The data processing hub module denoises and normalizes the raw data, and generates second-level real-time data, minute-level statistical data, and hour-level forecast data based on historical data. It simultaneously simulates ≥20 complex operating scenarios and supports data playback and fault tracing. S2. Hierarchical Multi-Time-Scale Control Strategy Generation: The strategy generation unit uses a deep reinforcement learning (DRL) algorithm that integrates physical constraints to capture the dynamic characteristics of the system for different emergency scenarios. It uses a 100ms control cycle and a 30-second prediction time domain to generate real-time adjustment instructions. Based on minute-level statistical data and 15-minute new energy output and load forecasts, it combines the attention mechanism to build a spatiotemporal prediction model. Rolling optimization is performed through the multi-time-scale control module, updating the distributed energy scheduling plan every 5 minutes to coordinate electric vehicle charging, load start-stop adjustment, and microgrid power interaction. Based on hourly forecast data and current market transaction results, the strategy generation unit models the grid topology, formulates unit combinations, inter-provincial power transmission, and energy storage charging and discharging plans for the next 24 hours, and makes strategy adjustments in 15-minute cycles. S3. Hierarchical collaborative execution of control strategies The collaborative control center module distributes the generated strategy through the multi-timescale control module. The execution monitoring unit tracks the instruction execution status with a delay of ≤50ms. The built-in priority scheduler ensures zero-delay queue execution of emergency fault instructions. A dual-ring buffer is used to implement instruction interaction. The real-time control layer prioritizes responding to second-level instructions, directly acting on the fast-moving devices of energy storage converters and inverters. The minute-level optimization strategy adjusts the operating status of DERs through the EMS. The hour-level scheduling plan is synchronized to the dispatch centers at all levels. The real-time control is frequently triggered, sending warnings to the minute-level optimization layer to dynamically adjust the energy storage charging and discharging plan. S4. Dynamic evaluation of control effects and strategy optimization The effectiveness evaluation unit uses the Analytic Hierarchy Process (AHP) to construct a three-level evaluation system encompassing safety, economy, and environmental protection. It combines grey correlation analysis to calculate the correlation between the strategy and the ideal solution, identifying performance shortcomings. The strategy optimization unit, based on the non-dominated sorting genetic algorithm II, simultaneously optimizes the three objective functions of economy, stability, and low carbon. Based on the evaluation results, the real-time control layer adaptively adjusts the strategy parameters in a second-level cycle, the rolling optimization layer in a minute-level cycle, and the planning and scheduling layer in an hour-level cycle to generate a Pareto frontier solution set and improve the overall system performance. S5. Scenario simulation verification and safety assurance The scenario simulation unit reproduces typical working conditions of grid frequency fluctuations, relay protection actions, and energy storage strategy conflicts, injects historical fault data, and uses digital twin technology to simulate the effectiveness of control strategies, discovering potential risks of command conflicts and equipment overloads in advance; data transmission uses SSL / TLS encryption to prevent data tampering and attacks; the device hardware integrates overvoltage / overcurrent protection circuits, and the environmental monitoring module monitors the equipment operating environment in real time. In the event of an abnormality, it triggers an audible and visual alarm and links heat dissipation or power reduction protection to ensure safe and stable operation of the system. As a further technical solution of the present invention, the implementation steps of the deep reinforcement learning DRL algorithm to construct a multi-time scale control strategy template are: S1, algorithm framework construction The algorithm is based on the Deep Deterministic Policy Gradient (DDPG) framework, which consists of an Actor network and a Critic network. The Actor network takes the current state s of the system as input, calculates the corresponding control action a through internal network calculation, and its network parameters are expressed as ,Right now , which aims to generate effective strategies that can influence the operation of the system; the Critic network takes the state s and action a as input, evaluates the value of performing this action in the state, and outputs , used to guide Actor network optimization strategies and improve action value; S2. Physical Constraint Embedding In order to make the strategy generated by the algorithm meet the actual operation requirements of the power system, the adaptive hierarchical Lagrange multiplier method is used to incorporate physical constraints into the reward function R; the formula is: In formula (1), Represents the original reward of the task, used to measure the system's performance in terms of stability and economy; is a time-varying Lagrange multiplier, which is dynamically updated by the following formula: in To constrain the learning rate of i, control the multiplier update step size, is the gradient contribution of constraint i to the total loss; is the constraint priority weight; is the penalty function for violating the i-th constraint. When the system operation state or action violates the corresponding physical constraint, Producing non-zero values ​​to reduce the reward function R, thereby prompting the algorithm to avoid generating strategies that violate the constraints; As the integral penalty coefficient, the historical cumulative effect of constraint violations is introduced; The historical cumulative violation of constraint i to avoid short-term oscillation; S3, key constraint implementation 1) Voltage amplitude constraint In the power system, the node voltage amplitude needs to be maintained in a reasonable range; through the formula: In formula (2), the voltage amplitude constraint penalty term is calculated as: is the voltage amplitude of node i; is the reference voltage value, usually set to 1.0pu; The allowable voltage deviation range is generally 0.05pu; is the voltage fluctuation adjustment coefficient, which dynamically adjusts the penalty intensity according to the voltage change rate; 2) Phase angle difference constraint If the phase angle difference between nodes is too large, it will affect the stability of the system. Formula (3) is used to constrain the phase angle difference: In formula (3), are the voltage phase angles at nodes i and j respectively; is the maximum allowed phase angle difference; E is the system line set; 3) Thermal power unit ramp rate constraint There is a limit on the power change speed of thermal power units, and the ramp rate is constrained by formula (4): In formula (4), is the active power of thermal power unit i at time t; is the time interval; is the maximum ramp rate, which is 5% of the rated power per minute; is the total number of thermal power units; is the basic power of the thermal power unit; is the rated power of the thermal power unit; 4) PV abandonment rate constraint In order to improve the utilization rate of photovoltaic energy, the abandoned light rate is constrained by formula (5): In formula (5), is the amount of electricity that can be generated by the photovoltaic power station; is the actual photovoltaic power generation; The maximum allowable abandoned light rate is 5%; is the proportion of photovoltaics in total energy; S4, transfer learning mechanism 1) Scene similarity calculation Calculate the current scene With historical scenes The similarity formula is: In formula (6), the eigenvector x contains the number of system nodes , Number of lines , new energy penetration rate , and the load fluctuation standard deviation Key parameters of are the time spans of the current scene and the historical scene respectively; is the preset maximum time span; is the time weight coefficient, which adjusts the influence of time factors on similarity; 2) Parameter migration After triggering transfer learning, the following formula is used for parameter migration: In formula (7), is the migration ratio (0.7); is the historical training parameter; are random initialization parameters; S5. Multi-timescale strategy generation 1) Second-level strategy For the scenario of rapid frequency fluctuation of the system, the improved gated recurrent unit GRU combined with the attention mechanism is used to generate a second-level strategy; the system state of the past n time steps is calculated by formula (8). and the previous hidden state As input, the network automatically focuses on key state information through the attention mechanism, captures the dynamic characteristics of the system, and outputs second-level control actions. , to achieve fast response regulation; formula (8) is: 2) Hourly Strategy For long-term scenarios involving load peak and valley scheduling, an hourly strategy is constructed using an architecture combining graph convolutional networks (GCNs) and spatiotemporal transformers. The formula is: In formula (9), G is the graph structure representing the power system, which includes the power system node set V and the power system line edge set E; is the system state sequence of the past m moments; S6, training optimization 1) Target network update In order to improve the stability of algorithm training, the target network mechanism is adopted; the target network parameters of the Critic network are updated through formula (10): ; In formula (10), It is a dynamic soft update coefficient, which is dynamically adjusted according to the training error and gradient changes; is the current critic network parameter; 2) Critic network loss calculation The Critic network optimizes its own parameters by calculating the loss function. The formula is: In formula (11), r is the immediate reward; is a discount factor used to measure the importance of future rewards; is the regularization coefficient; is the L2 norm of the Critic network parameters.

[0013] Positive beneficial effects A multi-time-scale coordinated control device for new power systems achieves multi-dimensional positive and beneficial effects through deep integration of innovative architecture and intelligent algorithms: at the system operation level, the multi-time-scale coordinated control mechanism achieves full-cycle coverage from 100ms real-time adjustment to 24-hour planned scheduling, so that the system frequency deviation is controlled within ±0.2Hz, the voltage fluctuation is stabilized at ±5% of the rated value, and the compliance rate of the thermal power unit climbing rate is increased to 99%, significantly enhancing the stability of the power system; in terms of new energy consumption, the photovoltaic curtailment rate is reduced to below 3%, and the wind farm utilization rate is increased to 97%, greatly improving the utilization efficiency of clean energy.

[0014] In terms of strategy optimization and decision-making, the DRL algorithm that integrates physical constraints is combined with the transfer learning mechanism to increase strategy generation efficiency by 40% and shorten the training cycle by 35%. The multi-objective optimization model simultaneously optimizes economy, stability and low carbon performance, and is expected to reduce system operating costs by 12%-18% and reduce carbon emissions by more than 15%. At the same time, the device uses SSL / TLS encrypted communication and overvoltage / overcurrent protection hardware design, combined with real-time early warning of the environmental monitoring module, to improve data transmission security to financial-grade standards, shorten equipment failure response time to seconds, and fully guarantee the safe and reliable operation of the system. In addition, the three-dimensional visual interactive interface and remote collaborative control function have achieved an increase in operation and maintenance efficiency of more than 50%, effectively promoting the intelligent and efficient development of new power systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which: Figure 1 This is a device structure diagram of a multi-time scale coordinated control device for a new power system; Figure 2 This is the overall flow chart of a multi-time-scale coordinated control device for a new type of power system; Figure 3 This is a flow chart of the collaborative control hub module of a multi-time-scale coordinated control device for a new type of power system; Figure 4 Flowchart of the multi-time-scale coordinated control method for new power systems. DETAILED DESCRIPTION

[0016] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0017] according to Figure 1-4 , a multi-time-scale coordinated control device for a new type of power system, comprising: The main frame 1 of the device is a high-strength aluminum alloy rectangular parallelepiped structure, and a hydraulic lifting support leg 2 and a heavy-duty directional wheel 3 are provided at the bottom of the main frame 1; an environmental monitoring module 4 is provided at the top of the main frame 1; The environmental monitoring module integrates temperature, humidity, dust, vibration and light intensity sensors to achieve high-frequency acquisition of multiple parameters and transmit data via RS485 and Modbus-RTU protocols; A composite heat dissipation structure 5 is provided on the side of the main frame 1, and the composite heat dissipation structure 5 includes a variable frequency heat dissipation fan and a heat dissipation structure air inlet; The front of the main frame 1 is provided with a modular functional panel, equipped with a multi-time scale control module 6 and a system interaction module 7; the multi-time scale control module 6 has the functions of real-time control, rolling optimization and scheduling, supports the communication protocol transmission functions of IEC61850, Modbus-TCP, OPCUA and MQTT, and adapts to the access requirements of distributed energy DER equipment; the system interaction module 7 is equipped with a touch screen 8, which displays the system operation status and control strategy through a three-dimensional visual interface, and the remote collaboration interface enables real-time monitoring and collaborative control by remote terminals such as mobile phones and tablets; The main frame 1 includes a data processing central module, a scene simulation unit and a collaborative control central module; The data processing hub module generates electrical parameter data at different time scales from seconds to hours based on the real-time operation data and historical data of the new power system, including simulations of ≥20 complex operation scenarios. It also has data playback and comparison capabilities, supports historical scenario reproduction, and is used for strategy debugging and fault tracing. The scenario simulation unit simulates operating scenarios such as grid frequency fluctuations, relay protection actions, energy storage charging and discharging strategy conflicts, renewable energy output fluctuations, load mutations, and equipment failures to test the effectiveness of multi-timescale coordinated control strategies. The communication link uses SSL / TLS encryption to prevent data tampering and malicious attacks. The collaborative control center module is responsible for the formulation, execution and optimization of multi-time scale control strategies, with real-time control strategy generation, rolling optimization scheme adjustment, and planned scheduling arrangements; the collaborative control center module includes a strategy generation unit, an execution monitoring unit, a data acquisition unit, an effect evaluation unit, and a strategy optimization unit; the strategy generation unit is based on the new power system topology and operating characteristics, and adopts the deep reinforcement learning DRL algorithm that integrates physical constraints to construct a multi-time scale control strategy template, including a transfer learning mechanism, and uses historical scenario training data to pre-initialize strategy network parameters; the execution monitoring unit monitors the execution of the control strategy in real time; the data acquisition unit A distributed sensor network collects system operation data; the effect evaluation unit uses a hierarchical analysis method to perform a multi-dimensional evaluation of the control strategy effect, combines it with grey correlation analysis to locate performance shortcomings, and generates optimization suggestions; the strategy optimization unit constructs a multi-objective optimization model based on a non-dominated sorting genetic algorithm II, and simultaneously optimizes the three objective functions of economy, stability, and low carbon; the output end of the strategy generation unit is connected to the input end of the execution monitoring unit, the output end of the data acquisition unit is connected to the input end of the effect evaluation unit, and the strategy generation unit and the strategy optimization unit are interconnected; the output end of the effect evaluation unit is connected to the input end of the execution monitoring unit; The multi-time scale control module is interconnected with the collaborative control central module; the data processing central module is interconnected with the collaborative control central module; the scene simulation unit is interconnected with the collaborative control central module; and the system interaction module is interconnected with the collaborative control central module.

[0018] Furthermore, the environmental monitoring module includes multi-parameter sensors for temperature and humidity, dust concentration, vibration amplitude and light intensity, which realize data transmission with a frequency of more than 10Hz through multi-channel synchronous acquisition and RS485 and Modbus-RTU protocols; when the environment is abnormal, it triggers an audible and visual alarm and simultaneously starts heat dissipation or power reduction protection.

[0019] In a specific embodiment, the environmental monitoring module integrates multi-parameter sensors for temperature and humidity, dust concentration, vibration amplitude, and light intensity. Using multi-channel synchronous acquisition technology, it acquires real-time data on the equipment's operating environment at a frequency exceeding 10Hz. This data is then transmitted to the data processing hub via the RS485 communication interface and Modbus-RTU protocol. When an abnormal environmental parameter is detected, such as a temperature exceeding the normal operating threshold of 45°C, a dust concentration exceeding the safety standard of 10mg / m³, a vibration amplitude exceeding the permitted range of 5mm / s, or an abnormal change in light intensity, the monitoring module immediately triggers an audible and visual alarm and simultaneously sends a warning signal to the collaborative control hub. Upon receiving the signal, the collaborative control hub automatically initiates heat dissipation or power reduction protection procedures based on preset rules. If the temperature is too high, the variable-frequency cooling fan is controlled to increase speed and the heat dissipation structure's air inlet is opened to enhance air convection. If the dust concentration exceeds the standard, power reduction protection is activated to reduce the equipment's operating load and minimize the risk of failure. Through the above implementation steps, the environmental monitoring module can realize high-frequency, multi-dimensional monitoring of the equipment operating environment, perceive potential risks in advance, and reduce the probability of equipment failure due to environmental factors by more than 60%. At the same time, it ensures the timeliness and accuracy of data collection, provides reliable environmental data support for the formulation of multi-time scale coordinated control strategies, and significantly improves the safety and stability of device operation.

[0020] Furthermore, the strategy generation unit includes a DRL algorithm module integrating physical constraints, a transfer learning mechanism module, a multi-time scale strategy template generation module, and a strategy evaluation and screening module; The strategy generation unit includes a DRL algorithm module integrating physical constraints, a transfer learning mechanism module, a multi-time scale strategy template generation module, and a strategy evaluation and screening module; The DRL algorithm module integrating physical constraints adopts the deep deterministic policy gradient (DDPG) framework, in which the Actor network consists of six fully connected layers with the number of neurons in the order of 256-128-64-32-16-action space dimensions, and the Critic network adopts a dual-Q structure. The physical constraints are embedded in the reward function through the augmented Lagrange multiplier method, including: voltage amplitude constraint, phase angle difference constraint, thermal power unit ramp rate constraint, and photovoltaic curtailment rate constraint. The node voltage amplitude constraint dynamically adjusts the initial value to 10 and updates the Lagrange multiplier every 500 steps to ensure that the voltage operates within the range of 0.95-1.05. The phase angle difference constraint limits the phase angle difference to within 30 degrees, and a penalty mechanism is triggered when it exceeds the range. The thermal power unit ramp rate constraint is that the ramp rate does not exceed 5% of the rated power per minute, and the minimum start and stop time is 4 hours and 2 hours respectively. The photovoltaic curtailment rate constraint is that the photovoltaic power station curtailment rate does not exceed 5%, and the wind farm power generation utilization rate is not less than 95%. The implementation method of the transfer learning mechanism module includes scenario similarity calculation and parameter migration; the scenario similarity calculation is based on the system topology structure of the node number and line connection relationship, the load peak-valley difference, and the operating characteristic parameters of the proportion of new energy to calculate the cosine similarity, and migration is triggered when the similarity is ≥70%; the layer-by-layer migration strategy fully reuses the parameters of the underlying feature extraction layer, and inherits the parameters of the upper decision layer at a ratio of 0.7; the adaptive learning rate is the initial learning rate, set to 0.0005, and uses a cosine annealing scheduler with a decay of 0.98 per training cycle; The implementation process of the multi-timescale policy template generation module includes: second-level policy and hour-level policy. The second-level policy is for frequency fluctuation scenarios, using an improved gated recurrent unit (GRU) network combined with an attention mechanism to generate a second-level policy; the prediction time domain is 30 seconds and the control period is 100 milliseconds; the hour-level policy is for load peak and valley scheduling, using a graph convolutional network (GCN) combined with a spatiotemporal transformer architecture to construct an hour-level policy, with a prediction time domain of 24 hours and a control period of 15 minutes. The strategy evaluation and screening module generates 10 candidate strategy templates for each training cycle. Through Monte Carlo Tree Search (MCTS) evaluation, the top three with the highest average cumulative rewards are selected to enter the verification phase. Templates with a verification pass rate of ≥90% are confirmed as valid strategy templates. The output segment of the DRL algorithm module integrating physical constraints is connected to the input end of the transfer learning mechanism module, the output end of the transfer learning mechanism module is connected to the input end of the multi-time scale policy template generation module, and the output end of the multi-time scale policy template generation module is connected to the input end of the policy evaluation and screening module.

[0021] In a specific embodiment, the strategy generation unit achieves intelligent generation of multi-timescale control strategies for power systems through the collaboration of multiple modules. This system, centered on a DRL algorithm that incorporates physical constraints and employs a transfer learning mechanism to improve training efficiency, generates multi-timescale strategy templates to meet control requirements across the system's time dimensions. Strategy evaluation and screening ensure strategy effectiveness. First, the DRL algorithm module, which integrates physical constraints, is based on the DDPG framework. The physical constraints of voltage amplitude and phase angle difference are embedded in the reward function using the augmented Lagrange multiplier method. The actor and critic networks employ a specific structure to perform policy generation and value evaluation. Next, the transfer learning mechanism module calculates scenario cosine similarity based on system topology and operational characteristic parameters. When the similarity is ≥70%, the policy parameters are transferred, with the bottom layer fully reused and the upper layer 70% inherited. The learning rate is adjusted using a cosine annealing scheduler. The multi-timescale policy template generation module utilizes the GRU-attention network and GCN-spatiotemporal Transformer architectures, respectively, to generate second- and hour-level control policies. Finally, the policy evaluation and screening module evaluates 10 candidate templates in each training cycle using MCTS. The top three with the highest average cumulative reward are selected for verification, and the templates with a pass rate ≥90% are considered valid policies. This implementation improves policy generation efficiency by 40%, shortens training cycles by 35%, increases system voltage control accuracy by 20%, and increases the renewable energy consumption rate to 97%, effectively enhancing the operational stability and economic efficiency of the power system. Table 1 shows the core performance indicators and comprehensive optimization results of each module in the strategy generation unit. The physical constraint DRL module significantly improves constraint satisfaction while reducing the number of training steps through the augmented Lagrange multiplier method. The transfer learning mechanism effectively shortens the training cycle in scenarios with a similarity of 70% or more. Multi-timescale strategy templates achieve high-precision control in different time domains. Strategy evaluation and screening ensure a high pass rate. Data shows that this unit outperforms traditional methods in terms of convergence efficiency, constraint satisfaction, generalization capability, and resource consumption. The comprehensive optimization cycle is shortened by 40%, and the key constraint satisfaction rate is ≥99%, verifying its effectiveness and robustness in complex power systems.

[0022] Furthermore, the distributed sensor network of the data acquisition unit includes: fiber optic current sensors deployed on transmission lines with a measurement accuracy of 0.1 and a sampling frequency ≥ 200 Hz; meteorological sensor groups installed at new energy stations to collect wind speed, light intensity, and temperature data in real time; smart meters set on the load side to support three-phase four-wire measurement; all sensors transmit data through a wireless sensor network based on the IEEE802.15.4 protocol, with a transmission delay of ≤ 30 ms.

[0023] In a specific embodiment, the data acquisition unit achieves high-precision, low-latency acquisition and transmission of multi-dimensional data from the power system through a distributed sensor network. A variety of high-precision sensors are integrated to build a wireless sensor network based on the IEEE802.15.4 protocol, enabling real-time capture and efficient transmission of data from transmission lines, renewable energy stations, and loads. First, fiber optic current sensors are deployed on the transmission lines. Utilizing their 0.1-level measurement accuracy and ≥200Hz sampling frequency, they collect line current data in real time, accurately reflecting the power system's current distribution. Second, meteorological sensors, including wind speed, light intensity, and temperature sensors, are installed at renewable energy stations to acquire environmental parameters in real time, providing basic data for renewable energy output forecasts. Simultaneously, smart meters supporting three-phase, four-wire measurement are deployed on the load side to collect electrical parameters such as active / reactive power, voltage, and current, enabling refined monitoring of load status. Finally, all sensors transmit data through a wireless sensor network built using the IEEE802.15.4 protocol. The protocol's low power consumption and short latency ensure data transmission delays of ≤30ms, meeting real-time control requirements. Through this implementation, the data acquisition unit achieves high-performance data acquisition with a transmission line current measurement error of ≤0.1%, a new energy environmental parameter acquisition delay of ≤20ms, and a load data update frequency of seconds. This provides the collaborative control hub module with multi-source real-time data covering grid operation, new energy output, and load status, supporting the precise generation of control strategies at the second to hour level. Field tests have shown that this solution reduces system state estimation error by 45%, increases new energy output prediction accuracy to 92%, and improves load fluctuation response speed by 60%, significantly enhancing the perception and control efficiency of the new power system.

[0024] Furthermore, the effect evaluation unit includes a hierarchical analysis model, a grey relational analysis model, performance shortcoming location rules, and optimization suggestion generation; The AHP model is an evaluation framework that constructs a three-level indicator system. The first-level indicators include safety (weighted between 0.35 and 0.55), economy (weighted between 0.25 and 0.45), and environmental protection (weighted between 0.15 and 0.3). Second-level indicators include frequency stability (weighted between 0.4 and 0.6 for safety), voltage compliance (weighted between 0.3 and 0.5), operating cost (weighted between 0.5 and 0.7 for economy), network loss rate (weighted between 0.2 and 0.4), and carbon emission intensity (weighted between 0.5 and 0.7 for environmental protection). Third-level indicators are further refined to include equipment-level parameters such as transformer load factor and line transmission efficiency. The grey correlation analysis uses an improved Deng's correlation algorithm to locate performance shortcomings by calculating the correlation between each evaluation object and the ideal solution. When the correlation is less than 0.6, a deep analysis process is triggered to identify the specific indicators that cause the shortcomings. The performance shortcoming identification rule is as follows: if the frequency stability correlation is less than 0.5 and the frequency deviation standard deviation is greater than 0.2Hz, it is identified as insufficient frequency regulation capability; if the economic correlation is less than 0.5 and the proportion of purchased electricity costs is greater than 70%, it is identified as an unreasonable energy structure; if the environmental protection correlation is less than 0.5 and the carbon emission intensity is greater than 0.5kg / kWh, it is identified as a need to optimize the low-carbon scheduling strategy; The optimization recommendations are generated based on the shortcoming identification results, generating specific improvement measures. In a specific embodiment, the performance evaluation unit implements multi-dimensional quantitative analysis and precise optimization of the control strategy by constructing a hierarchical and logically rigorous evaluation system. By integrating a hierarchical analysis model with a gray correlation analysis algorithm, a hierarchical indicator system is used to comprehensively evaluate safety, economy, and environmental performance. Performance shortcomings are identified using correlation calculations, and targeted optimization recommendations are generated based on pre-set rules. First, a hierarchical analysis model is constructed based on a three-level indicator system, assigning weights to safety, economy, and environmental performance according to weight ranges. This is broken down into three levels of indicators, each representing device-level parameters, comprehensively covering key dimensions of system operation. Second, an improved Deng's correlation algorithm is used to calculate the correlation between each evaluation target and the ideal solution. When the correlation is less than 0.6, an in-depth analysis is initiated. Next, based on the performance shortcoming identification rules, the correlation is compared with key parameter thresholds to accurately identify specific shortcomings such as insufficient frequency regulation capacity and an unreasonable energy structure. Finally, based on the identification results, specific optimization recommendations are generated, such as adjusting energy storage charging and discharging strategies to improve frequency regulation capacity and increasing new energy consumption to optimize the energy structure. The implementation method can improve the efficiency of control strategy evaluation by 50%, locate performance shortcomings with an accuracy rate of over 92%, improve system frequency stability by 30%, reduce operating costs by 18%, and reduce carbon emission intensity by 22%, significantly enhancing the comprehensive operating performance and sustainability of the new power system.

[0025] Furthermore, in the strategy optimization unit, the working method of the multi-objective optimization model constructed based on the non-dominated sorting genetic algorithm II is as follows: the population size is set to 150, the maximum number of iterations is 80, the crossover probability is 0.8, and the mutation probability is 0.02; the economic objective function aims to minimize the system operating cost, including the electricity purchase cost, equipment operation and maintenance cost, and energy storage charging and discharging loss cost; the stability objective function aims to minimize the sum of squared frequency deviations and voltage offsets, of which the frequency deviation weight accounts for 0.4 and the voltage offset weight accounts for 0.6; the low-carbon objective function aims to minimize the total carbon emissions of the system; through congestion distance calculation and fast non-dominated sorting mechanism, a Pareto front solution set is generated, and when the system new energy penetration rate exceeds 60%, the weight of the low-carbon objective function is automatically increased to above 0.5 to adapt to the optimization needs under the scenario of high proportion of new energy access.

[0026] In a specific embodiment, the strategy optimization unit constructs a multi-objective optimization model using a non-dominated sorting genetic algorithm II to achieve a dynamic balance between economy, stability, and low-carbon performance. To address the multi-objective optimization challenges faced by new power systems, an improved genetic algorithm is employed, dynamically adjusting objective weights and employing a population evolution mechanism to generate a Pareto-optimal solution set. First, the population size is initialized to 150, with a maximum iteration count of 80, a crossover probability of 0.8, and a mutation probability of 0.02. An economy objective function is constructed, encompassing the costs of electricity purchase, operation and maintenance, and energy storage losses. A stability objective function integrates frequency deviation with a weight of 0.4 and voltage offset with a weight of 0.6. A low-carbon objective function focuses on minimizing total carbon emissions. During algorithm operation, individuals are divided into levels using a fast non-dominated sorting algorithm. The congestion distance is combined to maintain solution diversity, and the Pareto frontier is generated through generational evolution. When the system's renewable energy penetration rate exceeds 60%, the low-carbon objective weight is automatically increased to above 0.5 to enhance green scheduling. This implementation method achieved, under standard test scenarios, the following: 1) a 12%-18% reduction in operating costs, 2) frequency deviation within ±0.15Hz, and 3) a 15%-20% reduction in carbon emissions. In a high-renewable energy penetration scenario exceeding 60%, the increased weighting of low-carbon objectives reduced the system's wind and solar curtailment rate from 8% to below 3%, validating the algorithm's adaptability and robustness. Compared to traditional single-objective optimization, this solution improves multi-objective balancing by over 30%, providing flexible and efficient intelligent decision-making support for new power systems. Table 2 shows the performance of the NSGA-II-based multi-objective optimization model under different renewable energy penetration scenarios. In terms of economics, multi-objective optimization reduces total costs by 3.6% by coordinating energy storage losses with electricity purchase costs. Among stability indicators, frequency deviation and voltage offset decrease by 33.3% and 28.9%, respectively, significantly improving system operation quality. The low-carbon goal achieves a 21.6% emission reduction in a high-proportion renewable energy scenario, validating the effectiveness of dynamic adjustment of objective weights. Furthermore, the algorithm converges 10% faster than traditional methods, enhances solution diversity, and improves decision-making efficiency by 16.7%. The data demonstrates that the model outperforms traditional methods in multi-objective balancing, adaptation to high-proportion renewable energy, and decision-making efficiency, providing a more comprehensive solution for power system optimization.

[0027] Furthermore, a multi-time scale coordinated control device method for a new power system is characterized by comprising the following steps: S1, real-time acquisition and preprocessing of multi-source heterogeneous data The environmental monitoring module on top of the device's main frame uses temperature, humidity, dust, vibration, and light intensity sensors to collect equipment operating environment data at a frequency of more than 10Hz. At the same time, the data acquisition unit drives the distributed sensor network to obtain power system operating data. Environmental data is transmitted to the data processing hub module via RS485 and Modbus-RTU protocols. Power system data is uploaded through a multi-time scale control module that supports IEC61850, Modbus-TCP, OPCUA, and MQTT protocols. The data processing hub module denoises and normalizes the raw data, and generates second-level real-time data, minute-level statistical data, and hour-level forecast data based on historical data. It simultaneously simulates ≥20 complex operating scenarios and supports data playback and fault tracing. S2. Generation of hierarchical multi-timescale control strategies The strategy generation unit uses a deep reinforcement learning (DRL) algorithm that integrates physical constraints to capture the dynamic characteristics of the system for different emergency scenarios. It generates real-time adjustment instructions with a control cycle of 100ms and a prediction time domain of 30 seconds. Based on minute-level statistical data and new energy output and load forecasts within 15 minutes, it combines the attention mechanism to build a spatiotemporal prediction model. Through the multi-time scale control module, it performs rolling optimization and updates the distributed energy scheduling plan every 5 minutes to coordinate electric vehicle charging, load start and stop regulation, and microgrid power interaction. Based on hourly forecast data and current market transaction results, the strategy generation unit models the grid topology and formulates unit combination, inter-provincial power transmission, and energy storage charging and discharging plans for the next 24 hours, and makes strategy adjustments in a 15-minute cycle. S3. Hierarchical collaborative execution of control strategies The collaborative control center module distributes the generated strategy through the multi-timescale control module. The execution monitoring unit tracks the instruction execution status with a delay of ≤50ms. The built-in priority scheduler ensures zero-delay queue execution of emergency fault instructions. A dual-ring buffer is used to implement instruction interaction. The real-time control layer prioritizes responding to second-level instructions, directly acting on the fast-moving devices of energy storage converters and inverters. The minute-level optimization strategy adjusts the operating status of DERs through the EMS. The hour-level scheduling plan is synchronized to the dispatch centers at all levels. The real-time control is frequently triggered, sending warnings to the minute-level optimization layer to dynamically adjust the energy storage charging and discharging plan. S4. Dynamic evaluation of control effects and strategy optimization The effectiveness evaluation unit uses the Analytic Hierarchy Process (AHP) to construct a three-level evaluation system encompassing safety, economy, and environmental protection. It combines grey correlation analysis to calculate the correlation between the strategy and the ideal solution, identifying performance shortcomings. The strategy optimization unit, based on the non-dominated sorting genetic algorithm II, simultaneously optimizes the three objective functions of economy, stability, and low carbon. Based on the evaluation results, the real-time control layer adaptively adjusts the strategy parameters in a second-level cycle, the rolling optimization layer in a minute-level cycle, and the planning and scheduling layer in an hour-level cycle to generate a Pareto frontier solution set and improve the overall system performance. S5. Scenario simulation verification and safety assurance The scenario simulation unit reproduces typical operating conditions of grid frequency fluctuations, relay protection actions, and energy storage strategy conflicts, injects historical fault data, and uses digital twin technology to simulate the effectiveness of control strategies, discovering potential risks of command conflicts and equipment overloads in advance; data transmission uses SSL / TLS encryption to prevent data tampering and attacks; the device hardware integrates overvoltage / overcurrent protection circuits, and the environmental monitoring module monitors the equipment operating environment in real time. In the event of an abnormality, it triggers an audible and visual alarm and links heat dissipation or power reduction protection to ensure safe and stable operation of the system. Furthermore, the implementation steps of the deep reinforcement learning DRL algorithm to construct a multi-time scale control strategy template are: S1, algorithm framework construction The algorithm is based on the Deep Deterministic Policy Gradient (DDPG) framework, which consists of an Actor network and a Critic network. The Actor network takes the current state s of the system as input, calculates the corresponding control action a through internal network calculation, and its network parameters are expressed as ,Right now , which aims to generate effective strategies that can influence the operation of the system; the Critic network takes the state s and action a as input, evaluates the value of performing this action in the state, and outputs , used to guide Actor network optimization strategies and improve action value; S2. Physical Constraint Embedding In order to make the strategy generated by the algorithm meet the actual operation requirements of the power system, the adaptive hierarchical Lagrange multiplier method is used to incorporate physical constraints into the reward function R; the formula is: In formula (1), Represents the original reward of the task, used to measure the system's performance in terms of stability and economy; is a time-varying Lagrange multiplier, which is dynamically updated by the following formula: in To constrain the learning rate of i, control the multiplier update step size, is the gradient contribution of constraint i to the total loss; is the constraint priority weight; is the penalty function for violating the i-th constraint. When the system operation state or action violates the corresponding physical constraint, Producing non-zero values ​​to reduce the reward function R, thereby prompting the algorithm to avoid generating strategies that violate the constraints; As the integral penalty coefficient, the historical cumulative effect of constraint violations is introduced; The historical cumulative violation of constraint i to avoid short-term oscillation; S3, key constraint implementation 1) Voltage amplitude constraint In the power system, the node voltage amplitude needs to be maintained in a reasonable range; through the formula: In formula (2), the voltage amplitude constraint penalty term is calculated as: is the voltage amplitude of node i; is the reference voltage value, usually set to 1.0pu; The allowable voltage deviation range is generally 0.05pu; is the voltage fluctuation adjustment coefficient, which dynamically adjusts the penalty intensity according to the voltage change rate; 2) Phase angle difference constraint If the phase angle difference between nodes is too large, it will affect the stability of the system. Formula (3) is used to constrain the phase angle difference: In formula (3), are the voltage phase angles at nodes i and j respectively; is the maximum allowed phase angle difference; E is the system line set; 3) Thermal power unit ramp rate constraint There is a limit on the power change speed of thermal power units, and the ramp rate is constrained by formula (4): In formula (4), is the active power of thermal power unit i at time t; is the time interval; is the maximum ramp rate, which is 5% of the rated power per minute; is the total number of thermal power units; is the basic power of the thermal power unit; is the rated power of the thermal power unit; 4) PV abandonment rate constraint In order to improve the utilization rate of photovoltaic energy, the abandoned light rate is constrained by formula (5): In formula (5), is the amount of electricity that can be generated by the photovoltaic power station; is the actual photovoltaic power generation; The maximum allowable abandoned light rate is 5%; is the proportion of photovoltaics in total energy; S4, transfer learning mechanism 1) Scene similarity calculation Calculate the current scene With historical scenes The similarity formula is: In formula (6), the eigenvector x contains the number of system nodes , Number of lines , new energy penetration rate , and the load fluctuation standard deviation Key parameters of are the time spans of the current scene and the historical scene respectively; is the preset maximum time span; is the time weight coefficient, which adjusts the influence of time factors on similarity; 2) Parameter migration After triggering transfer learning, the following formula is used for parameter migration: In formula (7), is the migration ratio (0.7); is the historical training parameter; are random initialization parameters; S5. Multi-timescale strategy generation 1) Second-level strategy For the scenario of rapid frequency fluctuation of the system, the improved gated recurrent unit GRU combined with the attention mechanism is used to generate a second-level strategy; the system state of the past n time steps is calculated by formula (8). and the previous hidden state As input, the network automatically focuses on key state information through the attention mechanism, captures the dynamic characteristics of the system, and outputs second-level control actions. , to achieve fast response regulation; formula (8) is: 2) Hourly Strategy For long-term scenarios involving load peak and valley scheduling, an hourly strategy is constructed using an architecture combining graph convolutional networks (GCNs) and spatiotemporal transformers. The formula is: In formula (9), G is the graph structure representing the power system, which includes the power system node set V and the power system line edge set E; is the system state sequence of the past m moments; S6, training optimization 1) Target network update In order to improve the stability of algorithm training, the target network mechanism is adopted; the target network parameters of the Critic network are updated through formula (10): ; In formula (10), It is a dynamic soft update coefficient, which is dynamically adjusted according to the training error and gradient changes; is the current critic network parameter; 2) Critic network loss calculation The Critic network optimizes its own parameters by calculating the loss function. The formula is: In formula (11), r is the immediate reward; is a discount factor used to measure the importance of future rewards; is the regularization coefficient; is the L2 norm of the critic network parameters. In a specific embodiment, this embodiment constructs a multi-time-scale control strategy template based on the deep reinforcement learning (DRL) algorithm that integrates physical constraints. Based on the DDPG framework, combined with adaptive hierarchical Lagrange multipliers, transfer learning and multi-time-scale network architecture, multi-objective optimization control of the power system is realized. First, the DDPG framework is built, and the Actor network and the Critic network are responsible for strategy generation and value evaluation respectively; then, the adaptive hierarchical Lagrange multiplier method is used to embed the key physical constraints of voltage amplitude and phase angle difference into the reward function, and the penalty intensity is dynamically adjusted through time-varying multipliers and historical cumulative effects; in the transfer learning stage, the scenario similarity is calculated based on the system topology and operating characteristic parameters, and when the similarity meets the standard, the network parameters are migrated according to the hierarchical strategy; then, for second-level and hour-level scenarios, the GRU-attention network and GCN-space-time Transformer architectures are used to generate control strategies respectively; finally, the critic network training process is optimized by dynamically adjusting the target network update coefficient and introducing regularization terms. Field validation has shown that the algorithm has increased the system voltage compliance rate to 99.8%, reduced the number of thermal power unit ramp violations by 85%, and lowered the photovoltaic curtailment rate to less than 3%. Furthermore, the transfer learning mechanism has increased the algorithm's training efficiency by 40% in new scenarios. The multi-timescale strategy has effectively enhanced the system's responsiveness to frequency fluctuations and load peaks and valleys, significantly improving the operational stability and economic efficiency of the new power system. In Table 3, the dynamic penalty coefficient in formula (2) , adjust the voltage constraint strength in real time, significantly improve the voltage qualification rate and reduce the risk of over-limit. Formula (3) introduces the cosine function to strengthen the critical phase angle difference penalty, reducing the stability risk of the system caused by phase problems. Formula (4) combines the unit's basic power to dynamically adjust the penalty weight to effectively suppress ramp violations. Formula (5) adjusts the penalty intensity according to the proportion of photovoltaic power to promote the efficient consumption of new energy. The transfer learning mechanism reduces the training time of new scenarios through scene similarity calculation (Formula (6)) and hierarchical parameter migration (Formula (7)). Based on the NSGA-II algorithm, the economy, stability and low carbon are optimized simultaneously, and AHP and gray correlation analysis assist in accurately locating the optimization direction. The second-level strategy adopts the GRU-attention network (Formula (8)) to quickly capture the frequency fluctuation characteristics and achieve millisecond-level response. When the system topology or operating conditions change, the transfer learning mechanism ensures that the algorithm adapts quickly and maintains the control performance.

[0028] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these specific embodiments are merely illustrative, and that those skilled in the art may omit, substitute, and modify the details of the methods and systems described above without departing from the principles and spirit of the present invention. For example, combining the above method steps to perform substantially the same functions in substantially the same manner to achieve substantially the same results falls within the scope of the present invention. Accordingly, the scope of the present invention is limited solely by the appended claims.

Claims

1. A multi-time scale coordinated control device for a new type of power system, comprising a main frame (1), wherein the main frame (1) is a high-strength aluminum alloy rectangular parallelepiped structure, and a hydraulic lifting support leg (2) and a heavy-duty directional wheel (3) are provided at the bottom of the main frame (1); and an environmental monitoring module (4) is provided at the top of the main frame (1); Its characteristics are: The environmental monitoring module integrates temperature, humidity, dust, vibration and light intensity sensors to achieve high-frequency acquisition of multiple parameters and transmit data via RS485 and Modbus-RTU protocols; A composite heat dissipation structure (5) is provided on the side of the main frame (1), and the composite heat dissipation structure (5) comprises a variable frequency heat dissipation fan and a heat dissipation structure air inlet; The main frame (1) is provided with a modular function panel on the front, equipped with a multi-time scale control module (6) and a system interaction module (7); the multi-time scale control module (6) has the functions of real-time control, rolling optimization and scheduling, supports the communication protocol transmission functions of IEC61850, Modbus-TCP, OPC UA and MQTT, and adapts to the access requirements of distributed energy DER equipment; the system interaction module (7) is equipped with a touch screen (8), which displays the system operation status and control strategy through a three-dimensional visual interface, and the remote collaboration interface enables remote terminals of mobile phones and tablets to monitor and control in real time; The main frame (1) includes a data processing central module, a scene simulation unit and a collaborative control central module; The data processing hub module generates electrical parameter data at different time scales from seconds to hours based on the real-time operation data and historical data of the new power system, including simulations of ≥20 complex operation scenarios. It also has data playback and comparison capabilities, supports historical scenario reproduction, and is used for strategy debugging and fault tracing. The scenario simulation unit simulates operating scenarios such as grid frequency fluctuations, relay protection actions, energy storage charging and discharging strategy conflicts, renewable energy output fluctuations, load mutations, and equipment failures to test the effectiveness of multi-timescale coordinated control strategies. The communication link uses SSL / TLS encryption to prevent data tampering and malicious attacks. The collaborative control hub module includes a strategy generation unit, an execution monitoring unit, a data acquisition unit, an effect evaluation unit, and a strategy optimization unit; The strategy generation unit, based on the new power system topology and operating characteristics, uses a deep reinforcement learning (DRL) algorithm that integrates physical constraints to construct a multi-timescale control strategy template, including a transfer learning mechanism, and uses historical scenario training data to pre-initialize strategy network parameters. The execution monitoring unit monitors the execution of the control strategy in real time. The data acquisition unit collects system operating data through a distributed sensor network. The effect evaluation unit uses the hierarchical analysis method to conduct a multi-dimensional evaluation of the control strategy effect, combines it with grey correlation analysis to identify performance shortcomings, and generates optimization suggestions. The strategy optimization unit constructs a multi-objective optimization model based on the non-dominated sorting genetic algorithm II, and simultaneously optimizes the three objective functions of economy, stability and low carbon performance; the output end of the strategy generation unit is connected to the input end of the execution monitoring unit, the output end of the data acquisition unit is connected to the input end of the effect evaluation unit, and the strategy generation unit and the strategy optimization unit are interconnected; the output end of the effect evaluation unit is connected to the input end of the execution monitoring unit; The multi-time scale control module is interconnected with the collaborative control central module; the data processing central module is interconnected with the collaborative control central module; the scene simulation unit is interconnected with the collaborative control central module; and the system interaction module is interconnected with the collaborative control central module.

2. A multi-time-scale coordinated control device for a new power system according to claim 1, characterized in that: The environmental monitoring module includes multi-parameter sensors for temperature and humidity, dust concentration, vibration amplitude, and light intensity. It realizes data transmission with a frequency of more than 10Hz through multi-channel synchronous acquisition and RS485 and Modbus-RTU protocols. When the environment is abnormal, it triggers an audible and visual alarm and simultaneously starts heat dissipation or power reduction protection.

3. The multi-time scale coordinated control device for a new power system according to claim 1, characterized in that: The strategy generation unit includes a DRL algorithm module integrating physical constraints, a transfer learning mechanism module, a multi-time scale strategy template generation module, and a strategy evaluation and screening module; The strategy generation unit includes a DRL algorithm module integrating physical constraints, a transfer learning mechanism module, a multi-time scale strategy template generation module, and a strategy evaluation and screening module; The DRL algorithm module integrating physical constraints adopts the deep deterministic policy gradient (DDPG) framework, in which the Actor network consists of six fully connected layers with the number of neurons in the order of 256-128-64-32-16-action space dimensions, and the Critic network adopts a dual-Q structure. The physical constraints are embedded in the reward function through the augmented Lagrange multiplier method, including: voltage amplitude constraint, phase angle difference constraint, thermal power unit ramp rate constraint, and photovoltaic curtailment rate constraint. The node voltage amplitude constraint dynamically adjusts the initial value to 10 and updates the Lagrange multiplier every 500 steps to ensure that the voltage operates within the range of 0.95-1.

05. The phase angle difference constraint limits the phase angle difference to within 30 degrees, and a penalty mechanism is triggered when it exceeds the range. The thermal power unit ramp rate constraint is that the ramp rate does not exceed 5% of the rated power per minute, and the minimum start and stop time is 4 hours and 2 hours respectively. The photovoltaic curtailment rate constraint is that the photovoltaic power station curtailment rate does not exceed 5%, and the wind farm power generation utilization rate is not less than 95%. The implementation method of the transfer learning mechanism module includes scenario similarity calculation and parameter migration; the scenario similarity calculation is based on the system topology structure of the node number and line connection relationship, the load peak-valley difference, and the operating characteristic parameters of the proportion of new energy to calculate the cosine similarity, and migration is triggered when the similarity is ≥70%; the layer-by-layer migration strategy fully reuses the parameters of the underlying feature extraction layer, and inherits the parameters of the upper decision layer at a ratio of 0.7; the adaptive learning rate is the initial learning rate, set to 0.0005, and uses a cosine annealing scheduler with a decay of 0.98 per training cycle; The implementation process of the multi-timescale policy template generation module includes: second-level policy and hour-level policy. The second-level policy is for frequency fluctuation scenarios, using an improved gated recurrent unit (GRU) network combined with an attention mechanism to generate a second-level policy; the prediction time domain is 30 seconds and the control period is 100 milliseconds; the hour-level policy is for load peak and valley scheduling, using a graph convolutional network (GCN) combined with a spatiotemporal transformer architecture to construct an hour-level policy, with a prediction time domain of 24 hours and a control period of 15 minutes. The strategy evaluation and screening module generates 10 candidate strategy templates for each training cycle. Through Monte Carlo Tree Search (MCTS) evaluation, the top three with the highest average cumulative rewards are selected to enter the verification phase. Templates with a verification pass rate of ≥90% are confirmed as valid strategy templates. The output segment of the DRL algorithm module integrating physical constraints is connected to the input end of the transfer learning mechanism module, the output end of the transfer learning mechanism module is connected to the input end of the multi-time scale policy template generation module, and the output end of the multi-time scale policy template generation module is connected to the input end of the policy evaluation and screening module.

4. The multi-time scale coordinated control device for a new power system according to claim 1, characterized in that: The distributed sensor network of the data acquisition unit includes: fiber optic current sensors deployed on transmission lines with a measurement accuracy of 0.1 and a sampling frequency of ≥200Hz; a meteorological sensor group installed at new energy stations to collect wind speed, light intensity, and temperature data in real time; and smart meters set on the load side that support three-phase four-wire measurement. All sensors transmit data through a wireless sensor network based on the IEEE802.15.4 protocol, with a transmission delay of ≤30ms.

5. The multi-time-scale coordinated control device for a new power system according to claim 1, characterized in that: The effect evaluation unit includes a hierarchical analysis model, a grey correlation analysis model, performance short board positioning rules and optimization suggestion generation; The AHP model is an evaluation framework that constructs a three-level indicator system. The first-level indicators include safety (weighted between 0.35 and 0.55), economy (weighted between 0.25 and 0.45), and environmental protection (weighted between 0.15 and 0.3). Second-level indicators include frequency stability (weighted between 0.4 and 0.6 for safety), voltage compliance (weighted between 0.3 and 0.5), operating cost (weighted between 0.5 and 0.7 for economy), network loss rate (weighted between 0.2 and 0.4), and carbon emission intensity (weighted between 0.5 and 0.7 for environmental protection). Third-level indicators are further refined to include equipment-level parameters such as transformer load factor and line transmission efficiency. The grey correlation analysis uses an improved Deng's correlation algorithm to locate performance shortcomings by calculating the correlation between each evaluation object and the ideal solution. When the correlation is less than 0.6, a deep analysis process is triggered to identify the specific indicators that cause the shortcomings. The performance shortcoming identification rule is as follows: if the frequency stability correlation is less than 0.5 and the frequency deviation standard deviation is greater than 0.2Hz, it is identified as insufficient frequency regulation capability; if the economic correlation is less than 0.5 and the proportion of purchased electricity costs is greater than 70%, it is identified as an unreasonable energy structure; if the environmental protection correlation is less than 0.5 and the carbon emission intensity is greater than 0.5kg / kWh, it is identified as a need to optimize the low-carbon scheduling strategy; The optimization suggestion generation is to generate optimization suggestions for specific improvement measures based on the short board positioning results.

6. According to claim 1, a multi-time scale coordinated control device for a new power system is characterized in that, in the strategy optimization unit, the working method of the multi-objective optimization model constructed based on the non-dominated sorting genetic algorithm II is: setting the population size to 150, the maximum number of iterations to 80, the crossover probability to 0.8, and the mutation probability to 0.02; the economic objective function aims to minimize the system operating cost, including the electricity purchase cost, equipment operation and maintenance cost, and energy storage charging and discharging loss cost; the stability objective function aims to minimize the sum of squared frequency deviations and the voltage offset, where the frequency deviation weight accounts for 0.4 and the voltage offset weight accounts for 0.6; the low-carbon objective function aims to minimize the total carbon emissions of the system; through congestion distance calculation and fast non-dominated sorting mechanism, a Pareto front solution set is generated, and when the system new energy penetration rate exceeds 60%, the weight of the low-carbon objective function is automatically increased to above 0.5 to adapt to the optimization needs under the high proportion of new energy access scenarios.

7. A multi-time scale coordinated control method for a new power system, characterized in that: A multi-time-scale coordinated control device for a new power system according to any one of claims 1 to 7 is characterized in that: The following steps are involved: S1. Real-time collection and preprocessing of multi-source heterogeneous data The environmental monitoring module on top of the device's main frame uses temperature, humidity, dust, vibration, and light intensity sensors to collect equipment operating environment data at a frequency of more than 10Hz. At the same time, the data acquisition unit drives the distributed sensor network to obtain power system operating data. Environmental data is transmitted to the data processing hub module via RS485 and Modbus-RTU protocols. Power system data is uploaded through a multi-time scale control module that supports IEC61850, Modbus-TCP, OPCUA, and MQTT protocols. The data processing hub module denoises and normalizes the raw data, and generates second-level real-time data, minute-level statistical data, and hour-level forecast data based on historical data. It simultaneously simulates ≥20 complex operating scenarios and supports data playback and fault tracing. S2. Hierarchical multi-timescale control strategy generation The strategy generation unit uses a deep reinforcement learning (DRL) algorithm that integrates physical constraints to capture the dynamic characteristics of the system for different emergency scenarios. It generates real-time adjustment instructions with a control cycle of 100ms and a prediction time domain of 30 seconds. Based on minute-level statistical data and new energy output and load forecasts within 15 minutes, it combines the attention mechanism to build a spatiotemporal prediction model. Through the multi-time scale control module, it performs rolling optimization and updates the distributed energy scheduling plan every 5 minutes to coordinate electric vehicle charging, load start and stop regulation, and microgrid power interaction. Based on hourly forecast data and current market transaction results, the strategy generation unit models the grid topology and formulates unit combination, inter-provincial power transmission, and energy storage charging and discharging plans for the next 24 hours, and makes strategy adjustments in a 15-minute cycle. S3. Hierarchical collaborative execution of control strategies The collaborative control center module distributes the generated strategy through the multi-timescale control module. The execution monitoring unit tracks the instruction execution status with a delay of ≤50ms. The built-in priority scheduler ensures zero-delay queue execution of emergency fault instructions. A dual-ring buffer is used to implement instruction interaction. The real-time control layer prioritizes responding to second-level instructions, directly acting on the fast-moving devices of energy storage converters and inverters. The minute-level optimization strategy adjusts the operating status of DERs through the EMS. The hour-level scheduling plan is synchronized to the dispatch centers at all levels. The real-time control is frequently triggered, sending warnings to the minute-level optimization layer to dynamically adjust the energy storage charging and discharging plan. S4. Dynamic evaluation of control effects and strategy optimization The effectiveness evaluation unit uses the Analytic Hierarchy Process (AHP) to construct a three-level evaluation system encompassing safety, economy, and environmental protection. It combines grey correlation analysis to calculate the correlation between the strategy and the ideal solution, identifying performance shortcomings. The strategy optimization unit, based on the non-dominated sorting genetic algorithm II, simultaneously optimizes the three objective functions of economy, stability, and low carbon. Based on the evaluation results, the real-time control layer adaptively adjusts the strategy parameters in a second-level cycle, the rolling optimization layer in a minute-level cycle, and the planning and scheduling layer in an hour-level cycle to generate a Pareto frontier solution set and improve the overall system performance. S5. Scenario simulation verification and safety assurance The scenario simulation unit reproduces typical operating conditions of grid frequency fluctuations, relay protection actions, and energy storage strategy conflicts, injects historical fault data, and uses digital twin technology to simulate the effectiveness of control strategies, thereby discovering potential risks of command conflicts and equipment overloads in advance; data transmission uses SSL / TLS encryption to prevent data tampering and attacks; the device hardware integrates overvoltage / overcurrent protection circuits, and the environmental monitoring module monitors the equipment operating environment in real time. In the event of an abnormality, it triggers an audible and visual alarm and links heat dissipation or power reduction protection to ensure safe and stable operation of the system.

8. The multi-time-scale coordinated control method for a novel power system according to claim 7, characterized in that: The implementation steps of the deep reinforcement learning (DRL) algorithm to construct a multi-timescale control strategy template are as follows: S1. Algorithm framework construction The algorithm is based on the Deep Deterministic Policy Gradient (DDPG) framework, which consists of an Actor network and a Critic network. The Actor network takes the current state s of the system as input, calculates the corresponding control action a through internal network calculation, and its network parameters are expressed as ,Right now , which aims to generate effective strategies that can influence the operation of the system; the Critic network takes the state s and action a as input, evaluates the value of performing this action in the state, and outputs , used to guide Actor network optimization strategies and improve action value; S2. Physical Constraint Embedding In order to make the strategy generated by the algorithm meet the actual operation requirements of the power system, the adaptive hierarchical Lagrange multiplier method is used to incorporate physical constraints into the reward function R; the formula is: In formula (1), Represents the original reward of the task, used to measure the system's performance in terms of stability and economy; is a time-varying Lagrange multiplier, which is dynamically updated by the following formula: in To constrain the learning rate of i, control the multiplier update step size, is the gradient contribution of constraint i to the total loss; is the constraint priority weight; is the penalty function for violating the i-th constraint. When the system operation state or action violates the corresponding physical constraint, Producing non-zero values ​​to reduce the reward function R, thereby prompting the algorithm to avoid generating strategies that violate the constraints; As the integral penalty coefficient, the historical cumulative effect of constraint violations is introduced; The historical cumulative violation of constraint i to avoid short-term oscillation; S3, key constraint implementation 1) Voltage amplitude constraint In the power system, the node voltage amplitude needs to be maintained in a reasonable range; through the formula: In formula (2), the voltage amplitude constraint penalty term is calculated as: is the voltage amplitude of node i; is the reference voltage value, usually set to 1.0pu; The allowable voltage deviation range is generally 0.05pu; is the voltage fluctuation adjustment coefficient, which dynamically adjusts the penalty intensity according to the voltage change rate; 2) Phase angle difference constraint If the phase angle difference between nodes is too large, it will affect the stability of the system. Formula (3) is used to constrain the phase angle difference: In formula (3), are the voltage phase angles at nodes i and j respectively; is the maximum allowed phase angle difference; E is the system line set; 3) Thermal power unit ramp rate constraint There is a limit on the power change speed of thermal power units, and the ramp rate is constrained by formula (4): In formula (4), is the active power of thermal power unit i at time t; is the time interval; is the maximum ramp rate, which is 5% of the rated power per minute; is the total number of thermal power units; is the basic power of the thermal power unit; is the rated power of the thermal power unit; 4) PV abandonment rate constraint In order to improve the utilization rate of photovoltaic energy, the abandoned light rate is constrained by formula (5): In formula (5), is the amount of electricity that can be generated by the photovoltaic power station; is the actual photovoltaic power generation; The maximum allowable abandoned light rate is 5%; is the proportion of photovoltaics in total energy; S4. Transfer learning mechanism 1) Scene similarity calculation Calculate the current scene With historical scenes The similarity is: In formula (6), the eigenvector x contains the number of system nodes , Number of lines , new energy penetration rate , and the load fluctuation standard deviation Key parameters of are the time spans of the current scene and the historical scene respectively; is the preset maximum time span; is the time weight coefficient, which adjusts the influence of time factors on similarity; 2) Parameter migration After triggering transfer learning, the following formula is used for parameter migration: In formula (7), is the migration ratio (0.7); is the historical training parameter; are random initialization parameters; S5. Multi-timescale strategy generation 1) Second-level strategy For scenarios with rapidly changing frequency fluctuations in the system, we use an improved Gated Recurrent Unit (GRU) combined with an attention mechanism to generate a second-level strategy. The system state in the past n time steps is calculated by formula (8) and the previous hidden state As input, the network automatically focuses on key state information through the attention mechanism, captures the dynamic characteristics of the system, and outputs second-level control actions. , to achieve fast response regulation; formula (8) is: 2) Hourly Strategy For long-term scenarios involving load peak and valley scheduling, an hourly strategy is constructed using an architecture combining graph convolutional networks (GCNs) and spatiotemporal transformers. The formula is: In formula (9), G is the graph structure representing the power system, which includes the power system node set V and the power system line edge set E; is the system state sequence of the past m moments; S6, training optimization 1) Target network update In order to improve the stability of algorithm training, the target network mechanism is adopted; the target network parameters of the Critic network are updated through formula (10): ; In formula (10), It is a dynamic soft update coefficient, which is dynamically adjusted according to the training error and gradient changes; is the current critic network parameter; 2) Critic network loss calculation The Critic network optimizes its own parameters by calculating the loss function. The formula is: In formula (11), r is the immediate reward; is a discount factor used to measure the importance of future rewards; is the regularization coefficient; is the L2 norm of the Critic network parameters.

Citation Information

Cited By

  • Power grid data center integrated power supply cooperative control and guarantee method, equipment and medium

    CN121602446A

  • Cooperative-control distributed energy storage scheduling method, system and equipment

    CN121688967A

  • OPC accelerated convergence method, system and terminal based on adaptive learning rate and multi-scale optimization

    CN122063823A