Distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving

By introducing a cooperative strategy network and a safety monitor into distributed power source control, combined with dynamic context awareness technology, the problems of single objective and rigid control in existing technologies are solved, realizing adaptive cooperative control of distributed power sources and improving the transient stability and security of the power system.

CN121749403APending Publication Date: 2026-03-27STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing distributed power source control strategies in new power systems suffer from problems such as single objective, lack of dynamic perception, and rigid control modes. They are unable to effectively cope with the challenges of grid transient stability, resulting in weakened frequency disturbance immunity and a prominent contradiction between control performance and security.

Method used

A distributed power supply cooperative control method based on transient dynamic characteristics is adopted. By embedding a cooperative policy network, a safety monitor and a robust backup controller in the local controller, and utilizing multi-dimensional transient performance coupling evaluation functions and dynamic context awareness technology, millisecond-level adaptive cooperative control is achieved.

Benefits of technology

It significantly improves the transient stability margin of the new power system, realizes the transformation of distributed power source control from passive and local to active and coordinated, enhances the global optimization capability of frequency, voltage, power angle and operating economy, and ensures the high safety and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121749403A_ABST
    Figure CN121749403A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving, which adopts a multi-dimensional transient performance coupling evaluation function, unifies multiple targets including frequency, voltage, power angle and operation economy into a global optimization target, and fundamentally solves the defect of single target in the prior art. A collaborative strategy network based on dynamic context awareness is adopted, and a time sequence processing technology based on an attention mechanism is utilized, so that an intelligent agent can extract key transient dynamic characteristics from a local observation sequence, and the limitation of no memory and non-self-adaption in the prior art is broken through; according to the method, a cooperative control flow of a cooperative control strategy is provided, a high-performance learning type strategy is combined with a high-reliability safety monitor, the problem that safety and performance are difficult to be compatible is solved, fundamental transformation of distributed power supply control from passive, local and static rule-based modes is achieved, and the transient stability margin of a novel power system is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power system operation and control, in particular to a distributed power source cooperative control method based on transient dynamic characteristics adaptive driving. BACKGROUND

[0002] The fundamental restructuring of power supply structure brings double core challenges to the safe and stable operation of power system: firstly, the physical support characteristics of the system change profoundly. Power electronic converters themselves do not have the inherent rotational inertia and physical damping characteristics of traditional synchronous generators, and can only provide limited virtual support through control strategies. With the continuous increase of its replacement ratio, the equivalent rotational inertia of the power grid is significantly reduced, which leads to a significant weakening of the frequency disturbance resistance of the system when it is subjected to large disturbances such as short-circuit of transmission lines, and the specific performance is that the frequency change rate increases sharply and the lowest frequency point is too low, which is easy to trigger the cascading protection action and cause the system frequency collapse. Secondly, there is a significant difference in control response mechanism. Power electronic converters have millisecond-level fast control response capability, which is essentially different from the inertia response and excitation response of synchronous generators on the scale of seconds. This characteristic can be a technical advantage for transient support, but the existing control strategy fails to fully exploit its potential, resulting in a great limitation of the overall support of distributed power sources in the system transient process.

[0003] In recent years, the data-driven intelligent technology path centered on reinforcement learning has begun to become the research focus for solving such complex dynamic regulation and control problems. At present, the commonly used is a distributed power source fault ride-through control strategy based on instantaneous measurement and static rules. The strategy is a passive and localized control scheme, and its core logic is: through local sensors to monitor key electrical quantities, once the monitored value instantaneously reaches the preset voltage drop static threshold, the controller will immediately trigger the rigid mode switching from the normal operation mode to the fault ride-through mode. In the fault ride-through mode, the controller strictly follows the pre-off-line setting and fixed control logic, that is, according to the current voltage drop degree, a fixed preset function is used to inject reactive current, and the active output is limited accordingly. The design goal of this strategy is to ensure that the power supply does not be disconnected from the grid under large disturbances, and to provide a certain and single reactive power support for the local node.

[0004] However, although the existing technology provides basic grid-connected safety protection, its control idea is limited and cannot meet the transient stability challenges of new power systems. The specific defects are as follows:

[0005] (1) Single goal, lack of coordination: the objective function of the existing fault ride-through strategy is single and local, that is, to maximize local voltage support, which ignores the multi-dimensional key performance in frequency stability, power angle damping and control economy in the system transient process. This local optimization often leads to global suboptimization or even worsens the system state.

[0006] (2) No memory, non-adaptive strategy: the existing strategy lacks the ability to adapt to the dynamic timing characteristics of the system, and the decision-making completely depends on the instantaneous measurement value at the current moment, which cannot perceive and utilize the dynamic timing characteristics of the transient process, resulting in blind and lagging response, and cannot realize proactive or adaptive active support. In addition, its fixed parameters are based on offline tuning, which cannot adapt to the real-time changes of the operation conditions of the power grid, resulting in lack of robustness.

[0007] (3) There is a contradiction between control performance and safety: the existing technical path faces the dilemma of performance and safety. The traditional fixed rule control has high reliability, but the logic is fixed, and the performance has obvious upper limit, which is a suboptimal solution at the cost of stability margin. While the purely data-driven strategy has high performance potential, but its black box characteristics and unpredictability make it unable to meet the high reliability standard of power system. Therefore, the existing technology has obvious fault between the verifiable suboptimal safe strategy and the unreliable high-performance strategy, and lacks means to integrate the advantages of both. SUMMARY

[0008] In order to solve the technical problems existing in the prior art, the application proposes a distributed power source cooperative control method based on transient dynamic characteristic adaptive driving, which combines the cooperative adaptive ability of reinforcement learning with the safety constraints of power system, and provides new technical support for new power system.

[0009] The application is implemented by the following technical solutions:

[0010] A distributed power source cooperative control method based on transient dynamic characteristic adaptive driving, comprising:

[0011] A plurality of distributed power source agents are used to realize millisecond-level adaptive cooperative control based on local timing observation;

[0012] The agent is formed by embedding a cooperative strategy network, a safety monitor and a robust backup controller in the local controller;

[0013] The cooperative strategy network is a cooperative strategy network based on dynamic context perception which is pre-constructed and trained to converge, which takes the global transient performance evaluation function shared by all agents as the optimization target of the transient process, and provides the optimal control strategy; the global transient performance evaluation function decouples the system transient stability into quantifiable multi-dimensional transient performance indicators;

[0014] The robust backup controller is a dynamic reactive power support controller based on local voltage measurement, which is used to output backup control strategy;

[0015] The safety monitor is a lightweight monitoring module which pre-defines a safe operating area of the local controller and performs hybrid control arbitration according to the safe operating area to output a final control strategy.

[0016] The control strategy comprises active power regulation instructions and reactive power regulation instructions.

[0017] In some embodiments, the global transient performance evaluation function is defined as a negative inner product of a multi-dimensional transient penalty vector and a priority weight vector.

[0018] Each component of the multi-dimensional transient penalty vector is a non-negative penalty term calculated based on a global state of the system or a joint action, and specifically comprises a frequency structure maintaining term, a voltage resilience maintaining term, an oscillation suppression term and a control economy constraint term, and the priority weight vector comprises weights corresponding to each component of the multi-dimensional transient penalty vector.

[0019] In some embodiments, the frequency structure maintaining term is used to quantify a frequency response characteristic of the system, and is defined as a weighted sum of squares of a frequency deviation of a key bus and a frequency change rate of an inertia center of the system.

[0020] The voltage resilience maintaining term is used to quantify a voltage out-of-limit degree of a key bus of the system, and is defined by using a quadratic penalty function.

[0021] The oscillation suppression term is used to quantify a power angle separation tendency between generator units, and is achieved by penalizing a relative angular velocity difference between key units.

[0022] The control economy constraint term is used to quantify a total cost of all agent control actions, and is defined as a weighted L2 norm sum of squares.

[0023] In some embodiments, the construction and training process of the cooperative strategy network comprises:

[0024] A cooperative strategy network based on dynamic context perception is constructed as a strategy function of the agent, the cooperative strategy network receives a local observation sequence within a time window as an input, and the local observation sequence is composed of a multi-dimensional local observation vector of a time window before a current time;

[0025] The constructed cooperative strategy network is trained, and in the training stage, each agent is equipped with a centralized value network, the centralized value network inputs a global state and a joint action to evaluate an expected return of a joint strategy, and the global state contains local observations of all agents and additional global information.

[0026] The centralized value network is iteratively updated by minimizing a time difference error loss.

[0027] The cooperative strategy network is iteratively updated by maximizing the expected return until the parameters of the cooperative strategy network of all agents converge.

[0028] In some embodiments, the cooperative strategy network comprises:

[0029] a transient dynamic feature encoder for extracting key dynamic indicators of transient processes from the observation sequence to generate a dynamic context vector;

[0030] and a decision-making cooperative generator that employs a multi-layer perception to combine the dynamic context vector output by the transient dynamic feature encoder with the most recent observation to generate a final control strategy.

[0031] In some embodiments, the transient dynamic feature encoder is implemented using a multi-head attention mechanism, and the implementation process is as follows:

[0032] The input observation sequence is first mapped to query, key, and value spaces to obtain query vectors, key vectors, and value vectors;

[0033] The weight of the value vector is obtained by calculating the scaled dot product of the query vector and the key vector;

[0034] The weighted sum of the value vector and its weight is calculated to obtain the dynamic context vector.

[0035] In some embodiments, the robust backup controller complies with grid fault ride-through specifications, and the backup control strategy output by the robust backup controller is determined according to the following rules:

[0036] The active power regulation instruction in the backup control strategy is set to zero or limited according to the specification during voltage drop;

[0037] The reactive power regulation instruction in the backup control strategy is calculated according to the deviation of the local terminal voltage from the nominal voltage through a pre-set piecewise linear gain;

[0038] The piecewise linear gain is pre-tuned to ensure that the output of the robust backup controller is still within the maximum rated value of the converter under the maximum voltage deviation.

[0039] In some embodiments, the millisecond-level adaptive cooperative control based on local time-series observations comprises:

[0040] Transient disturbance perception is performed;

[0041] When it is perceived that the system is subjected to a transient disturbance, a millisecond-level response is performed, and the local measurement unit of each distributed power source collects the changes in local electrical quantities caused by the transient disturbance, thereby constructing an observation vector at the current time and updating the observation sequence;

[0042] performing millisecond-level hybrid control arbitration to output a final control strategy;

[0043] performing millisecond-level physical layer execution, superimposing the final control strategy and the original steady-state power given value of the converter to form a total instantaneous power reference value, and generating a final control pulse signal based on the instantaneous power reference value to control the power electronic switch action.

[0044] In some embodiments, the performing millisecond-level hybrid control arbitration to output a final control strategy comprises:

[0045] inputting the updated observation sequence into the cooperative strategy network to perform a forward propagation to calculate an optimal control strategy;

[0046] inputting the observation vector at the current time and the optimal control strategy into the safety monitor to perform two checks: a state check to check whether the observation vector has entered or is about to enter an unsafe area; and an action check to check whether the optimal control strategy exceeds the preset physical executor constraint;

[0047] if the safety monitor check passes, adopting the optimal control strategy as the final control strategy;

[0048] if the safety monitor check fails, rejecting the optimal control strategy and starting the output of the robust backup controller as the final control strategy.

[0049] In a second aspect, the application provides a distributed power supply cooperative control system based on transient dynamic characteristics adaptive driving, comprising:

[0050] a plurality of distributed power supply intelligent agents for realizing millisecond-level adaptive cooperative control based on local time sequence observation;

[0051] wherein the intelligent agents are formed by embedding a cooperative strategy network, a safety monitor, and a robust backup controller in a local controller;

[0052] the cooperative strategy network is a cooperative strategy network based on dynamic context awareness that is pre-constructed and trained to converge, which takes a global transient performance evaluation function shared by all intelligent agents as an optimization target of the transient process and provides an optimal control strategy; the global transient performance evaluation function decouples the system transient stability into quantifiable multi-dimensional transient performance indicators;

[0053] the robust backup controller is a dynamic reactive power support controller based on local voltage measurement, which is used to output a backup control strategy;

[0054] The safety monitor is a lightweight monitoring module which pre-defines a safety operation area of the local controller and performs hybrid control arbitration according to the safety operation area, and outputs a final control strategy.

[0055] The control strategy includes active power regulation instructions and reactive power regulation instructions.

[0056] The application provides a distributed power source cooperative control method based on transient dynamic characteristics adaptive driving, which adopts a multi-dimensional transient performance coupling evaluation function, unifies multiple targets including frequency, voltage, power angle and operation economy as a global optimization target, and fundamentally solves the defect of single target in the prior art; a cooperative strategy network based on dynamic context perception is also adopted, a timing processing technology based on an attention mechanism is used, so that an intelligent agent can extract key transient dynamic characteristics from a local observation sequence, and the limitation of "no memory and non-adaptive" in the prior art is broken; and a cooperative control process of the cooperative control strategy is provided, a high-performance learning strategy is combined with a high-reliability safety monitor, the problem of difficulty in compatibility of safety and performance is solved, a fundamental change of distributed power source control from "passive, local and based on static rules" is realized, and the transient stability margin of a new power system is significantly improved.

[0057] Correspondingly, a distributed power source cooperative control system based on transient dynamic characteristics adaptive driving also has the same technical effects. BRIEF DESCRIPTION OF DRAWINGS

[0058] The accompanying drawings, which are included to provide a further understanding of the embodiments of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the principles of the application. In the drawings:

[0059] Figure 1 A cooperative control method principle schematic diagram provided by the embodiments of the application;

[0060] Figure 2 A cooperative control system architecture schematic diagram provided by the embodiments of the application. DETAILED DESCRIPTION

[0061] Hereinafter, the term "include" or "may include" used in various embodiments of the present application indicates the existence of the invented function, operation or element, and does not limit addition of one or more functions, operations or elements. In addition, as used in various embodiments of the present application, the terms "include", "have" and their synonyms merely indicate the presence of a specific feature, number, step, operation, element, component or combination of the foregoing, and should not be understood as first excluding the presence or addition of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing.

[0062] In various embodiments of the present application, the expression "or" or "at least one of A or / and B" includes any and all combinations of the listed terms. For example, the expression "A or B" or "at least one of A or / and B" can include A, can include B, or can include both A and B.

[0063] The expressions used in various embodiments of the present application, such as "first", "second", etc., can modify various constituent elements in various embodiments, but can not limit the corresponding constituent elements. For example, the above expressions do not limit the order and / or importance of the elements. The above expressions are only for the purpose of distinguishing one element from other elements. For example, the first user device and the second user device indicate different user devices, although both are user devices. For example, without departing from the scope of various embodiments of the present application, a first element can be referred to as a second element, and likewise, a second element can be referred to as a first element.

[0064] It should be noted that if a description connects one constituent element to another constituent element, the first constituent element can be directly connected to the second constituent element, and a third constituent element can be "connected" between the first constituent element and the second constituent element. Conversely, when one constituent element is "directly connected" to another constituent element, it can be understood that there is no third constituent element between the first constituent element and the second constituent element.

[0065] The terms used in various embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit various embodiments of the present application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Unless otherwise defined, all terms used herein, including technical terms and scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which various embodiments of the present application belong. The terms such as those defined in a generally used dictionary will be interpreted as having the same meaning as the context in the relevant technical field and will not be interpreted as having an idealized or overly formal meaning, unless clearly defined in various embodiments of the present application.

[0066] In order to make the purposes, technical solutions and advantages of the present application clearer, further detailed description of the present application is made below in combination with embodiments and drawings, the illustrative embodiments of the present application and the description thereof are only for the purpose of explaining the present application and do not limit the present application.

[0067] The existing distributed power supply control technology has problems such as single target, lack of dynamic perception, and rigid control mode. In view of the above problems, the embodiment of the present application proposes a distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving, as shown in Figure 1As shown, the method constructs a complete technical closed loop of "multi-dimensional target definition-dynamic strategy learning-mixed security execution", which designs a multi-dimensional transient performance coupling evaluation function, unifies the multi-element targets including frequency, voltage, power angle and operation economy as a global optimization target, and fundamentally solves the defect of single target in the prior art; at the same time, a collaborative strategy network based on dynamic context perception is designed, a timing processing module based on an attention mechanism is used, so that the intelligent agent can extract key transient dynamic characteristics from the local observation sequence, and the limitation of "no memory, non-adaptive" in the prior art is broken through; and a mixed execution and security deployment architecture of the collaborative control strategy is also proposed, which combines the high-performance learning strategy with the high-reliability security monitor, and solves the problem that safety and performance are difficult to be considered. The collaborative control method proposed in the embodiment of the application realizes a fundamental change of distributed power supply control from "passive, local, based on static rules" to "active, collaborative, based on dynamic characteristics", and significantly improves the transient stability margin of the new power system.

[0068] The collaborative control method proposed in the embodiment of the application includes the following steps:

[0069] A plurality of distributed power supply intelligent agents are used to realize millisecond-level adaptive collaborative control based on local timing observation; wherein the intelligent agent is formed by embedding a collaborative strategy network, a security monitor and a robust backup controller in a local controller; the collaborative strategy network is a collaborative strategy network based on dynamic context perception which is pre-constructed and trained to converge, which takes a global transient performance evaluation function shared by all intelligent agents as an optimization target of the transient process, and provides an optimal control strategy, the global transient performance evaluation function decouples the system transient stability into quantifiable multi-dimensional transient performance indexes; the robust backup controller is a dynamic reactive power support controller based on local voltage measurement, which is used to output a backup control strategy; and the security monitor is a lightweight monitoring module, which pre-defines a safe operation area of the local controller and performs mixed control arbitration accordingly to output a final control strategy.

[0070] Further, in order to guide the collaborative behavior of the system The embodiment of the application constructs a global transient performance evaluation function shared by all intelligent agents (the decision-making subject is a local controller) , which aims to decouple the complex multi-dimensional target of system transient stability into quantifiable multi-dimensional transient performance indexes, specifically:

[0071] The embodiment of the application defines the global transient performance evaluation function as the negative inner product of a multi-dimensional transient penalty vector and a priority weight vector :

[0072]

[0073] wherein, is a positive weight vector that can be offline adjusted to indicate the relative priority of different transient stability targets; is a transient penalty vector that characterizes the instantaneous stability state of the system at time , which is defined as:

[0074]

[0075] Each component of the transient penalty vector is a non-negative penalty term calculated based on the global state or joint action , which specifically includes:

[0076] a frequency structure preserving term , which is used to quantify the frequency response characteristics of the system, and is defined as the weighted sum of squared deviations of the key bus frequency and the rate of change of frequency (ROCOF) of the center of inertia (COI) of the system:

[0077]

[0078] wherein, is the nominal frequency, is the frequency of the key bus at time , and is the frequency of the COI of the system at time , and is the ROCOF penalty coefficient, is the set of system key buses; this term guides the agent to provide virtual inertia support, and the corresponding weight is .

[0079] a voltage resilience maintaining term , which is used to quantify the voltage excursion of the key buses of the system, and is defined using a quadratic penalty function:

[0080]

[0081] wherein, is the preset voltage safe operating interval, is the voltage of the key bus at time , and this term guides the agent to provide reactive power support to maintain voltage resilience, and the corresponding weight is .

[0082] a power angle oscillation suppression term , which is used to quantify the power angle separation trend between generator units, and is defined by penalizing the key units for The relative angular velocity difference (i.e., the relative work angle oscillation rate) between them is realized as follows:

[0083]

[0084] in, This is a set of key generator pairs used for monitoring power angle stability. , Generators The rotor angular velocity, this guides the agent to inject active damping through active modulation, and its corresponding weight is... .

[0085] Controlling economic constraints This item is used to quantify the control actions of all intelligent agents. The total cost is defined as the weighted sum of squares of the L2 norm:

[0086]

[0087] in, , The economic weighting coefficients correspond to the active and reactive power control actions, respectively. , These are active power control actions and active power control actions, respectively. This item is used to constrain the smoothness and economy of control actions, avoiding excessive wear on the actuators. Its corresponding weight is... .

[0088] Furthermore, in this embodiment, the construction and training process of the dynamic context-aware cooperative policy network is as follows:

[0089] First, a cooperative policy network based on dynamic context awareness is constructed. Specifically, to achieve adaptive driving of transient dynamic features, this application embodiment constructs a cooperative policy network based on dynamic context awareness. As an intelligent agent The policy function of the network Receive a time window Local observation sequence within As input, where for Moment Maintain local electrical quantities, such as The network structure contains two core modules: a transient dynamic feature encoder and a decision co-generator.

[0090] Among them, the transient dynamic feature encoder is used to extract from the observation sequence To extract key dynamic features of the transient process, this application embodiment employs a multi-head attention mechanism to achieve this function. Input sequence Firstly, the query ( ), key ( ), and value ( ) are mapped into three spaces:

[0091]

[0092] where is a learnable parameter matrix, are the query vector, key vector, and value vector, respectively. The attention mechanism obtains the weight of by computing the scaled dot-product of :

[0093]

[0094] where is the dimension of the key vector . The output of this encoder, i.e., the dynamic context vector, is the weighted sum of the value vector :

[0095]

[0096] This output vector dynamically aggregates the most crucial transient information for the current decision-making within the past time steps, forming an adaptive representation of the transient dynamic characteristics.

[0097] The collaborative decision generator is a multi-layer perceptron (MLP) that combines the adaptively encoded dynamic context vector with the latest observation to generate the final control action . To ensure the instant responsiveness of the decision, the two are concatenated and input into the MLP:

[0098]

[0099]

[0100] where is the set of all learnable parameters in the generator and encoder, and the output is defined as the active and reactive power regulation instructions that the agent should execute at time .

[0101] Then, the constructed dynamic context-aware collaborative strategy network is trained offline. In the offline training phase, each agent is equipped with a centralized value network​ with parameters . The network input is the global state and joint action to evaluate the expected return of the joint policy. Where:

[0102] The global state contains the local observations of all agents and additional global information i.e. , contains non-local observable information such as the states of traditional synchronous units, key line power flow, or global frequency, to alleviate the non-stationarity of the multi-agent environment.

[0103] The value network is updated by minimizing the temporal difference error loss . To support sequence input, the experience replay buffer stores trajectory segments containing historical observations is the sequential state information up to time , and is the sequential state information up to time , and is the reward, which is specific to a certain time .

[0104]

[0105] where is the target value, is the discount factor, and are the instantaneous global state and joint action at time extracted from :

[0106]

[0107] where and are the target network parameters, is the global state at time , and is the local observation sequence at time .

[0108] The cooperative policy network is then updated by maximizing the expected return , whose policy gradient is calculated as follows:

[0109]

[0110] in, All from the experience replay pool Obtained from [the source].

[0111] The training process is conducted in a high-fidelity simulation environment covering various operating conditions and fault scenarios, until all policy networks are trained. The parameters converge.

[0112] Furthermore, in this embodiment of the application, to ensure that the high performance of millisecond-level response is consistent with the high security required by critical infrastructure, this embodiment of the application adopts a hybrid execution architecture and security deployment with a collaborative control strategy, the specific implementation of which is as follows:

[0113] The three core components—cooperative policy network, security monitor, and robust backup controller—are solidified and embedded. The firmware of the local controller of a distributed power source enables the local controller to simultaneously perform neural network inference, safety rule verification, and control policy arbitration.

[0114] Among them, the collaborative policy network is the dynamic context-aware policy network that has been trained and converged. The parameters of this network are extracted and solidified, and it serves as the main controller of the system, responsible for providing optimal performance.

[0115] The robust backup controller is a dynamic reactive power support controller based on local voltage measurements. Its core logic conforms to the grid fault ride-through specification, and its output... Determined according to the following rules:

[0116] Active power regulation During voltage dips, active power output is either set to zero or limited according to specifications (such as priority reactive power injection). .

[0117] Reactive power regulation According to the local terminal voltage With nominal voltage deviation Through a preset piecewise linear gain To calculate the reactive power command, i.e. .

[0118] The parameters of the controller It is pre-tuned to ensure that its output remains within the converter's maximum rating even under maximum voltage deviation, thus providing a deterministic and safe fallback voltage support operation.

[0119] The safety monitor is a lightweight monitoring module that pre-defines a safe operating region for the local agent, whose boundary is consistent with the penalty term in the multi-dimensional transient performance coupling evaluation function and the physical actuator constraints (physical constraints of the converter).

[0120] The local inference-based transient closed-loop control is performed, and the specific control flow is as follows:

[0121] When the system encounters a transient disturbance, the online hybrid execution flow is activated according to the following millisecond-level time sequence:

[0122] [ ]Disturbance awareness: the system suffers a transient disturbance, and the state quantity starts to deviate from the nominal operating point;

[0123] [ ]Local measurement: the local measurement unit of each distributed power source collects the dramatic change of the local electrical quantity caused by the disturbance, and the local controller constructs the observation vector at the moment and updates the observation sequence ;

[0124] [ ]Hybrid control arbitration: the observation data is sent into three embedded modules in parallel: strategy network inference: is input to , and a forward propagation is performed to calculate the optimal recommended action :

[0125]

[0126] Safety monitor check: the safety monitor performs two checks based on and :

[0127] State check: check whether has entered or is about to enter an unsafe region;

[0128] Action check: check whether has exceeded the preset physical actuator constraints;

[0129] Final action arbitration: if the safety monitor check passes, the decision of the neural network is adopted, that is:

[0130] If the safety monitor check fails, the decision of the neural network is rejected , and the output of the robust backup controller is immediately started as the bottom action, that is:

[0131]

[0132] This arbitration process is completely based on local information, without communication, ensuring the safety and instantaneity of the decision. Physical layer execution: by the arbitration module in The control action of the final decision at the moment is superimposed on the original steady-state power given value of the converter. Among them, and are the steady-state active and reactive power given values of the converter before the transient occurs. Superimposed to form the total transient power reference value and :

[0133]

[0134]

[0135] The inner loop current control and pulse width modulation layer of the converter track and , and through the action of power electronic switches, the transient support power after safety verification is quickly injected or absorbed to the grid. It should be noted that the millisecond-level timing given in the above process is only an exemplary description, and does not limit the specific timing.

[0136] Through this embedded hybrid execution architecture, the agents in the distributed power source realize millisecond-level adaptive cooperative control based on local timing observation under the premise of ensuring the safety boundary.

[0137] Based on the same technical concept described above, the embodiment of the application also proposes a distributed power source cooperative control system based on transient dynamic characteristic adaptive driving. The cooperative control system proposed by the embodiment of the application comprises:

[0138] A plurality of distributed power source agents for realizing millisecond-level adaptive cooperative control based on local timing observation; wherein the agent is formed by embedding a cooperative strategy network, a safety monitor and a robust backup controller in a local controller; the cooperative strategy network is a cooperative strategy network based on dynamic context perception that is pre-constructed and trained to converge, which takes a global transient performance evaluation function shared by all agents as the optimization target of the transient process, and provides an optimal control strategy, and the global transient performance evaluation function decouples the system transient stability into quantifiable multi-dimensional transient performance indicators; the robust backup controller is a dynamic reactive power support controller based on local voltage measurement, which is used to output a backup control strategy; the safety monitor is a lightweight monitoring module, which pre-defines the safe operation area of the local controller and performs hybrid control arbitration accordingly to output the final control strategy. Figure 2A single-agent architecture is shown.

[0139] It should be noted that the construction and training process of the above-mentioned cooperative strategy network, the construction method of the global transient performance evaluation function, and the millisecond-level adaptive cooperative control process are as described in the above-mentioned cooperative control method, and will not be repeated here.

[0140] In addition, the cooperative control system proposed in the embodiments of the present application further comprises:

[0141] A local measurement unit of the plurality of distributed power sources is configured to collect local electrical quantities and send the local electrical quantities into the local controller to construct an observation vector.

[0142] A power superposition unit of the plurality of distributed power sources is configured to superimpose the final decision output by the local controller and the original steady-state power given value of the converter to form a transient power reference value.

[0143] In addition, an inner loop current control and pulse width modulation unit of the converter is configured to track the transient power reference value, generate a control instruction, and output the control instruction to the power electronic switch for corresponding control.

[0144] The local measurement unit, the power superposition unit, and the inner loop current control and pulse width modulation unit all belong to conventional control technology of a power system, and are not the core of the present application. Therefore, the specific implementation manner will not be repeated here.

[0145] The embodiments of the present application compare the above-mentioned cooperative control method (referred to as TDA-DCC) with two existing schemes, namely, a reference scheme (static rule FRT) and an ablation experiment (referred to as RL-DCC). The static rule FRT adopts rigid rule switching logic based on instantaneous measurement and a preset static threshold, and only takes local voltage support as a single target, without cooperation or dynamic adaptability. The RL-DCC is a reinforcement learning (RL) scheme for ablation experiment comparison. The RL-DCC shares the same multi-dimensional transient performance coupling evaluation function as the TDA-DCC proposed in the embodiments of the present application , but the strategy network is a simple multi-layer perceptron, and the input is only the instantaneous measurement at the current time , which is a "memoryless" network, used to verify the necessity and superiority of the transient dynamic feature encoder based on the attention mechanism (i.e., processing time series ) proposed in the embodiments of the present application.

[0146] By setting a three-phase short-circuit fault of a key transmission line as a reference fault, and comparing the comprehensive transient performance and stability margin of the system under a plurality of working conditions including different DER penetration rates and different fault locations, the detailed comparison is shown in Tables 1-3.

[0147] Table 1 Comparison of instantaneous performance indicators

[0148]

[0149] Table 2 Comparison of Cumulative Performance Indicators

[0150]

[0151] Table 3 Comparison of robustness indices (CCT: critical resection time)

[0152]

[0153] The simulation results above demonstrate that the TDA-DCC method proposed in this application effectively overcomes the three core defects of the prior art:

[0154] (1) It achieves synergistic optimization of multi-dimensional transient indicators, overcoming the defects of "single objective and lack of coordination". As shown in Tables 1 and 2, the static rule FRT performance is comprehensively lagging behind: its instantaneous indicators such as the lowest frequency point and the maximum power angle offset are the worst, and the accumulated frequency, voltage, and power angle penalties are all at the highest level. This proves that its local and single objective leads to global suboptimality. In contrast, both reinforcement learning schemes (RL-DCC and TDA-DCC) have achieved significant improvements in all indicators. This is due to the multi-dimensional transient performance coupled evaluation function, which guides the agent to learn the globally optimal synergistic strategy that takes into account frequency, voltage, power angle and economy, fundamentally solving the problem of lack of coordination in the FRT strategy.

[0155] (2) Adaptive decision-making based on dynamic features is achieved, overcoming the limitations of "memorylessness and non-adaptability". The core innovation of this invention lies in utilizing dynamic context awareness. Although the RL-DCC scheme is superior to FRT, its "memorylessness" characteristic makes its performance lag behind TDA-DCC across the board. As shown in the table, the embodiment of this application achieves the best results in all instantaneous performance indicators, all cumulative performance indicators, and all robust scenarios. This is because RL-DCC, like FRT, relies on instantaneous measurements, resulting in decision-making lag. However, the TDA-DCC embodiment of this application, through a transient dynamic feature encoder, can extract key dynamic features from time-series observations, achieving a leap from passively responding to instantaneous values ​​to actively predicting dynamic trends, demonstrating stronger adaptability to operating conditions.

[0156] (3) The dilemma between control performance and safety is solved, and the deployment bottleneck of black-box strategy is overcome. Scenarios 1-3 in Tables 1-3 prove the high performance of the TDA-DCC strategy proposed in the embodiments of the present application. On this basis, the sensor failure test in scenario 4 in Table 3 verifies the hybrid execution and safety deployment architecture proposed in the embodiments of the present application. In this scenario, the RL-DCC scheme relying purely on sensor data driving fails to make decisions due to false input, and the stability deteriorates seriously, with the CCT falling to 125 ms, exposing the safety hazard of the black-box strategy. In contrast, the safety monitor module in the embodiments of the present application successfully identifies the anomaly and switches to the robust backup controller to perform bottom support, keeping the CCT at 185 ms, close to the FRT performance of the static rule that does not rely on sensors at all, effectively preventing the risk of system instability caused by purely data-driven reinforcement learning methods in extreme cases. It is proved that the embodiments of the present application combine the high performance of data-driven reinforcement learning with the high-reliability safety bottom mechanism, and overcome the technical fault line that performance and safety are difficult to balance.

[0157] In summary, through the complete technical closed loop of “multi-dimensional target definition-dynamic strategy learning-hybrid safety execution”, the embodiments of the present application successfully solve the core defects of the prior art, such as non-collaboration, non-adaptability, and split safety performance, and realize the fast, collaborative, adaptive, and high-safety optimization control of distributed power sources in transient processes, thereby improving the transient stability and operation robustness of power systems with high proportion of DERs to a new level.

[0158] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0159] The embodiments of the present application are described with reference to flowcharts and / or block diagrams according to the methods, devices (systems), and computer program products of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices generate a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks. Figure 1 The device that implements the functions specified in one flow or multiple flows and / or blocks.

[0160] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0162] The specific implementation described above is further explained in detail for the purpose of purposes, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific implementation of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A distributed power supply cooperative control method based on transient dynamic characteristics adaptive driving, characterized in that, include: By utilizing multiple distributed power intelligent agents, millisecond-level adaptive cooperative control based on local time-series observation is achieved; The intelligent agent is formed by embedding a cooperative policy network, a security monitor, and a robust backup controller in a local controller; The cooperative policy network is a pre-constructed and trained convergent network based on dynamic context awareness. It uses a global transient performance evaluation function shared by all agents as the optimization objective of the transient process and provides the optimal control policy. The global transient performance evaluation function decouples the transient stability of the system into quantifiable multidimensional transient performance indicators. The robust backup controller is a dynamic reactive power support controller based on local voltage measurement, used to output backup control strategies. The security monitor is a lightweight monitoring module that predefines the safe operating zone of the local controller and performs hybrid control arbitration accordingly, outputting the final control strategy. The control strategy includes active power regulation commands and reactive power regulation commands.

2. The distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving according to claim 1, characterized in that, The global transient performance evaluation function is defined as the negative inner product of the multi-dimensional transient penalty vector and the priority weight vector; Each component of the multi-dimensional transient penalty vector is a non-negative penalty term calculated based on the global state or joint actions of the system, specifically including: frequency structure preservation term, voltage resilience preservation term, power angle oscillation suppression term, and control economy constraint term. The priority weight vector includes the weights corresponding to each component of the multi-dimensional transient penalty vector.

3. The distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving according to claim 2, characterized in that, The frequency structure preservation term is used to quantify the system frequency response characteristics, and it is defined as the weighted sum of squares of the critical bus frequency deviation and the rate of change of the system inertial center frequency. The voltage resilience maintenance term is used to quantify the degree of voltage overrun on the system's critical bus and is defined using a quadratic penalty function; The power angle oscillation suppression term is used to quantify the power angle separation trend between generator sets, which is achieved by penalizing the relative angular velocity difference between key generator pairs. The control economy constraint term is used to quantify the total cost of all agent control actions and is defined as the weighted sum of squared L2 norms.

4. A distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving according to any one of claims 1-3, characterized in that, The construction and training process of the cooperative policy network includes: A dynamic context-aware cooperative policy network is constructed as the policy function of the agent. The cooperative policy network receives a local observation sequence within a time window as input. The local observation sequence is composed of a multi-dimensional local observation vector of a time window before the current time. The constructed cooperative policy network is trained. During the training phase, each agent is equipped with a centralized value network. The input of the centralized value network is the global state and joint action, which are used to evaluate the expected reward of the joint policy. The global state includes the local observations of all agents and additional global information. The centralized value network is iteratively updated by minimizing the loss of temporal difference error; The cooperative policy network is iteratively updated by maximizing the expected reward until the parameters of the cooperative policy network of all agents converge.

5. The distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving according to claim 4, characterized in that, The collaborative strategy network includes: Transient dynamic feature encoder is used to extract key dynamic indicators of transient processes from observation sequences and generate dynamic context vectors; In addition, a decision co-generation generator, which employs a multilayer perceptron, combines the dynamic context vector output by the transient dynamic feature encoder with the latest observations to generate the final control strategy.

6. The distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving according to claim 5, characterized in that, The transient dynamic feature encoder is implemented using a multi-head attention mechanism, and the specific implementation process is as follows: The input observation sequence is first mapped to three spaces: query, key, and value, resulting in query vector, key vector, and value vector; The weights of the value vectors are obtained by calculating the scaled dot product of the query vector and the key vector; The dynamic context vector is obtained by calculating the weighted sum of the value vector and its weights.

7. The distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving according to claim 4, characterized in that, The robust backup controller conforms to the power grid fault ride-through specification, and its output backup control strategy is determined according to the following rules: The backup control strategy includes an active power regulation command, which is either set to zero or restricted according to specifications during voltage dips. The reactive power adjustment command in the backup control strategy is calculated based on the deviation between the local terminal voltage and the nominal voltage using a preset piecewise linear gain. The piecewise linear gain is pre-tuned to ensure that the robust backup controller output remains within the converter's maximum ratings even under maximum voltage deviation.

8. The distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving according to claim 7, characterized in that, The millisecond-level adaptive cooperative control based on local time-series observation includes: Perform transient disturbance sensing; When a transient disturbance is detected in the system, a millisecond-level response is initiated. The local measurement units of each distributed power source collect the changes in local electrical quantities caused by the transient disturbance, construct the observation vector at the current moment, and update the observation sequence accordingly. Perform millisecond-level hybrid control arbitration and output the final control strategy; The millisecond-level physical layer execution is performed, and the final control strategy is superimposed on the original steady-state power setpoint of the converter to form a total instantaneous power reference value. Based on the instantaneous power reference value, the final control pulse signal is generated to control the operation of the power electronic switch.

9. A distributed power supply cooperative control method based on transient dynamic characteristic adaptive driving according to claim 8, characterized in that, The aforementioned millisecond-level hybrid control arbitration, outputting the final control strategy, includes: The updated observation sequence is input into the cooperative policy network, a forward propagation is performed, and the optimal control policy is calculated. The current observation vector and optimal control strategy are input into the safety monitor to perform two checks: a state check, which checks whether the observation vector is already in or about to enter an unsafe area; and an action check, which checks whether the optimal control strategy exceeds the preset physical actuator constraints. If the security monitor passes the verification, the optimal control strategy will be adopted as the final control strategy. If the security monitor fails the verification, the optimal control strategy is rejected, and the output of the robust backup controller is activated as the final control strategy.

10. A distributed power supply cooperative control system based on transient dynamic characteristic adaptive drive, characterized in that, include: Multiple distributed power agents are used to achieve millisecond-level adaptive cooperative control based on local time-series observations; The intelligent agent is formed by embedding a cooperative policy network, a security monitor, and a robust backup controller in a local controller; The cooperative policy network is a pre-constructed and trained convergent network based on dynamic context awareness. It uses a global transient performance evaluation function shared by all agents as the optimization objective of the transient process and provides the optimal control policy. The global transient performance evaluation function decouples the transient stability of the system into quantifiable multidimensional transient performance indicators. The robust backup controller is a dynamic reactive power support controller based on local voltage measurement, used to output backup control strategies. The security monitor is a lightweight monitoring module that predefines the safe operating zone of the local controller and performs hybrid control arbitration accordingly, outputting the final control strategy. The control strategy includes active power regulation commands and reactive power regulation commands.

Citation Information

Cited By

  • A power system inertia constant adaptive control method and system

    CN122159242A