Power distribution network flexible management method, system, program and storage medium

By combining lightweight neural networks and virtual synchronous machines in the distribution network, the inverter output voltage is dynamically adjusted, which solves the problem of insufficient adaptive capability of existing governance technologies under complex operating conditions. This enables precise and flexible governance of the distribution network, improving system stability and power quality for equipment.

CN122118709APending Publication Date: 2026-05-29SICHUAN ZHONGDIAN AOSTAR INFORMATION TECHNOLOGIES CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN ZHONGDIAN AOSTAR INFORMATION TECHNOLOGIES CO LTD
Filing Date
2026-04-22
Publication Date
2026-05-29

Smart Images

  • Figure CN122118709A_ABST
    Figure CN122118709A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of power distribution network management, and discloses a power distribution network flexible management method, system, program and storage medium. The method comprises the following steps: inputting a target power distribution network real-time feature vector into an updated lightweight neural network model to perform nonlinear mapping calculation, and outputting a virtual synchronous machine reference reference voltage at a current time; determining a virtual angular displacement at the current time based on a virtual angular velocity and displacement, an active deviation and an action parameter space vector of the virtual synchronous machine at the current time; determining a sinusoidal reference voltage according to the space vector, the initial bias angle of each phase, the reference reference voltage, the virtual angular displacement and the real-time voltage effective value of the low-voltage trapped phase; determining a modulation ratio control instruction according to the action, the sinusoidal reference voltage and the actual output electrical parameter of the inverter; and driving the inverter to convert the DC bus voltage into an AC compensation voltage and inject the AC compensation voltage into the low-voltage trapped phase. The application can enhance the adaptive adjustment capability of the power distribution network under complex scenario power fluctuations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power distribution network governance technology, specifically to a flexible governance method, system, program, and storage medium for power distribution networks. Background Technology

[0002] Modern power distribution networks are evolving from traditional unidirectional power receiving networks to complex networks with high interaction between sources, grids, and loads. In this process, the widespread integration of massive distributed energy resources and the frequent start-ups and shutdowns of various nonlinear, high-power alternating loads (such as charging facilities for new energy vehicles and large industrial equipment) have led to extremely complex power flow distribution at the end of the distribution network. This bidirectional, dynamic, and unpredictable tidal power fluctuation can easily trigger severe power quality problems at the end of the distribution network, such as serious three-phase load imbalances and localized low-voltage drops, posing significant challenges to the operational safety of end-user equipment and the stability of the entire distribution network.

[0003] To alleviate the aforementioned power supply difficulties, existing distribution network end-point management technologies typically rely on traditional static reactive power compensation devices or conventional control models based on preset fixed parameters for unidirectional voltage and current regulation. However, these conventional management methods are often static control strategies based on specific operating conditions or historical experience, and can only function within extremely limited settings.

[0004] One of the most prominent common problems with existing technologies is the lack of dynamic adaptive capabilities and continuous intelligent optimization capabilities in the face of complex and ever-changing operating conditions. In practical applications, the topology, environmental conditions, and source-load fluctuations at the end of the distribution network are changing rapidly in real time. Traditional devices that rely on fixed parameters, fixed dead zones, or finite static rules cannot perceive this environmental evolution in multiple dimensions, let alone learn and correct control deviations autonomously based on the real-time grid status without human intervention.

[0005] When the power grid experiences frequent fluctuations or encounters sudden disturbances, the response of existing power management equipment is severely lagging, easily leading to insufficient or excessive compensation. Due to a lack of deep environmental adaptability and self-evolution mechanisms, existing technologies not only struggle to accurately and flexibly address the rapidly changing low-voltage situations at the end of the distribution network, but may even trigger secondary oscillations in the system under extreme conditions, ultimately failing to effectively guarantee the long-term stability of the end-point power supply system and the absolute safety of equipment power consumption. Summary of the Invention

[0006] The purpose of this application is to provide a flexible management method, system, program, and storage medium for distribution networks, in order to solve the problem of poor adaptability of distribution network management under changing operating conditions in the prior art.

[0007] To achieve the above objectives, the first aspect of this application provides a method for flexible management of distribution networks, the method comprising: The real-time feature vector of the target distribution network is input into the updated lightweight neural network model for nonlinear mapping calculation, and the reference reference voltage of the virtual synchronizing machine at the current moment is output. Based on the virtual angular velocity and virtual angular displacement of the virtual synchronizer at the previous moment, the active power deviation, the preset reference angular velocity, and the pre-confirmed motion parameter space vector, the virtual angular displacement of the virtual synchronizer at the current moment is determined. Based on the action parameter space vector, the initial bias angle of each phase in the virtual synchronizing machine, the reference reference voltage, the virtual angular displacement, and the real-time effective voltage value of the low-voltage trapped phase of the target distribution network, the sinusoidal reference voltage of each phase is determined. Based on the action parameter space vector, sinusoidal reference voltage, and the inverter's actual output voltage and current at the current moment, determine the modulation ratio control command for each phase; The inverter is driven by the modulation ratio control command to convert the DC bus voltage into AC compensation voltage and inject the AC compensation voltage into the low-voltage trapped phase.

[0008] The second aspect of this application provides a flexible distribution network management system, including a sampling module, a rectifier module, a DC bus, a control module, and an inverter module. The sampling module's acquisition terminal is electrically connected to the target distribution network, and its signal output terminal is communicatively connected to the control module. The inverter module's DC input terminal is electrically connected to the DC bus, its control terminal is communicatively connected to the control module, and its AC output terminal is electrically connected to the low-voltage trapped phase of the target distribution network. The control module is used to input the real-time feature vector of the target distribution network acquired by the acquisition module into an updated lightweight neural network model for nonlinear mapping calculation, and output the reference reference voltage of the virtual synchronizer at the current moment. Based on the virtual synchronizer's virtual... The virtual angular displacement of the virtual synchronous machine at the current moment is determined by the angular velocity, virtual angular displacement, active power deviation, preset reference angular velocity, and pre-confirmed action parameter space vector. Based on the action parameter space vector, the initial bias angle of each phase within the virtual synchronous machine, the reference voltage, the virtual angular displacement, and the real-time effective voltage value of the low-voltage trapped phase in the target distribution network, the sinusoidal reference voltage of each phase is determined. Based on the action parameter space vector, the sinusoidal reference voltage, and the actual output voltage and actual output current of the inverter module at the current moment, the modulation ratio control command for each phase is determined. The modulation ratio control command drives the inverter module, controlling it to convert the DC bus voltage into an AC compensation voltage and inject the AC compensation voltage into the low-voltage trapped phase.

[0009] A third aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0010] A fourth aspect of this application provides a machine-readable storage medium storing instructions that cause a machine to perform the methods described above.

[0011] Through the above technical solution, the system first uses an updated lightweight neural network model to perform deep nonlinear mapping on the real-time feature vector of the target distribution network, overcoming the limitation of traditional governance methods that heavily rely on preset fixed standard voltages. This allows for the precise and dynamic derivation of the most suitable reference voltage for the current operating conditions, combining real-time grid fluctuations with the energy storage status of the equipment itself. Subsequently, by injecting the pre-confirmed action parameter space vector into the virtual synchronous machine dynamics model, the dynamically updated damping characteristics effectively suppress system frequency fluctuations and secondary oscillations caused by low voltage. Supported by virtual inertia, the system ensures the smoothness of the adjustment process, thus adaptively deriving an extremely stable real-time virtual angular displacement as the synchronization reference. Based on this, the system accurately captures the transient state between the ideal target and the actual drop state. The deviation is addressed by combining dynamically optimized compensation amplitude with real-time synchronous phase information to create an independent ideal sinusoidal reference voltage waveform for each AC phase. This enables the system to precisely manage and bias the low-voltage trapped phase. Next, by constructing a dual closed-loop feedback architecture for voltage and current, the ideal target is compared in real time with the actual output state of the inverter, significantly improving the system's response speed to sudden disturbances and waveform tracking accuracy. Finally, by driving the inverter through modulation ratio control commands to perform a flexible conversion and injection of digital decision-making into physical energy, the system directly and effectively resolves the three-phase load imbalance and local voltage drop problems at the end of the distribution network. While ensuring the power quality of terminal equipment, this significantly enhances the adaptive adjustment capability and long-term operational stability of the entire distribution network under complex power fluctuation scenarios.

[0012] Other features and advantages of the embodiments of this application will be described in detail in the following detailed description section. Attached Figure Description

[0013] The accompanying drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the following detailed description to explain the embodiments of this application, but do not constitute a limitation on the embodiments of this application. In the drawings: Figure 1 A flowchart illustrating a flexible management method for a power distribution network according to an embodiment of this application is shown schematically. Figure 2 A flowchart illustrating another flexible management method for power distribution networks according to an embodiment of this application is shown schematically; Figure 3 The diagram illustrates the architecture of a flexible power distribution network management system according to an embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustration and explanation of the embodiments of this application and are not intended to limit the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0015] It should be noted that if the embodiments of this application involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0016] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0017] The acquisition, transmission, storage, use, and processing of data in this application comply with relevant laws and regulations. Furthermore, it should be noted that certain software, components, models, and other existing industry solutions may be mentioned in the embodiments of this application. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0018] It should be noted that all data involved in this application (including but not limited to data used for analysis, data stored, data displayed, etc.) are information and data authorized by the client or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0019] Figure 1 A flowchart illustrating a flexible management method for a power distribution network according to an embodiment of this application is shown schematically. Figure 1As shown in the figure, this application provides a method for flexible management of a power distribution network, which may include the following steps.

[0020] Step 101: Input the real-time feature vector of the target distribution network into the updated lightweight neural network model for nonlinear mapping calculation, and output the reference voltage of the virtual synchronizing machine at the current moment.

[0021] In this embodiment, the target distribution network refers to a controlled power network system currently undergoing flexible power quality management and real-time monitoring. The real-time feature vector refers to a data set obtained through the underlying sensing module that reflects the transient operating status of the controlled power network system and related information such as the energy storage capacity of internal equipment in multiple dimensions. The lightweight neural network model refers to an artificial intelligence algorithm structure with low computational complexity, easily deployed and executed on limited computing power or edge devices, used for efficient data processing and feedforward prediction. The internal weights and bias parameters of the relevant network architecture of the lightweight neural network model here have been adaptively optimized and iterated based on recent actual power grid operating conditions or reinforcement learning algorithm training results. Nonlinear mapping calculation refers to using mathematical mechanisms such as activation functions within the aforementioned artificial intelligence algorithm structure to perform complex spatial transformations and feature fitting on the input multi-dimensional underlying operating data to accurately capture the nonlinear correlation between the complex alternating power grid state and the desired output target. A virtual synchronous machine refers to an inverter control algorithm or logical entity that simulates the rotor kinematics and electrical characteristics of a traditional synchronous generator, used to provide dynamic voltage and frequency support for the power network. The current moment represents the instantaneous discrete calculation cycle in which the system control algorithm is executing and responding. The reference voltage, distinct from the traditional pre-set fixed dead zone value, is the expected target ideal voltage parameter dynamically derived by the aforementioned intelligent structure based on real-time system operating conditions. It serves as the standard basis for the subsequent control system to generate compensation commands. Through this step, the system overcomes the limitations of traditional governance methods that heavily rely on preset fixed standard voltages. By fully considering the real-time fluctuations and deviations of the power grid and the energy storage status of the equipment itself through intelligent algorithms, it accurately, quickly, and adaptively calculates the most suitable expected target voltage under the current operating conditions. This lays an accurate computational foundation for subsequent precise and flexible compensation for low-voltage conditions in complex and variable power grids.

[0022] Step 102: Based on the virtual angular velocity and virtual angular displacement of the virtual synchronizer at the previous moment, the active power deviation, the preset reference angular velocity, and the pre-confirmed motion parameter space vector, determine the virtual angular displacement of the virtual synchronizer at the current moment.

[0023] In this embodiment, the previous moment refers to the data sampling or processing cycle immediately preceding the current real-time computing node in the discrete digital control system. Virtual angular velocity refers to the dynamic state variable used to simulate the rotation speed of a traditional rotating generator rotor, reflecting the transient frequency characteristics and trends of the system. Virtual angular displacement refers to the simulated current position or phase deflection angle of the generator rotor during rotation, serving as a phase reference benchmark characterizing the system's operating state and maintaining synchronization with the power grid. Active power deviation refers to the numerical difference between the system's expected output reference power target and the total power actually measured and calculated by the inverter system within a specific evaluation period, representing the current energy supply and demand imbalance state of the system. Preset reference angular velocity refers to the rated or nominal rotational speed benchmark pre-set within the system to align with the standard operating frequency of the power grid. Pre-confirmed action parameter space vector refers to the combined parameters, including multi-dimensional control coefficients and variables such as moment of inertia adjustment values ​​and damping coefficient adjustment values, determined or generated by the intelligent inference model during the pre-processing stage. By combining pre-optimized dynamic control variables with the real-time power imbalance state, the rotor dynamic evolution characteristics of the generator are accurately reconstructed and simulated. It can not only effectively suppress system frequency fluctuations and secondary oscillations caused by low voltage by utilizing dynamically updated damping characteristics, but also ensure the smoothness of the response process through virtual inertia, thereby adaptively deduce and iterate an extremely stable internal phase reference at the current moment, laying a core phase reference foundation for promoting phase-to-phase balance and subsequent flexible compensation in the distribution network.

[0024] Step 103: Determine the sinusoidal reference voltage for each phase based on the action parameter space vector, the initial bias angle of each phase in the virtual synchronizing machine, the reference reference voltage, the virtual angular displacement, and the real-time effective voltage value of the low-voltage trapped phase of the target distribution network.

[0025] In this embodiment, the initial bias angle of each phase refers to the inherent static basic electrical angle difference between each AC line in a multi-phase power system, used to establish an independent and staggered phase calculation reference for each phase within the algorithm. The real-time effective voltage value of the low-voltage trapped phase refers to the actual root mean square voltage value extracted instantaneously by the underlying hardware on the specific transmission line currently experiencing insufficient power supply or a significant voltage drop anomaly in the overall controlled power network. The sinusoidal reference voltage refers to the ideal AC waveform target sequence constructed by using the dynamically adjusted synthetic voltage amplitude and the real-time phase angle containing system synchronization characteristics. This instruction follows a sinusoidal alternation law and is specifically used to indicate the expected final target voltage state. By accurately capturing the transient deviation between the ideal target state and the actual drop state, and introducing dynamically optimized proportional and integral gain and other action parameters for adaptive adjustment, the compensation amplitude that closely matches the actual needs can be flexibly derived. Moreover, by combining this amplitude with real-time synchronization phase information, an independent ideal tracking waveform is tailored for each AC phase. This mechanism enables the compensation control algorithm to have multi-dimensional perception and self-correction capabilities, and can lock onto and bias towards supporting those electrical phases that are truly in a low-voltage trapped state, providing the most core and accurate waveform control benchmark for resolving the variable power fluctuations at the end and achieving three-phase balance.

[0026] Step 104: Determine the modulation ratio control command for each phase based on the action parameter space vector, sinusoidal reference voltage, and the actual output voltage and current of the inverter at the current moment.

[0027] In this embodiment, the inverter refers to the core hardware device that performs AC / DC power conversion, responsible for converting internal DC power into AC power to output compensation energy. The actual output voltage refers to the physical quantity of the real potential difference instantaneously generated by the core hardware device on the AC side and applied to the external electrical node. The actual output current refers to the numerical quantity of the real charge flow instantaneously flowing out of the core hardware device on the AC side. The modulation ratio control command for each phase refers to the underlying control signal or drive parameter used to directly drive the underlying power switching devices to periodically switch on and off according to a specific duty cycle. By constructing a dual closed-loop feedback mechanism of voltage and current, the ideal reference waveform target is compared with the actual output state of the device in real time, and adaptive calculations are performed using dynamically optimized adjustment gain parameters, thereby enabling the rapid and accurate generation of the underlying switching drive signal. This effectively improves the system's response speed and tracking accuracy to sudden low-voltage disturbances, ensuring that the compensation device can quickly self-correct according to real-time deviations, providing reliable command guarantees for the final output of a high-quality compensation waveform and reducing low-voltage deviations.

[0028] Step 105: Drive the inverter through modulation ratio control command, control the inverter to convert the DC bus voltage into AC compensation voltage, and inject the AC compensation voltage into the low voltage trapped phase.

[0029] In this embodiment, the DC bus refers to a common conductive loop or node architecture in an AC / DC power conversion system used for collecting, transmitting, and buffering DC power, playing a key role in energy transfer and voltage stabilization storage. AC compensation voltage refers to a dynamic AC potential difference physical quantity, synthesized and output by the power conversion hardware based on a pre-control strategy, specifically used to fill the difference between the actual electrical state and the ideal target state of the controlled network. Low-voltage trapped phase refers to a specific AC transmission line in a multi-phase power supply network that is currently suffering from negative impacts such as insufficient energy or severe load fluctuations, causing the overall power supply state to drop below the normal baseline. By executing the underlying drive signal generated by the pre-calculation, the final closed loop of conversion from digital information instructions to actual physical power output is completed. This process inverts the previously extracted and stored DC energy into an AC form that meets the expected phase and amplitude requirements, and precisely injects it into the specific electrical node with insufficient power supply. This flexible energy transmission mechanism directly and effectively alleviates the problems of three-phase load imbalance and local voltage drop at the end of the distribution network, significantly enhances the system's dynamic adjustment capability to cope with power fluctuations in complex scenarios, and thus effectively ensures the power quality of terminal equipment and the long-term operational stability of the overall distribution network.

[0030] Through the above technical solution, the system first uses an updated lightweight neural network model to perform deep nonlinear mapping on the real-time feature vector of the target distribution network, overcoming the limitation of traditional governance methods that heavily rely on preset fixed standard voltages. This allows for the precise and dynamic derivation of the most suitable reference voltage for the current operating conditions, combining real-time grid fluctuations with the energy storage status of the equipment itself. Subsequently, by injecting the pre-confirmed action parameter space vector into the virtual synchronous machine dynamics model, the dynamically updated damping characteristics effectively suppress system frequency fluctuations and secondary oscillations caused by low voltage. Supported by virtual inertia, the system ensures the smoothness of the adjustment process, thus adaptively deriving an extremely stable real-time virtual angular displacement as the synchronization reference. Based on this, the system accurately captures the transient state between the ideal target and the actual drop state. The deviation is addressed by combining dynamically optimized compensation amplitude with real-time synchronous phase information to create an independent ideal sinusoidal reference voltage waveform for each AC phase. This enables the system to precisely manage and bias the low-voltage trapped phase. Next, by constructing a dual closed-loop feedback architecture for voltage and current, the ideal target is compared in real time with the actual output state of the inverter, significantly improving the system's response speed to sudden disturbances and waveform tracking accuracy. Finally, by driving the inverter through modulation ratio control commands to perform a flexible conversion and injection of digital decision-making into physical energy, the system directly and effectively resolves the three-phase load imbalance and local voltage drop problems at the end of the distribution network. While ensuring the power quality of terminal equipment, this significantly enhances the adaptive adjustment capability and long-term operational stability of the entire distribution network under complex power fluctuation scenarios.

[0031] In this embodiment, the method further includes: determining an instant reward value based on a reference voltage, a real-time effective voltage value, the state information matrix of the target distribution network at the current moment, the total harmonic distortion rate, and waveform stability; constructing an optimized loss function for the evaluator model based on the state information matrix of the target distribution network at the current moment, the state information matrix at the next moment, the action parameter space vector, and the instant reward value; calculating the gradient of the optimized loss function and updating the network parameters in the evaluator model based on the gradient; determining the action value function predicted by the updated evaluator model for the state information matrix and action parameter space vector at the current moment; constructing an optimized objective function for the actor model based on the action value function with the goal of maximizing the expected return; calculating the gradient of the optimized objective function and updating the network parameters of the actor model based on the gradient and a preset learning rate.

[0032] In this embodiment, the total harmonic distortion rate (THD) refers to the ratio of the root mean square value of each harmonic component in the AC power grid to the effective value of the fundamental wave, used to quantitatively characterize the degree of distortion of the power signal deviating from the ideal sinusoidal state. Waveform stability refers to the characteristic index of the AC electrical signal maintaining its physical form and amplitude stability within a certain time period. The instant reward value refers to the quantitative evaluation feedback score calculated within the intelligent learning framework based on the multi-dimensional power quality improvement effect after the current compensation action is executed, used to measure the actual quality of the pre-control action. The state information matrix refers to a data structure that deeply integrates multi-dimensional data such as external environmental characteristics, real-time power indicators, current equipment settings, and system global plans. The next moment refers to the new round of data sampling and evaluation cycle in the discrete digital control system after the current action command is actually applied to the power grid environment and the system state transitions. The Critic model is a neural network module specifically responsible for predicting and evaluating the long-term benefit quality of the system taking a certain action in a specific state. The optimization loss function is a mathematical formula used to measure the error between the value prediction result of the Critic model and the target value derived based on the actual reward feedback. The gradient is the direction vector of a mathematical function with the largest rate of change at a specific point, often used as a derivative to guide the iterative evolution of artificial intelligence algorithms towards a better solution. Network parameters refer to core data variables such as the weight matrix and bias matrix that constitute the connection strength of nodes within an artificial intelligence model. The action value function is a quantitative indicator, predicted by the evaluator model, representing the total accumulated benefit that a specific combination of actions in a given state can obtain over a future period. Expected return refers to the statistically optimal expectation of the overall benefit or comprehensive governance effect that the system strives to achieve in the long-term flexible governance process. The actor model is a decision network module responsible for reasoning and directly generating the optimal action strategy combination based on the acquired system environment state information. The optimization objective function is a mathematical guiding formula aimed at maximizing the expected benefit of the control strategy while maintaining a certain entropy characteristic in the strategy distribution. The preset learning rate is a hyperparameter pre-set during algorithm training iterations to control the step size and speed of iterative updates of the policy network weights.

[0033] Through this step, the system constructs a closed-loop deep reinforcement learning mechanism based on the collaborative interaction between actors and commentators. This mechanism can extract objective quantitative evaluation feedback based on the actual voltage deviation, harmonic conditions, and stability characteristics after each governance step, and use this feedback to drive the continuous self-updating of the computational weights of the internal decision-making and evaluation networks. This autonomous evolutionary mechanism effectively endows the device with dynamic adaptability to complex and ever-changing power conditions, enabling it to operate independently of human intervention and continuously correct control deviations through trial and error evaluation. Furthermore, guided by long-term benefits, it adaptively deduces flexible governance strategies suitable for the current rapidly changing power grid environment.

[0034] In this embodiment of the application, the instant reward value is determined according to the following formula:

[0035] in, This is an instant reward value; Reference voltage; This is the real-time effective voltage value; Total harmonic distortion (THD); For waveform stability; and These are weight parameters; This is the state information matrix at the current moment.

[0036] In this embodiment, the optimized loss function satisfies the following formula:

[0037] in, To optimize the loss function; For expected operators; For the network parameters of the evaluator model; It is a vector of action parameter space; This is an instant reward value; Discount factor; This is the state information matrix for the next time step; This is the state information matrix at the current moment; This is the action value function.

[0038] In this embodiment of the application, the objective function satisfies the following formula:

[0039] in, To optimize the objective function; For the network parameters of the actor model; The policy probability is to output the action parameter space vector under the state information matrix at the current moment; The hyperparameters for controlling entropy weights; Entropy for the strategy dimension.

[0040] In one embodiment, the governance effect of the current control action is first quantitatively evaluated using a reward function. The specific formula for calculating the immediate reward value is as follows:

[0041] in, and These are weighting parameters used to balance the importance of various governance indicators.

[0042] Subsequently, an experience replay mechanism was employed to train the system using a small batch of datasets. In the value assessment logic of the evaluator model, a "post-event reflection" mechanism for the recently completed control cycle was introduced. Specifically, the system is based on the newly arrived state information matrix... Combined with the action parameter space vector of the previous moment that prompted the system to complete the state transition (i.e., reach the new state). The effectiveness of the action is evaluated to obtain the predicted value function of the new state. The reflection process satisfies the following intermediate expectation formula:

[0043]

[0044] in, The new state information matrix of the system at the next time step. Then, the action parameter space vector of the previous time step in the new state information matrix. The reflective value assessment function is to evaluate whether the action just taken was good or bad from the perspective of the current new result.

[0045] Based on this, the optimal loss function of the evaluator model is constructed by minimizing the Bellman error:

[0046] To update the network parameters, the gradient of the optimization loss function is calculated, and the specific formula is as follows:

[0047] in, To optimize the gradient of the loss function, the network parameters of the evaluator model are updated through backpropagation based on the obtained gradient.

[0048] Next, with the goal of maximizing expected return, we construct the optimization objective function for the Actor model:

[0049] in, This is the entropy of the strategy dimension. The specific intermediate calculation process for this entropy is as follows:

[0050] in, The standard deviation of the action parameters generated based on the state information matrix; is the dimension of the action space vector; Pi is a constant.

[0051] To update the Actor model, the gradient of the objective function is calculated, and the intermediate derivation process is as follows:

[0052] in, This indicates the presence of the experience replay dataset. Calculate the expectation of a state-action combination. Let be the logarithmic gradient of the policy probability.

[0053] Finally, the network parameters of the Actor model are iteratively updated according to the gradient and the preset learning rate using the following formula:

[0054] in, This represents the network parameter matrix before the actor model is updated; This represents the updated network parameter matrix; The preset learning rate is used to determine the model's iteration step size.

[0055] In practical engineering deployments, considering that the aforementioned series of expectation calculations and gradient matrix differentiation require a large amount of tensor computing resources, when the multi-dimensional state perception module detects that the distribution network is in a relatively stable period and the edge computing device has computing power redundancy, the system will complete the aforementioned batch gradient updates locally in a closed loop. If the network fluctuates drastically and the local computing power load is too high, the system can switch to edge inference mode, relying solely on local caching. Once the triggering conditions are met, the data is uploaded and processed by the cloud based on the gradient formula described above. The updated network parameters of the evaluator model and the actor model are then distributed.

[0056] In this embodiment of the application, the step of confirming the action parameter space vector includes: determining the weighted Euclidean distance between the state information matrix of the target distribution network at the current moment and the applicable environment description matrix corresponding to each model in the preset model database, and taking the model with the smallest distance as the target actor model; inputting the state information matrix into the target actor model for inference calculation, and outputting the action parameter space vector.

[0057] In this embodiment, the preset model database refers to a data collection library that is pre-built and centrally stored in physical storage media or a cloud platform, covering various alternative system control strategy models accumulated from historical operation or evolved through training under multiple typical operating conditions. The applicable environment description matrix refers to a benchmark data body that is mapped and bound one-to-one with the aforementioned alternative control strategy models, used to structurally characterize the multi-dimensional reference features such as external temperature and humidity, alternating load levels, and time-period global plans most suitable for the specific model to handle. The weighted Euclidean distance is a mathematical geometric statistical algorithm used to measure the logical similarity between two feature sets in a high-dimensional data space. Based on calculating the straight-line distance in conventional space, it assigns differentiated product coefficients to different dimensional features such as electricity consumption plans and time slicing that have different impact proportions on the overall power grid state evolution. The target actor model refers to a specific decision network module extracted from the vast candidate library after rigorous comparison and screening of the aforementioned spatial similarity, currently determined by the system to have the highest matching degree with the actual transient operating conditions of the power grid. Inference computation refers to the process of feeding real-time acquired multiple power grid operation status features into a selected decision network module, and using its internally fixed neuron connection weights and activation logic to perform forward mapping to quickly derive specific control parameters.

[0058] In this embodiment, the motion parameter space vector is determined according to the following formula:

[0059] in, This is the space vector of action parameters at the current time t; This is the state information matrix at the current time t; These are policy parameters obtained based on the state information matrix; The standard deviation is obtained based on the state information matrix; The parameters are random noise and follow a standard normal distribution. ; This is the activation function used to restrict the output to the range [-1, 1]. This indicates element-wise multiplication.

[0060] In one embodiment, when the system detects events requiring model changes, such as the entry into a new time period or modifications to the electricity consumption plan, the system initiates an edge intelligent inference and model selection mechanism. A preset model database stores multiple system models, each associated with an applicable environment description matrix, as follows:

[0061] The state information matrix is ​​as follows:

[0062] in, This is an environmental description matrix; The environmental feature vector contains real-time temperature T, humidity H, and partial discharge signal PD, which is used to characterize the coupled influence of the external environment on load fluctuations. This is a characteristic vector of electrical quantities, containing the effective values ​​of the three-phase voltages. Three-phase current RMS value and total harmonic distortion It is used to characterize the current voltage deviation and waveform distortion state of the power grid.

[0063] The current setting vector for the device includes the virtual synchronizer VSG control parameters (moment of inertia). Damping coefficient ), voltage and current dual-loop PI controller gain The reference voltages of each phase output by the MLP neural network and the network weight parameters of the lightweight neural network model MLP , . This is the reactive power compensation ratio gain; This is the integral gain for reactive power compensation; This is the excitation gain; This is the voltage loop proportional gain; This is the voltage loop integral gain; This is the proportional gain of the current loop; This is the integral gain of the current loop.

[0064] This is a system-wide vector containing the electricity consumption plan identifier, distributed power source access coefficient, and time sharding index; for the electricity consumption plan... and the number of distributed power sources connected The expected power consumption and the degree of disorder in distributed power sources are represented by continuous values ​​within a certain range; for time slicing... Considering the varying degrees of chaos and numerous influencing factors during peak and off-peak periods, the system is divided into half-hour time slices, with each time slice containing a separate model, represented by discrete slice numbers.

[0065] The system uses fast clustering algorithms such as weighted k-means to calculate the weighted Euclidean distance between the current state information matrix and the description matrices of each applicable environment. In calculating this distance, considering that electricity usage plans and time slicing have a more significant impact on the future state of the end-point power grid, the system assigns higher weight parameters to these global vector features. After calculation, the system selects the model with the smallest weighted Euclidean distance as the system model with the most similar scenario, i.e., the target actor model.

[0066] After determining the target actor model, the system enters the action learning and parameter generation phase. The current state information matrix is ​​input into the target actor model for forward inference, and the network outputs the corresponding policy parameters. and standard deviation Subsequently, combining the Gaussian distribution's random exploration mechanism, the motion parameter space vector under the current working condition is generated using the following specific mathematical formula:

[0067] The generated motion parameter space vector encompasses multi-dimensional system control commands, and its complete vector form is as follows:

[0068] in, This is the adjustment value for the moment of inertia; This is the adjustment value for the damping coefficient; This is the reactive power compensation ratio gain adjustment value; This is the reactive power compensation integral gain adjustment value; This is the excitation gain adjustment value; Voltage loop proportional gain adjustment value; This is the voltage loop integral gain adjustment value; This is the proportional gain adjustment value for the current loop; This is the adjustment value for the current loop integral gain; and Adjust the values ​​for the neural network parameters.

[0069] These fine-tuned gain values ​​will subsequently be mapped and applied to the adaptive updates of virtual synchronizer parameters, dual-loop control parameters, and weights of the lightweight neural network model (MLP).

[0070] To prevent destructive physical impacts on the inverter's AC output during transient processes such as direct model switching or drastic updates to operating parameters, a weighted moving average scheme can be introduced for smooth transition when loading the generated operating parameters. Specifically, the following formula can be used to perform a sliding iteration over time on the new and old control parameters:

[0071] in, (Left side) shows the updated target evaluator model weights; It is a smoothing factor (soft update coefficient), with a value range of (0,1), which determines the proportion of the new parameters in the fusion process; For the network parameters of the evaluator model; (Right side) shows the current weights of the target evaluator model.

[0072] In this embodiment, the motion parameter space vector includes the moment of inertia adjustment value, damping coefficient adjustment value, reactive power compensation proportional gain, reactive power compensation integral gain, excitation gain, voltage loop proportional gain, voltage loop integral gain, current loop proportional gain, current loop integral gain, and neural network parameter adjustment value; the neural network parameter adjustment value is used to update the lightweight neural network model.

[0073] In this embodiment, the neural network parameter adjustment values ​​are as described above. .

[0074] In one embodiment, the system parses the action parameter space vector using a pre-processing algorithm to extract neural network parameter adjustment values ​​for fine-tuning the decision network structure, thereby updating the lightweight neural network model. Specifically, this model employs a shallow multilayer perceptron architecture trained using a flexible policy actor and commentator algorithm. Its structure includes an input layer, one or more hidden layers, and an output layer, where the hidden layers use a linear rectified activation function to perform efficient nonlinear mapping calculations.

[0075] During the data input phase, the system's multidimensional state perception module and underlying hardware acquire relevant real-time operating status data to construct the real-time feature vector of the multilayer perceptron model (lightweight neural network model). The specific derivation and construction formula of this real-time feature vector are as follows:

[0076] in, For real-time feature vectors; The voltage of the DC bus; The connection capacitor capacity status; The deviation between the real-time effective voltage value of the low-voltage trapped phase (the current effective voltage value of the phase to be treated) and the previous ideal voltage represents the current gap in treatment needs.

[0077] During the forward propagation phase of the hidden layers, the network performs layer-by-layer dimensionality reduction and feature extraction using dynamically updated weights and biases. Each hidden layer is configured with a weight matrix updated from the action parameter space vector. Set initial transfer characteristics. The core inter-layer iterative mapping formula is expanded as follows:

[0078] in, The output feature matrix of the nth hidden layer after linear transformation and activation; This is the weight matrix for this layer; This is the output feature matrix of the previous hidden layer; This is the bias matrix for this layer.

[0079] In the output stage, after multiple layers of implicit mapping, the output layer integrates the final features through a specific activation function, dynamically deriving and generating the fundamental target values ​​used to guide inverter control. The specific calculation formula is as follows:

[0080] in, The activation function operator used by the output layer; This is the comprehensive feature matrix passed from the terminal hidden layer.

[0081] The above steps not only take into account the voltage drop that urgently needs compensation, but also deeply integrate the actual physical energy storage potential on the DC bus side, thereby avoiding waveform distortion caused by commands exceeding hardware output limits. Simultaneously, its internal weight and bias matrices are adaptively evaluated and updated by the reinforcement learning algorithm in each control cycle, ensuring that this lightweight neural network model maintains excellent nonlinear fitting accuracy even under long-term operating conditions such as grid topology changes or aging energy storage capacitors. This provides highly dynamic algorithmic support for generating a flexible and realistic reference voltage.

[0082] In this embodiment, the step of determining the sinusoidal reference voltage of each phase based on the action parameter space vector, the initial bias angle of each phase in the virtual synchronous machine, the reference reference voltage, the virtual angular displacement, and the real-time effective voltage value of the low-voltage trapped phase of the target distribution network includes: determining the voltage sag deviation value based on the reference reference voltage and the real-time effective voltage value; adaptively calculating the voltage sag deviation value using the reactive power compensation proportional gain and reactive power compensation integral gain to obtain the reactive power compensation reference value; determining the voltage amplitude based on the reactive power compensation reference value, the excitation gain, and the reference reference voltage; and determining the sinusoidal reference voltage of the corresponding phase based on the initial bias angle of each phase, the virtual angular displacement at the current moment, and the voltage amplitude.

[0083] In this embodiment, the voltage sag deviation value refers to the transient difference between the ideal target voltage parameter and the actual root-mean-square value obtained by instantaneous sampling of the underlying hardware during the operation of the distribution network, reflecting the degree of voltage deficit that the system urgently needs to address. The reactive power compensation proportional gain refers to the product coefficient used in the control logic to linearly amplify the aforementioned voltage deficit degree in real time, determining the initial response strength of the system to voltage errors. The reactive power compensation integral gain refers to the product coefficient used to amplify the cumulative amount of the aforementioned voltage deficit degree over time, aiming to eliminate the steady-state error of the system and improve control accuracy. Adaptive calculation refers to the digital calculation process by which the system automatically matches and adjusts the input deviation using dynamically updated control parameters based on changes in the external environment and internal state. The reactive power compensation reference value refers to the intermediate guidance parameter output by the system after the above logical processing, used to indicate the expected reactive power regulation demand that the inverter needs to inject externally. The excitation gain refers to the product coefficient simulating the internal magnetic field regulation characteristics of a traditional synchronous generator, used to convert the aforementioned reactive power regulation demand into a proportional factor corresponding to the voltage regulation scale. Voltage amplitude refers to the peak value that an AC signal can reach within one cycle. Here, it is used as the ideal compensation contour envelope synthesized by the system after integrating static basic targets and dynamic reactive power regulation requirements.

[0084] By accurately extracting the difference between the actual and ideal states of the power grid and introducing proportional-integral calculations to simulate the excitation regulation mechanism of a generator, the regulation amount and envelope required to compensate for this difference can be dynamically derived. This design not only endows the compensation control process with smooth and steady-state error-free regulation characteristics, but also successfully constructs an ideal waveform reference that combines dynamic response and accurate tracking by synthesizing the regulated amplitude with the independent synchronous phase of each phase. This significantly improves the compensation accuracy and flexibility of the flexible management device in response to voltage drops in the distribution network.

[0085] In this embodiment of the application, the reactive power compensation reference value is determined according to the following formula:

[0086] in, This is the reference value for reactive power compensation at the current time t; This is the reactive power compensation ratio gain; This is the integral gain for reactive power compensation; This represents the voltage drop deviation value at the current time t. This represents the integral time variable from the initial time 0 to the current time t, used to represent the historical accumulation process; Indicates the integral time The differential.

[0087] In this embodiment, the voltage amplitude is determined according to the following formula:

[0088] in, The voltage amplitude at the current time t; The reference voltage at the current time t; This is the excitation gain; This is the reference value for reactive power compensation at the current time t.

[0089] In the embodiments of this application, the sinusoidal reference voltage of each phase is determined according to the following formula:

[0090]

[0091] in, The subscript indicates the phase sequence, representing phase a, phase b, or phase c. The composite phase angle of phase x is obtained based on the virtual angular displacement and initial offset angle of phase x. This represents the virtual angular displacement at the current time t. Let x be the initial offset angle of phase x; The sinusoidal reference voltage of phase x at the current time t; The voltage amplitude at the current time t; Pi is a constant.

[0092] In one embodiment, the transient difference between the ideal controlled voltage and the actual voltage drop is first calculated. Based on the reference voltage dynamically output by the lightweight neural network model and the real-time effective voltage value of the low-voltage trapped phase in the target distribution network, the voltage drop deviation at the current time t is determined using the following formula:

[0093] in, The voltage drop deviation value of the low-voltage trapped phase in the target distribution network; The reference voltage for the low-voltage trapped phase of the target distribution network; The real-time effective voltage value of the low-voltage trapped phase in the target distribution network.

[0094] Subsequently, an adaptive calculation is performed using a PI mechanism to obtain the expected required adjustment amount. The voltage sag deviation is processed using the reactive power compensation proportional gain and reactive power compensation integral gain to obtain the reactive power compensation reference value. The derivation formula is as follows:

[0095] Next, the system uses excitation gain to adjust the voltage amplitude, achieving flexible amplitude compensation. Based on the reactive power compensation reference value, excitation gain, and reference voltage, the voltage amplitude at the current time t is determined. The calculation formula is as follows:

[0096] Finally, the system uses the updated phase information to synthesize a sinusoidal reference voltage, ensuring synchronization with the grid and bias towards the lower voltage phase. Based on the initial bias angle of each phase, the virtual angular displacement at the current moment, and the voltage amplitude, the sinusoidal reference voltage for the corresponding phase is determined. The detailed calculation formulas are as follows:

[0097]

[0098] In the above formula 、 and The core control gain is not a traditional static fixed empirical value, but rather a space vector of action parameters generated by inference and dynamic optimization of the target actor model in the reinforcement learning (SAC algorithm) model in the previous step. This allows the reactive voltage control loop of the virtual synchronous machine to adaptively optimize according to the actual severity of grid voltage drops, prioritizing the improvement of the lowest phase voltage. Simultaneously, this process overcomes the limitations of relying on traditional fixed standard voltages of 220V or 380V. By introducing a reference voltage dynamically mapped by a lightweight neural network, it fully considers the capacitor energy storage state of the DC bus, thereby ensuring the accuracy and dynamic adaptability of the inverter's flexible compensation command output at the underlying electrical calculation level.

[0099] In this embodiment, the step of determining the virtual angular displacement of the virtual synchronizer at the current moment based on the virtual angular velocity and virtual angular displacement, active power deviation, preset reference angular velocity, and motion parameter space vector of the virtual synchronizer at the previous moment includes: updating the moment of inertia and damping coefficient of the virtual synchronizer through the moment of inertia adjustment value and damping coefficient adjustment value; determining the virtual angular velocity of the virtual synchronizer at the current moment based on the updated moment of inertia and damping coefficient, the virtual angular velocity at the previous moment, active power deviation, and preset reference angular velocity; and determining the virtual angular displacement of the virtual synchronizer at the current moment based on the virtual angular velocity at the current moment and the virtual angular displacement at the previous moment.

[0100] In this embodiment, the moment of inertia adjustment value refers to a specific numerical variable in the action parameter space vector output by the preceding deep reinforcement learning model, specifically used for dynamically fine-tuning the strength of the system's inertial response. The damping coefficient adjustment value refers to a specific numerical variable generated by the aforementioned decision network model, specifically used for dynamically fine-tuning the system's ability to suppress power oscillations. Moment of inertia is a physical inertial parameter used to simulate the original motion characteristics of a traditional rotating generator rotor in a rotating state; its core function is to ensure a smooth and stable dynamic frequency response process when the system responds to external disturbances. The damping coefficient is an electrical parameter used to simulate the natural attenuation and physical dissipation of frequency fluctuations within a traditional generator. This parameter directly affects the system's frequency deviation term, aiming to effectively suppress local grid oscillations induced by low voltage drops and promote inter-phase power balance in power supply lines.

[0101] By introducing dynamically optimized fine-tuning variables to update the core physical parameters of the simulated rotor in real time, a virtual rotor dynamics model with adaptive capabilities was successfully constructed. This mechanism enables the compensation device to quickly quell system frequency fluctuations and potential oscillations when facing complex active power imbalances in the distribution network, utilizing updated damping characteristics, while ensuring a smooth overall adjustment trajectory through reasonable inertia configuration. Based on this, the system iteratively derives the required phase reference information layer by layer according to the calculated angular velocity increments, thus establishing a stable and reliable synchronous reference support for driving the low-voltage phase compensation of subsequent three-phase signals.

[0102] In this embodiment of the application, the active power deviation is determined according to the following formula:

[0103] in, The active power deviation at the current time t; Preset desired power; The period constant of the power frequency alternating current; , and Let be the actual output voltage of each of the three phases of the inverter at the current time t; , and Let be the actual output current of each of the three phases of the inverter at the current time t.

[0104] In this embodiment, the virtual angular velocity is determined according to the following formula:

[0105]

[0106] in, The updated damping coefficient; The updated moment of inertia; This is the virtual angular velocity at the previous moment; Preset reference angular velocity; The active power deviation at the current time t; This is the virtual angular velocity at the current time t; This is the constant for the sampling period of discrete control.

[0107] In this embodiment, the virtual angular displacement is determined according to the following formula:

[0108] in, Let be the virtual angular displacement at time t; This represents the virtual angular displacement at the time preceding time t; Let be the virtual angular velocity at time t; .

[0109] In one embodiment, firstly, the system extracts the motion parameter space vector output by the actor model in the preceding steps, and obtains the moment of inertia adjustment value and damping coefficient adjustment value from it. Using these two adjustment values, the original basic parameters of the virtual synchronizer are superimposed and dynamically updated to obtain the updated moment of inertia and damping coefficient suitable for the current cycle.

[0110] Before frequency iteration, the system needs to calculate the difference between the preset expected power and the actual output power. The system determines the active power deviation based on the time-domain integrals of the actual output voltage and actual output current of each of the three phases of the inverter at the current moment. The calculation formula is as follows:

[0111] Next, the system combines the updated damping coefficient and moment of inertia to construct a virtual rotor dynamics model and derive the dynamic evolution process of angular velocity. Based on the updated moment of inertia and damping coefficient, the virtual angular velocity at the previous moment, the active power deviation, and the preset reference angular velocity, the virtual angular velocity of the virtual synchronizer at the current moment is determined. The specific calculation formula for the intermediate process is as follows:

[0112]

[0113] in, The differential rate of change of angular velocity; This is the damping coefficient, which acts on the frequency deviation term to suppress low-voltage-induced oscillations and promote interphase balance; It represents the moment of inertia, used to ensure a smooth inertial response.

[0114] In one embodiment, the active power deviation calculation is not limited to static evaluation of the steady-state period, but is obtained by acquiring transient electrical parameters through a sampler module and continuously performing sliding integration. Simultaneously, the moment of inertia and damping coefficient are no longer limited to traditional preset fixed physical parameters, but are transformed into adaptive software variables driven by the SAC reinforcement learning algorithm. This allows the virtual synchronous machine to autonomously adjust virtual damping to quell secondary oscillations in the system frequency when the end-point grid encounters sudden disturbances or frequent switching of high-power nonlinear loads, and to provide smooth transient power support in conjunction with the virtual moment of inertia. This significantly enhances the robustness and high synchronization characteristics of the three-phase balanced flexible compensation process at the underlying physical control logic.

[0115] In this embodiment of the application, the step of determining the modulation ratio control command based on the action parameter space vector, the sinusoidal reference voltage, the actual output voltage of the inverter at the current moment, and the actual output current includes: determining the reference current command based on the sinusoidal reference voltage, the actual output voltage, the voltage loop proportional gain, and the voltage loop integral gain; and determining the modulation ratio control command based on the reference current command, the actual output current, the current loop proportional gain, and the current loop integral gain.

[0116] In this embodiment, the voltage loop proportional gain refers to the product factor used in the outer voltage loop control logic to linearly amplify the transient deviation between the target ideal waveform and the actual output potential difference, aiming to accelerate the system's response speed to waveform errors. The voltage loop integral gain refers to the product factor used to amplify the cumulative evolution of the aforementioned voltage deviation over time, aiming to eliminate long-term steady-state tracking errors and thus improve waveform fitting accuracy. The reference current command refers to the intermediate dynamic target parameter output by the system after outer loop voltage proportional and integral adjustment, specifically used to provide an ideal charge flow tracking benchmark for the underlying inner loop control. The current loop proportional gain refers to the product factor used in the inner current loop control logic to linearly amplify the instantaneous error between the aforementioned target benchmark and the actual physical charge flow value, ensuring the agility of the underlying control signal tracking. The current loop integral gain refers to the product factor used to comprehensively amplify the historical accumulation of errors in the current dimension, aiming to eliminate current steady-state errors and ensure smooth follow-up of output commands.

[0117] This step constructs a dual closed-loop regulation architecture with nested and coordinated outer voltage and inner current loops. By introducing dynamically optimized multidimensional gain parameters, the ideal compensation target is decoupled layer by layer and accurately transformed into the underlying physical drive signal. This suppresses sudden low-voltage disturbances in the power grid, minimizes mitigation deviations, and endows the inverter system with excellent waveform tracking accuracy and anti-interference stability under complex load conditions. It provides a solid computational and command conversion guarantee for the reliable injection of high-quality AC compensation energy.

[0118] In this embodiment, the reference current command is determined according to the following formula:

[0119] in, The subscript indicates the phase sequence, representing phase a, phase b, or phase c. The reference current command for phase x at the current time t; This is the voltage loop proportional gain; This is the voltage loop integral gain; The sinusoidal reference voltage of phase x at the current time t; This represents the actual output voltage of phase x of the inverter at the current time t. For integration variables The reference voltage of phase x; For integration variables The actual instantaneous voltage of phase x.

[0120] In this embodiment, the modulation ratio control command is determined according to the following formula:

[0121] in, The subscript indicates the phase sequence, representing phase a, phase b, or phase c. The modulation ratio control command for phase x at the current time t; This is the proportional gain of the current loop; This is the integral gain of the current loop; The reference current command for phase x at the current time t; This represents the actual output current of phase x of the inverter at the current time t. For integration variables The reference current of phase x; For integration variables The actual instantaneous current of phase x.

[0122] In one embodiment, firstly, for each phase, the system compares the transient potential difference between the target reference waveform and the actual output voltage at the lower level using the voltage outer loop, and generates a current command using an enhanced PI algorithm. Based on the sinusoidal reference voltage, the actual output voltage, the voltage loop proportional gain, and the voltage loop integral gain, the reference current command is adaptively determined. Taking any phase x (representing phase a, b, or c) as an example, the specific calculation formula for the intermediate process is as follows:

[0123] Subsequently, the system transmits the generated reference current command to the inner current loop for further adjustment and conversion. The current loop compares this current command with the actual current to generate the PWM modulation ratio for the underlying hardware driver. Specifically, the modulation ratio control command is determined based on the reference current command, the actual output current, the proportional gain of the current loop, and the integral gain of the current loop. The corresponding calculation formula is as follows:

[0124] In the above formula 、 、 All of these are issued through real-time adaptive optimization using the aforementioned target actor model combined with the current power grid state information matrix. Voltage loop gain and Used to minimize voltage deviation. Current loop gain. It is specifically designed for rapid response to low voltage disturbances. This enhanced PI dual-loop architecture based on reinforcement learning parameter injection enables extremely low latency operations in a digital signal processor (DSP), effectively ensuring the system's accurate tracking and anti-interference stability under rapidly changing power grid conditions.

[0125] Figure 2 A flowchart illustrating another distribution network flexibility management method according to an embodiment of this application is shown schematically. Figure 2As shown in this embodiment, the device first connects to the power grid and collects status data in real time, forming a current-moment status information matrix containing features such as environment and power consumption. Then, based on K-means or fast clustering algorithms, the most similar model in the preset model database is selected as the target actor model according to the current state. Next, based on the local state, the Actor model of the deep reinforcement learning algorithm SAC is invoked for inference to obtain a system parameter adjustment strategy containing various gain fine-tuning values. Using this strategy, the lightweight neural network MLP model updates its own parameters and calculates a dynamic reference target voltage based on the DC bus voltage and the connected capacitor capacity. Simultaneously, the virtual synchronous machine (VSG) updates its internal parameters such as moment of inertia and damping coefficient, simulating the generator rotor dynamics and outputting the current sinusoidal reference voltage for each phase. Afterward, the voltage and current dual closed-loop control unit updates the corresponding proportional and integral gain parameters, converting the deviation between the reference voltage and the actual output into a low-level PWM modulation ratio control command. The IGBT control module then uses this PWM... The M-parameter performs high-frequency switching actions to precisely regulate and output AC compensation voltage to the low-voltage trapped phase. After a single control operation is completed, the system collects the new state after governance, makes an immediate evaluation using a reward function that includes voltage deviation and harmonic distortion rate, and saves a governance record array containing state transitions and rewards locally. During the model evolution and update phase, the system first determines whether there is computing power redundancy. If conditions permit and the governance array length reaches a preset threshold, the SAC upgrade evaluation algorithm is run locally for evaluation and iterative evolution. Next, the system determines whether there is network connectivity. If the network is unobstructed, the local governance history dataset is transmitted to the cloud after a certain period. The Actor and Critic models for similar scenarios are globally learned, updated, and distributed using the cloud-based SAC reinforcement learning algorithm. Finally, the system determines whether to enter the next time slice or be deployed to a new environment. If so, it returns to the initial step to reselect the most similar model based on K-means. If not, it directly returns to the state data collection step and continues to perform high-frequency closed-loop bottom-level regulation under the current applicable model.

[0126] Figure 3 This diagram schematically illustrates the architecture of a flexible power distribution network management system according to an embodiment of this application. Figure 3As shown in this embodiment, the flexible power distribution network management system mainly consists of two parts: a cloud-based backend and local hardware devices, which work together to serve the three-phase power end management of the 0.4KV power distribution network. The cloud-based backend is equipped with a communication module to monitor the electrical parameters, environmental parameters, and equipment status of the current cabinet. Internally, it runs a "three-phase balanced" power end quality management model system based on the SAC algorithm, which can use historical state parameter matrices for model evaluation and centralized calculation upgrades. In terms of energy flow in the local hardware devices, the system includes a rectifier side and an inverter side. The front end extracts electrical energy from the AC side of the power distribution network through a rectifier containing a single-phase turn-off three-phase full-bridge rectifier. After LC filtering and energy storage, it maintains the stability of the DC bus voltage. The back end uses a three-phase full-bridge IGBT inverter as the core component to invert the energy from the DC bus into AC power and inject it directionally into the trapped phase of the grid. In terms of core control signal flow... The sampling module obtains three-phase voltage information from the power grid, which is then fed into the MLP reference voltage model and the PLL phase-frequency analysis unit to dynamically deduce the voltage reference and extract key frequency and phase information. The aforementioned reference characteristics and the voltage data obtained by the DC bus voltage monitoring unit are synchronously input to the specialized virtual synchronous machine (VSG) control unit to synthesize and output reference voltage commands for each phase with dynamic support capabilities. Subsequently, the voltage / current dual closed-loop control unit calculates and generates PWM duty cycle control commands based on the reference voltage and the actual electrical parameters fed back from the inverter side, and sends them to the underlying PWM control unit to drive the IGBT to perform precise switching actions. In addition, the external communication module and model storage module configured inside the device work together to not only ensure parameter interaction between the control logic and the cloud-based large model, but also realize persistent and secure caching of local system operation records and governance scenario models.

[0127] This application also provides a flexible distribution network management system, including a sampling module, a rectifier module, a DC bus, a control module, and an inverter module. The sampling module's acquisition terminal is electrically connected to the target distribution network, and its signal output terminal is communicatively connected to the control module. The inverter module's DC input terminal is electrically connected to the DC bus, its control terminal is communicatively connected to the control module, and its AC output terminal is electrically connected to the low-voltage trapped phase of the target distribution network. The control module is used to input the real-time feature vector of the target distribution network acquired by the acquisition module into an updated lightweight neural network model for nonlinear mapping calculation, and output the reference voltage of the virtual synchronizer at the current moment; based on the virtual angular velocity of the virtual synchronizer at the previous moment... The virtual angular displacement of the virtual synchronous machine at the current moment is determined by the degree and virtual angular displacement, active power deviation, preset reference angular velocity, and pre-confirmed action parameter space vector. Based on the action parameter space vector, the initial bias angle of each phase within the virtual synchronous machine, the reference voltage, the virtual angular displacement, and the real-time effective voltage value of the low-voltage trapped phase of the target distribution network, the sinusoidal reference voltage of each phase is determined. Based on the action parameter space vector, the sinusoidal reference voltage, the actual output voltage and actual output current of the inverter module at the current moment, the modulation ratio control command for each phase is determined. The inverter module is driven by the modulation ratio control command to convert the DC bus voltage into an AC compensation voltage and inject the AC compensation voltage into the low-voltage trapped phase.

[0128] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0129] This application also provides a machine-readable storage medium storing instructions that cause a machine to perform the above-described method.

[0130] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0131] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0132] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0134] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0135] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0136] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0137] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0138] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A flexible management method for power distribution networks, characterized in that, The method includes: The real-time feature vector of the target distribution network is input into the updated lightweight neural network model for nonlinear mapping calculation, and the reference reference voltage of the virtual synchronizing machine at the current moment is output. Based on the virtual angular velocity and virtual angular displacement, active power deviation, preset reference angular velocity, and pre-confirmed motion parameter space vector of the virtual synchronizer at the previous moment, the virtual angular displacement of the virtual synchronizer at the current moment is determined. Based on the action parameter space vector, the initial bias angle of each phase in the virtual synchronizing machine, the reference reference voltage, the virtual angular displacement, and the real-time effective voltage value of the low-voltage trapped phase of the target distribution network, the sinusoidal reference voltage of each phase is determined. Based on the action parameter space vector, the sinusoidal reference voltage, the actual output voltage and actual output current of the inverter at the current moment, the modulation ratio control command for each phase is determined. The inverter is driven by the modulation ratio control command to convert the DC bus voltage into an AC compensation voltage and inject the AC compensation voltage into the low-voltage trapped phase.

2. The method according to claim 1, characterized in that, The method further includes: The instantaneous reward value is determined based on the reference voltage, the real-time effective voltage value, the state information matrix of the target distribution network at the current moment, the total harmonic distortion rate, and the waveform stability. Based on the state information matrix of the target distribution network at the current moment, the state information matrix at the next moment, the action parameter space vector, and the instant reward value, construct the optimized loss function of the evaluator model; The gradient of the optimized loss function is calculated, and the network parameters in the evaluator model are updated based on the gradient. Determine the action value function predicted by the updated evaluator model for the current state information matrix and the action parameter space vector; With the goal of maximizing expected return, an optimization objective function for the actor model is constructed based on the action value function. The gradient of the optimization objective function is calculated, and the network parameters of the actor model are updated based on the gradient and the preset learning rate.

3. The method according to claim 2, characterized in that, The instant reward value is determined according to the following formula: in, The instant reward value; The reference voltage is the reference voltage. The real-time voltage RMS value; The total harmonic distortion rate is mentioned above; For the stability of the waveform; and These are weight parameters; This is the state information matrix at the current moment.

4. The method according to claim 3, characterized in that, The optimized loss function satisfies the following formula: in, The optimized loss function is... For expected operators; These are the network parameters of the evaluator model; The action parameter space vector; The instant reward value; Discount factor; This is the state information matrix for the next time step; This is the state information matrix at the current moment; Let be the action value function.

5. The method according to claim 4, characterized in that, The optimization objective function satisfies the following formula: in, Let the optimization objective function be... These are the network parameters of the actor model; To output the policy probability of the action parameter space vector under the state information matrix at the current moment; The hyperparameters for controlling entropy weights; Entropy for the strategy dimension.

6. The method according to any one of claims 1 to 5, characterized in that, The steps for confirming the motion parameter space vector include: Determine the weighted Euclidean distance between the state information matrix of the target distribution network at the current moment and the applicable environment description matrix of each model in the preset model database, and take the model with the smallest distance as the target actor model; The state information matrix is ​​input into the target actor model for inference calculation, and the action parameter space vector is output.

7. The method according to claim 6, characterized in that, The motion parameter space vector is determined according to the following formula: in, This is the space vector of action parameters at the current time t; This is the state information matrix at the current time t; These are the strategy parameters obtained based on the state information matrix; The standard deviation is obtained based on the state information matrix; The parameters are random noise and follow a standard normal distribution. ; This is the activation function used to restrict the output to the range [-1, 1]. This indicates element-wise multiplication.

8. The method according to claim 7, characterized in that, The motion parameter space vector includes the moment of inertia adjustment value, damping coefficient adjustment value, reactive power compensation proportional gain, reactive power compensation integral gain, excitation gain, voltage loop proportional gain, voltage loop integral gain, current loop proportional gain, current loop integral gain, and neural network parameter adjustment value; the neural network parameter adjustment value is used to update the lightweight neural network model.

9. The method according to claim 8, characterized in that, The steps for determining the sinusoidal reference voltage for each phase based on the action parameter space vector, the initial bias angle of each phase in the virtual synchronizer, the reference voltage, the virtual angular displacement, and the real-time effective voltage value of the low-voltage trapped phase of the target distribution network include: The voltage drop deviation value is determined based on the reference voltage and the real-time effective voltage value. The voltage drop deviation value is adaptively calculated using the reactive power compensation proportional gain and the reactive power compensation integral gain to obtain a reactive power compensation reference value. The voltage amplitude is determined based on the reactive power compensation reference value, the excitation gain, and the reference voltage. The sinusoidal reference voltage for each phase is determined based on the initial bias angle of each phase, the virtual angular displacement at the current moment, and the voltage amplitude.

10. The method according to claim 9, characterized in that, The reactive power compensation reference value is determined according to the following formula: in, This is the reference value for reactive power compensation at the current time t; The reactive power compensation ratio gain; The reactive power compensation integral gain; This represents the voltage drop deviation value at the current time t. This represents the integral time variable from the initial time 0 to the current time t, used to represent the historical accumulation process; Indicates the integral time The differential.

11. The method according to claim 10, characterized in that, The voltage amplitude is determined according to the following formula: in, The voltage amplitude at the current time t; The reference voltage at the current time t; The excitation gain is mentioned above; This is the reference value for reactive power compensation at the current time t.

12. The method according to claim 11, characterized in that, The sinusoidal reference voltage for each phase is determined using the following formula: in, The subscript indicates the phase sequence, representing phase a, phase b, or phase c. The composite phase angle of phase x is obtained based on the virtual angular displacement and initial offset angle of phase x. This represents the virtual angular displacement at the current time t. Let x be the initial offset angle of phase x; The sinusoidal reference voltage of phase x at the current time t; The voltage amplitude at the current time t; Pi is a constant.

13. The method according to claim 8, characterized in that, The steps for determining the virtual angular displacement of the virtual synchronizer at the current moment, based on the virtual angular velocity and virtual angular displacement, active power deviation, preset reference angular velocity, and motion parameter space vector of the virtual synchronizer at the previous moment, include: The moment of inertia and damping coefficient of the virtual synchronizer are updated using the aforementioned adjustment values ​​for moment of inertia and damping coefficient. The virtual angular velocity of the virtual synchronizer at the current moment is determined based on the updated moment of inertia and damping coefficient, the virtual angular velocity at the previous moment, the active power deviation, and the preset reference angular velocity. The virtual angular displacement of the virtual synchronizer at the current moment is determined based on the virtual angular velocity at the current moment and the virtual angular displacement at the previous moment.

14. The method according to claim 13, characterized in that, The active power deviation is determined according to the following formula: in, The active power deviation at the current time t; Preset desired power; The period constant of the power frequency alternating current; , and Let be the actual output voltage of each of the three phases of the inverter at the current time t; , and Let be the actual output current of each of the three phases of the inverter at the current time t.

15. The method according to claim 14, characterized in that, The virtual angular velocity is determined according to the following formula: in, The updated damping coefficient; The updated moment of inertia; This is the virtual angular velocity at the previous moment; The preset reference angular velocity; The active power deviation at the current time t; This is the virtual angular velocity at the current time t; This is the constant for the sampling period of discrete control.

16. The method according to claim 15, characterized in that, The virtual angular displacement is determined according to the following formula: in, Let be the virtual angular displacement at time t; This represents the virtual angular displacement at the time preceding time t; Let be the virtual angular velocity at time t; .

17. The method according to claim 8, characterized in that, The steps for determining the modulation ratio control command based on the action parameter space vector, the sinusoidal reference voltage, and the inverter's actual output voltage and current at the current moment include: The reference current command is determined based on the sinusoidal reference voltage, the actual output voltage, the voltage loop proportional gain, and the voltage loop integral gain. The modulation ratio control command is determined based on the reference current command, the actual output current, the current loop proportional gain, and the current loop integral gain.

18. The method according to claim 17, characterized in that, The reference current command is determined according to the following formula: in, The subscript indicates the phase sequence, representing phase a, phase b, or phase c. The reference current command for phase x at the current time t; The voltage loop proportional gain; The voltage loop integral gain; The sinusoidal reference voltage of phase x at the current time t; The actual output voltage of phase x of the inverter at the current time t; For integration variables The reference voltage of phase x; For integration variables The actual instantaneous voltage of phase x.

19. The method according to claim 18, characterized in that, The modulation ratio control command is determined according to the following formula: in, The subscript indicates the phase sequence, representing phase a, phase b, or phase c. The modulation ratio control command for phase x at the current time t; The current loop proportional gain; The integral gain of the current loop; The reference current command for phase x at the current time t; Let x be the actual output current of the inverter in phase x at the current time t; For integration variables The reference current of phase x; For integration variables The actual instantaneous current of phase x.

20. A flexible management system for power distribution networks, characterized in that, It includes a sampling module, a rectifier module, a DC bus, a control module, and an inverter module; The sampling module's acquisition end is electrically connected to the target power distribution network, and the sampling module's signal output end is communicatively connected to the control module. The DC input terminal of the inverter module is electrically connected to the DC bus, the control terminal of the inverter module is communicatively connected to the control module, and the AC output terminal of the inverter module is electrically connected to the low-voltage trapped phase of the target distribution network. The control module is used to input the real-time feature vector of the target distribution network collected by the acquisition module into the updated lightweight neural network model for nonlinear mapping calculation, and output the reference reference voltage of the virtual synchronizing machine at the current moment. Based on the virtual angular velocity and virtual angular displacement, active power deviation, preset reference angular velocity, and pre-confirmed motion parameter space vector of the virtual synchronizer at the previous moment, the virtual angular displacement of the virtual synchronizer at the current moment is determined. Based on the action parameter space vector, the initial bias angle of each phase in the virtual synchronizing machine, the reference reference voltage, the virtual angular displacement, and the real-time effective voltage value of the low-voltage trapped phase of the target distribution network, the sinusoidal reference voltage of each phase is determined. Based on the action parameter space vector, the sinusoidal reference voltage, the actual output voltage and actual output current of the inverter module at the current moment, the modulation ratio control command for each phase is determined; the modulation ratio control command drives the inverter module to convert the DC bus voltage into an AC compensation voltage, and injects the AC compensation voltage into the low-voltage trapped phase.

21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 19.

22. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores instructions for causing the machine to perform the method according to any one of claims 1 to 19.