Intelligent regulation type low-voltage intelligent power distribution control system

By introducing time-varying priority and correlation weighting terms, and combining real-time guidance rewards and distributed power source uncertainty rewards, the control strategy of the low-voltage intelligent power distribution system is optimized, which solves the problems of lack of coordination in control actions and voltage and power flow oscillations, and achieves safe, stable and rapid response under sudden environmental changes.

CN120855557BActive Publication Date: 2026-01-13PRIMA NEW ENERGY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510945468.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2026-01-13
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Low-voltage intelligent power distribution systems suffer from a lack of coordinated control actions, causing voltage and power flow to oscillate near safety limits. They also neglect the safety margin requirements of the power distribution system when photovoltaic and wind power fluctuate significantly, resulting in poor regulation stability and difficulty in quickly responding to the effects of voltage deviation and load change rate.

Method used

By introducing time-varying priority and relevance weighting terms, and adding real-time guidance rewards and distributed power uncertainty rewards, the control strategy is optimized. Through load condition analysis, measurement uncertainty assessment, control strategy planning, control reward mechanism construction, control history experience review and system training, the control coordination and response speed are improved.

Benefits of technology

To maintain the safety and stability of the power distribution system under sudden environmental changes, improve regulation efficiency, respond quickly to voltage deviations and load changes, and enhance the ability to accommodate photovoltaic and wind power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120855557B_ABST
    Figure CN120855557B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent regulation and control type low-voltage intelligent power distribution control system, including load condition analysis module, measurement uncertainty evaluation module, regulation and control strategy planning module, regulation and control reward mechanism construction module, regulation and control historical experience review module, system training module and intelligent regulation and control module.The application belongs to the field of power distribution control, specifically refers to a kind of intelligent regulation and control type low-voltage intelligent power distribution control system, the scheme introduces time-varying priority, and adds correlation degree weighting item in strategy probability, automatically promotes control priority, avoids secondary disturbance caused by disorder switching;On the basis of termination reward, join real-time guidance reward and distributed power uncertainty reward, keep the safety and stability of power distribution system;Power quality reward factor emphasizes the more serious power distribution interaction sample that voltage deviates from nominal value;Load change rate factor focuses on sudden load event;Introduce sampling number attenuation, avoid excessive dependence on a small number of key power distribution interaction sample;Further improve the regulation and control effect of power distribution system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power distribution control, specifically to an intelligent control type low-voltage intelligent power distribution control system. Background Technology

[0002] Low-voltage intelligent distribution systems utilize advanced information technology and intelligent algorithms to automatically monitor, control, and protect low-voltage distribution networks. They can analyze load in real time, assess measurement uncertainties, coordinate the actions of various control units, and ensure the safe and stable operation of the distribution network under complex conditions. However, conventional distribution systems suffer from a lack of coordinated control actions, which can easily generate secondary disturbances when multiple nodes switch simultaneously. This causes voltage and power flow to oscillate near safety limits, neglecting the safety margin requirements of the distribution system when photovoltaic and wind power fluctuate significantly, leading to poor regulation stability. Furthermore, conventional distribution systems often ignore the impact of voltage deviation and load change rate on distribution safety and stability, making it difficult for control strategies to quickly optimize responses to voltage exceeding limits or sudden load increases. Summary of the Invention

[0003] To address the above issues and overcome the shortcomings of existing technologies, this invention provides an intelligent low-voltage intelligent power distribution control system. Addressing the problems of uncoordinated control actions in general power distribution systems, which easily lead to secondary disturbances when multiple nodes switch simultaneously, causing voltage and power flow to oscillate near safety limits, and neglecting the safety margin requirements of the power distribution system during large fluctuations in photovoltaic and wind power, resulting in poor control stability, this solution introduces time-varying priorities to each edge control unit and adds a correlation weighting term for other units to the strategy probability. This automatically increases the control priority for hotspot nodes where voltage and power flow exceed limits. When multiple units act, they consider the states of related units, avoiding secondary disturbances caused by disordered switching and improving overall control efficiency. In addition to the termination reward, real-time guidance rewards and distributed power source uncertainty rewards are added, ensuring non-zero feedback at each step. To avoid the sparsity problem caused by pure termination rewards, this solution simultaneously minimizes voltage deviation, line loss, and operational jitter, thereby maintaining the safety and stability of the distribution system even under sudden environmental changes. Addressing the common problem in general distribution systems where the impact of voltage deviation and load change rate on distribution safety and stability is ignored, leading to difficulties in quickly optimizing responses to voltage overruns or sudden load increases, this solution introduces a power quality reward factor, focusing on voltage overruns and emphasizing distribution interaction samples with more severe voltage deviations from nominal values. It also introduces a load change rate factor, focusing on sudden load events and enhancing the sampling probability of distribution interaction samples with instantaneous large load fluctuations. Furthermore, it introduces a sampling count decay to avoid over-reliance on a small number of key distribution interaction samples, balancing the reuse of high-priority experience with the continuous exploration of new distribution interaction samples, thereby improving the control effect of the distribution system.

[0004] The technical solution adopted by the present invention is as follows: The present invention provides an intelligent control type low-voltage intelligent power distribution control system, including a load condition analysis module, a measurement uncertainty assessment module, a control strategy planning module, a control reward mechanism construction module, a control historical experience review module, a system training module, and an intelligent control module;

[0005] The load condition analysis module collects load and energy storage output to obtain the net power of the low-voltage distribution network.

[0006] The measurement uncertainty assessment module collects bus voltage and branch power flow data and injects Gaussian noise to obtain the final observation status;

[0007] The control strategy planning module defines a set of actions for each edge control unit and allocates scheduling priorities based on voltage deviation and node correlation.

[0008] In addition to the safe termination reward, the regulation reward mechanism construction module constructs a real-time feedback signal through implicit reward function differential guidance reward and distributed power source uncertainty reward.

[0009] The regulation history experience review module calculates the priority of distribution interaction samples based on factors such as voltage deviation, load change rate and time difference error, and dynamically adjusts the sampling weight in combination with the attenuation mechanism;

[0010] The system training module is based on historical data and uses the DQN framework to sample and update Q network parameters from the experience replay buffer, thereby training the complete power distribution system.

[0011] The intelligent control module acquires real-time data on load, energy storage, bus voltage, and branch power flow, and performs intelligent control based on the selected control actions.

[0012] Furthermore, the load condition analysis module collects load power. and energy storage output Treating the load center as a scalable power region provides time-varying state input for intelligent scheduling, as follows: ; ; ;in, and These are load and distributed generation rate, respectively; among which, It is the change in load power; It is the change in energy storage output; It is a time interval; and It is the rate of change; and These are the net power of the low-voltage distribution network at time t+1 and time t, respectively.

[0013] Furthermore, the measurement uncertainty assessment module acquires the bus voltage. Branch Road Trend And add Gaussian noise, which is represented as: ;in, This represents the original measurement state of the i-th node; This is the final observed state; The mean is 0 and the variance is Gaussian noise.

[0014] Furthermore, the control strategy planning module deploys an edge control unit at each node of the power distribution system. The actions of each edge control unit include: adjusting voltage and reactive power; initiating energy storage charging and discharging commands; and dynamically allocating priority weights. Furthermore, an action coordination mechanism is introduced, and the action strategy is represented as follows: ; ;in, It is in state The probability that the i-th edge control unit selects action a; It is the Pearson correlation coefficient between the i-th and j-th edge control units; It is the reward function value of the i-th edge control unit selecting action a in state s; It is a priority adjustment factor; This is the rated bus voltage.

[0015] Furthermore, the regulation reward mechanism construction module introduces a real-time guidance reward, in addition to the termination reward when the power distribution system recovers to a safe operating condition, and also introduces an uncertainty reward for distributed power sources; the total reward... Represented as: Real-time guidance and rewards Represented as: Uncertainty-based rewards for distributed power sources Represented as: M is the set of distributed power sources, and m is the index of the distributed power source. It is the weighting coefficient of distributed power sources; It is the actual output power of the distributed power source; It predicts output power; implicit reward function. Defined as: ;in, and These are the power distribution system states at time t and time t+1, respectively. This refers to the action taken by the edge control unit at time t. It is the termination reward for the power distribution system to return to a safe operating condition; It is a discount factor; , , and This is the reward weighting coefficient; voltage compliance reward. Represented as: Line loss penalty Represented as: ; It is the branch resistance; the smoothness of the action is rewarded. Represented as: ; It is the threshold for changes in action; It is the reward value when the action is smooth; the reward for stable scheduling. Represented as: U is the set of scheduling variables, and k is the index of the scheduling variables. and These are the values ​​of the k-th scheduling variable at time t and time t-1, respectively. It is the maximum change of the k-th scheduling variable.

[0016] Furthermore, the historical experience review module for regulation, for each interaction experience, uses the absolute value of the voltage deviation of the node corresponding to the interaction experience as... Calculate the power quality reward factor , represented as: Where g is the interaction experience index, and each interaction experience corresponds to a power distribution interaction sample; It is the absolute value of the voltage deviation of the node corresponding to the h-th interaction experience; It is a smoothing term; it represents the rate of change of nodal load. Discretize into low, medium, and high and assign weights Calculate the load change rate factor , represented as: ;in, It is the discrete node load change rate corresponding to the h-th interaction experience; the priority weight of the initial distribution interaction sample is obtained. , represented as: After normalization, use Indicates; among which, , and It is the priority weight coefficient; It is the standard time difference error factor; the final power distribution interaction sample priority weight. Represented as: ;in, It is the number of samples. ; It is a regulatory factor; That is the standard number of samplings.

[0017] Furthermore, the system training module continuously optimizes decisions based on environmental feedback, thereby achieving the training of the power distribution system, specifically including the following:

[0018] Environment setup: Load power, energy storage output, bus voltage, and branch power flow are used as input data for the environment; simultaneously, Gaussian noise is added to the edge control unit to simulate measurement uncertainties and obtain the observed state; the data is then integrated into the system state. , represented as: ,in, pass Perform status updates; determine the action space of each edge control unit in the power distribution system, including voltage adjustment, reactive power, and energy storage charging and discharging commands, and determine the actions to be taken. Indicates; use of total reward To evaluate the merits of each action; collect historical data as power distribution interaction samples, and divide the power distribution interaction sample set into a test set and a training set for power distribution system training;

[0019] Initialization; Based on DQN, initialize DQN parameters and initialize the action selection strategy to a greedy strategy;

[0020] Training loop; based on the training set, at the beginning of each training loop, obtain the initial state of the power distribution system; according to the current state and the action selection strategy, each edge control unit selects an action from the power distribution system action space. The selected action is applied to the power distribution system environment, and the power distribution system updates its state accordingly based on the action to obtain the next state. Meanwhile, the environment calculates and returns a reward based on the reward function. ; and the current experience tuple The data is stored in the experience replay buffer. Distribution interaction samples are replayed based on their priority weights. The distribution system updates its parameters according to the current experience tuples and DQN, minimizing the loss function, L, which is expressed as: ;in, It is the loss correction factor; It is in state Lower edge control unit selects action The reward value is calculated; the priority weight of each node is updated according to the formula for dynamically allocating priority weights, which is used for subsequent action selection; a reward threshold is set to determine whether the current round has ended. If the preset time step is reached or the reward value fluctuation is less than the reward threshold, the round ends and the next round of training begins; otherwise, t+1 is added, and training continues for the next time step; a loss threshold is set. If the average loss of the power distribution system on the test set after all rounds is lower than the loss threshold, the power distribution system training is completed; otherwise, the initial parameters of DQN are adjusted and retraining is performed.

[0021] Furthermore, the intelligent control module acquires load power, energy storage output, bus voltage and branch power flow data in real time. Each node applies the action selected by the power distribution control unit to the power distribution system environment, thereby realizing power distribution control.

[0022] The beneficial effects achieved by the present invention using the above solution are as follows:

[0023] (1) In view of the problem that the control actions of general power distribution systems lack coordination and are prone to secondary disturbances when multiple nodes switch at the same time, causing voltage and power flow to oscillate near the safety limit, ignoring the safety margin requirements of the power distribution system when photovoltaic and wind power fluctuate greatly, and thus resulting in poor control stability, this solution introduces time-varying priority for each edge control unit and adds a correlation weighting term of other units to the strategy probability. It can automatically increase the control priority for hot nodes that exceed the voltage and power flow limit. When multiple units act, they take into account the state of related units, avoid disordered switching and secondary disturbances caused by disorder, and improve the overall control efficiency. On the basis of termination reward, real-time guidance reward and distributed power uncertainty reward are added. Each step receives non-zero feedback, avoiding the sparsity problem caused by pure termination reward, and simultaneously minimizing voltage deviation, line loss and action jitter. Thus, the power distribution system can still maintain safety and stability under sudden environmental changes.

[0024] (2) In view of the problem that the general power distribution system ignores the impact of voltage deviation and load change rate on power distribution safety and stability, which makes it difficult for the control strategy to quickly optimize the response to voltage over-limit or sudden load increase, this scheme introduces a power quality reward factor, focuses on voltage over-limit, and emphasizes the distribution interaction samples with more severe voltage deviation from the nominal value; introduces a load change rate factor, focuses on sudden load events, and strengthens the sampling probability of distribution interaction samples with instantaneous large load fluctuations; and introduces sampling number decay to avoid over-reliance on a small number of key distribution interaction samples, balance the repeated use of high priority experience and the continuous exploration of new distribution interaction samples; thereby improving the control effect of the power distribution system. Attached Figure Description

[0025] Figure 1 This is a flowchart illustrating an intelligent control type low-voltage intelligent power distribution control system provided by the present invention.

[0026] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0027] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0028] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0029] Example 1, see Figure 1 The present invention provides an intelligent control type low-voltage intelligent power distribution control system, including a load condition analysis module, a measurement uncertainty assessment module, a control strategy planning module, a control reward mechanism construction module, a control historical experience review module, a system training module, and an intelligent control module;

[0030] The load condition analysis module collects load and energy storage output to obtain the net power of the low-voltage distribution network; and sends the data to the measurement uncertainty assessment module.

[0031] The measurement uncertainty assessment module collects bus voltage and branch power flow data, injects Gaussian noise to obtain the final observation state, and sends the data to the control strategy planning module.

[0032] The control strategy planning module defines an action set for each edge control unit and allocates scheduling priorities based on voltage deviation and node correlation; and sends the data to the control reward mechanism construction module.

[0033] In addition to the safety termination reward, the regulation reward mechanism construction module constructs a real-time feedback signal through implicit reward function differential guidance reward and distributed power source uncertainty reward; and sends the data to the regulation historical experience review module.

[0034] The historical experience review module for regulation calculates the priority of distribution interaction samples based on factors such as voltage deviation, load change rate, and time difference error, and dynamically adjusts the sampling weights in conjunction with the attenuation mechanism; and sends the data to the system training module.

[0035] The system training module, based on historical data and the DQN framework, samples and updates the Q-network parameters from the experience replay buffer to train the complete power distribution system; and sends the data to the intelligent control module.

[0036] The intelligent control module acquires real-time data on load, energy storage, bus voltage, and branch power flow, and performs intelligent control based on the selected control actions.

[0037] Example 2, see Figure 1 This embodiment is based on the above embodiment, and the load condition analysis module collects load power based on edge-side IoT sensors and AMI. and energy storage output State updates are performed on the edge gateway using a time-series model; load centers are treated as scalable power regions to characterize the dynamic power distribution characteristics in the low-voltage distribution network, providing time-varying state input for intelligent dispatch, as expressed as: ; ; ;in, and These are load and distributed generation rate, respectively; compared with traditional static load forecasting, it can reflect short-term load / output fluctuations; among them, It is the change in load power; It is the change in energy storage output; It is a time interval; and It is the rate of change; and These are the net power of the low-voltage distribution network at time t+1 and time t, respectively.

[0038] Example 3, see Figure 1 This embodiment is based on the above embodiment, and the measurement uncertainty assessment module uses the distribution automation SCADA / PMU to collect the bus voltage. Branch Road Trend Gaussian noise is added to the edge control unit to simulate measurement errors and communication packet loss; this closely resembles the measurement errors introduced by voltage sensors and communication links in real-world applications, expressed as: Even in noisy environments, the voltage / power can still be kept within a safe range; among which, This represents the original measurement state of the i-th node; This is the final observed state; The mean is 0 and the variance is Gaussian noise.

[0039] Example 4, see Figure 1 This embodiment, based on the above embodiment, deploys an edge control unit at each node of the power distribution system. The actions of each edge control unit include: adjusting voltage reactive power, including switching OLTC and Static Var Compensator (SVC); initiating energy storage charging and discharging commands; and dynamically allocating priority weights. For critical fault nodes and voltage over-limit hotspots, higher decision-making priority is given to avoid secondary disturbances caused by simultaneous actions of multiple nodes. An action coordination mechanism is introduced to enable multiple edge control units to coordinate their actions, improving the overall control efficiency of the power distribution system. The action strategy is expressed as follows: ; ;in, It is in state The probability that the i-th edge control unit selects action a; It is the Pearson correlation coefficient between the i-th and j-th edge control units; It is the reward function value of the i-th edge control unit selecting action a in state s; It is a priority adjustment factor; It is the rated bus voltage; to avoid oscillation caused by disordered switching, priority should be given to solving voltage / power flow exceeding the limit area.

[0040] Example 5, see Figure 1 This embodiment, based on the above embodiments, introduces a real-time guiding reward in addition to the termination reward when the distribution system recovers to a safe operating condition, accelerating strategy learning. While ensuring the optimal control strategy remains unchanged, it encourages evolution towards voltage compliance and power loss minimization at each step through implicit reward function differentiation, and additionally introduces a reward for distributed source uncertainty. This allows the control strategy to maintain the safe and stable operation of the distribution system even when the output of distributed sources is unstable, improving the distribution system's ability to accept renewable energy. The total reward... Represented as: Real-time guidance and rewards Represented as: Uncertainty-based rewards for distributed power sources Represented as: M is the set of distributed power sources, and m is the index of the distributed power source. It is the weighting coefficient of distributed power sources; It is the actual output power of the distributed power source; It predicts output power, achieved through a Long Short-Term Memory network trained on historical data; implicit reward function. Defined as: ;in, and These are the power distribution system states at time t and time t+1, respectively. This refers to the action taken by the edge control unit at time t. It is the termination reward for the power distribution system to return to a safe operating condition; It is a discount factor; , , and This is the reward weighting coefficient; voltage compliance reward. Represented as: Line loss penalty Represented as: ; It is the branch resistance; the smoothness of the action is rewarded. Represented as: ; It is the threshold for changes in action; It is the reward value when the action is smooth; the reward for stable scheduling. Represented as: U is the set of scheduling variables, and k is the index of the scheduling variables. and These are the values ​​of the k-th scheduling variable at time t and time t-1, respectively. It is the maximum change of the k-th scheduling variable; each step has non-zero feedback to avoid sparsity problems caused by pure termination rewards, while optimizing voltage deviation, loss and smoothness.

[0041] By performing the above operations, this solution addresses the common problems in general power distribution systems, such as a lack of coordinated control actions, secondary disturbances arising from simultaneous switching of multiple nodes, voltage and power flow oscillations near safety limits, and neglect of the safety margin requirements of power distribution systems during significant fluctuations in photovoltaic and wind power, leading to poor control stability. This solution introduces time-varying priorities to each edge control unit and adds a correlation weighting term for other units to the strategy probability. This automatically increases the control priority for hotspot nodes where voltage and power flow exceed limits. When multiple units act, they consider the states of related units, avoiding secondary disturbances caused by disordered switching and improving overall control efficiency. Furthermore, it adds real-time guidance rewards and distributed power source uncertainty rewards to the termination reward, ensuring non-zero feedback at each step, avoiding sparsity problems caused by pure termination rewards, and simultaneously minimizing voltage deviation, line losses, and action jitter. This allows the power distribution system to maintain safety and stability even under sudden environmental changes.

[0042] Example 6, see Figure 1 This embodiment is based on the above embodiment. The historical experience review module adjusts the absolute value of the voltage deviation of the node corresponding to each interaction experience for each interaction experience. Calculate the power quality reward factor , represented as: Where g is the interaction experience index, and each interaction experience corresponds to a power distribution interaction sample; It is the absolute value of the voltage deviation of the node corresponding to the h-th interaction experience; It is a smoothing term; it represents the rate of change of nodal load. Discretize into low, medium, and high and assign weights Calculate the load change rate factor , represented as: ;in, It is the discrete node load change rate corresponding to the h-th interaction experience; the priority weight of the initial distribution interaction sample is obtained. , represented as: After normalization, use Indicates; among which, , and It is the priority weight coefficient; This is the standard time difference error factor; to avoid multiple dependencies of the power distribution system on power distribution interaction samples, the priority weight of the power distribution interaction samples will decrease as the number of samplings increases, ultimately resulting in a lower priority weight for the power distribution interaction samples. Represented as: ;in, It is the number of samples. ; It is a regulatory factor; That is the standard number of samplings.

[0043] By performing the above operations, this solution addresses the problem in general power distribution systems where the impact of voltage deviation and load change rate on distribution safety and stability is ignored, making it difficult for control strategies to quickly optimize responses to voltage overruns or sudden load increases. This solution introduces a power quality reward factor, focusing on voltage overruns and emphasizing distribution interaction samples with more severe voltage deviations from nominal values; it also introduces a load change rate factor, focusing on sudden load events and strengthening the sampling probability of distribution interaction samples with instantaneous large load fluctuations; and it introduces a sampling count decay to avoid over-reliance on a small number of key distribution interaction samples, balancing the reuse of high-priority experience with the continuous exploration of new distribution interaction samples; thereby improving the control effect of the power distribution system.

[0044] Example 7, see Figure 1 This embodiment is based on the above embodiment. The system training module continuously optimizes decisions based on environmental feedback, thereby realizing the training of the power distribution system. Specifically, it includes the following:

[0045] Environment setup: Load power and energy storage output collected by edge-side IoT sensors and AMIs, as well as bus voltage and branch power flow collected by distribution automation SCADA / PMUs, are used as input data for the environment. Simultaneously, Gaussian noise is added to the edge control unit to simulate measurement uncertainties and obtain the observed state. The data is then integrated into the system state. , represented as: ,in, pass Perform status updates; determine the action space of each edge control unit in the power distribution system, including voltage adjustment, reactive power, and energy storage charging and discharging commands, and determine the actions to be taken. Indicates; use of total reward To evaluate the merits of each action and motivate the power distribution system to learn toward the desired goal; collect historical data as power distribution interaction samples, and divide the power distribution interaction sample set into a test set and a training set for power distribution system training;

[0046] Initialization; Based on DQN, initialize DQN parameters and initialize the action selection strategy to a greedy strategy;

[0047] Training loop; based on the training set, at the beginning of each training loop, obtain the initial state of the power distribution system; according to the current state and the action selection strategy, each edge control unit selects an action from the power distribution system action space. The selected action is applied to the power distribution system environment, and the power distribution system updates its state accordingly based on the action to obtain the next state. Meanwhile, the environment calculates and returns a reward based on the reward function. ; and the current experience tuple The data is stored in the experience replay buffer. Distribution interaction samples are replayed based on their priority weights. The distribution system updates its parameters according to the current experience tuples and DQN, minimizing the loss function, L, which is expressed as: ;in, It is the loss correction factor; It is in state Lower edge control unit selects action The reward value is calculated; the priority weight of each node is updated according to the formula for dynamically allocating priority weights, which is used for subsequent action selection; a reward threshold is set to determine whether the current round has ended. If the preset time step is reached or the reward value fluctuation is less than the reward threshold, the round ends and the next round of training begins; otherwise, t+1 is added, and training continues for the next time step; a loss threshold is set. If the average loss of the power distribution system on the test set after all rounds is lower than the loss threshold, the power distribution system training is completed; otherwise, the initial parameters of DQN are adjusted and retraining is performed.

[0048] Example 8, see Figure 1 This embodiment is based on the above embodiment. The intelligent control module acquires load power, energy storage output, bus voltage and branch power flow data in real time. Each node applies the action selected by the power distribution control unit to the power distribution system environment, thereby realizing power distribution control.

[0049] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0050] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

[0051] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A smart control type low-voltage intelligent power distribution control system, characterized in that: The system comprises a load condition analysis module, a measurement uncertainty evaluation module, a regulation strategy planning module, a regulation reward mechanism construction module, a regulation historical experience review module, a system training module and an intelligent regulation module. The load condition analysis module collects load and energy storage output to obtain net power of the low-voltage distribution network. The measurement uncertainty evaluation module collects bus voltage and branch power flow data and injects Gaussian noise to obtain the final observation state. The regulation strategy planning module defines an action set for each edge control unit and assigns a scheduling priority based on voltage deviation and node correlation. The regulation reward mechanism construction module constructs real-time feedback signals through implicit reward function difference guidance reward and distributed power uncertainty reward in addition to safe termination reward. The regulation historical experience review module calculates distribution interaction sample priority according to voltage deviation, load change rate and time difference error factor, and dynamically adjusts the sampling weight in combination with the decay mechanism. The system training module samples and updates Q network parameters from the experience replay buffer based on historical data and the DQN framework, thereby completing the training of the distribution system. The intelligent regulation module obtains load, energy storage, bus voltage and branch power flow observation data in real time and performs intelligent regulation based on regulation action selection. The load condition analysis module collects load power and energy storage output ; the load center is regarded as a scalable power area, and a time-varying state input is provided for intelligent scheduling, which is represented as: ; ; ; wherein, and are the load and distributed power generation rates respectively; wherein, is the change amount of load power; is the change amount of energy storage output; is the time interval; and are the change rates; and are the net powers of the low-voltage power distribution network at t+1 and t respectively; The regulation strategy planning module is to deploy an edge control unit for each node of the power distribution system, and the action execution of each edge control unit includes: adjusting voltage and reactive power; initiating energy storage charging and discharging instructions; dynamically allocating priority weights , and introducing an action coordination mechanism, and the action strategy is represented as: ; ; wherein, is the probability of the i-th edge control unit selecting action a in the state ; is the Pearson correlation coefficient of the i-th and j-th edge control units; is the reward function value of the i-th edge control unit selecting action a in the state s; is a priority adjustment factor; is a rated bus voltage.

2. The intelligent regulation type low-voltage intelligent power distribution control system according to claim 1, characterized in that: The measurement uncertainty evaluation module collects bus voltage , branch power flow , and adds Gaussian noise, expressed as: ; wherein, is the original measurement state of the i-th node; is the final observation state; is the Gaussian noise with mean 0 and variance .

3. The intelligent regulation type low-voltage intelligent power distribution control system according to claim 2, characterized in that: The regulation reward mechanism construction module introduces real-time guidance reward and additional distributed power uncertainty reward in addition to the termination reward of the power distribution system restored to the safe working condition; the total reward is expressed as: ; real-time guidance reward is expressed as: ; distributed power uncertainty reward is expressed as: ; M is a set of distributed power, and m is a distributed power index; is a distributed power weight coefficient; is the actual output power of the distributed power; is the predicted output power; implicit reward function is defined as: ; wherein and are the power distribution system states at time t and t+1, respectively; is the action taken by the edge control unit at time t; is the termination reward of the power distribution system restored to the safe working condition; is a discount factor; , , and are reward weight coefficients; voltage compliance reward is expressed as: ; line loss penalty is expressed as: ; is the branch resistance; action smoothness reward is expressed as: ; is the threshold value of action change; is the reward value when the action is smooth; scheduling stability reward is expressed as: ; U is a set of scheduling variables, and k is a scheduling variable index; and are the values of the kth scheduling variable at time t and t-1, respectively; is the maximum change amount of the kth scheduling variable.

4. The intelligent regulation type low-voltage intelligent power distribution control system according to claim 3, characterized in that: The regulation historical experience review module is for each interaction experience, with the voltage deviation absolute value of the interaction experience corresponding node as , calculates the power quality reward factor , expressed as: ; wherein g is the interaction experience index, each interaction experience corresponds to a power distribution interaction sample; is the voltage deviation absolute value of the hth interaction experience corresponding node; is a smoothing term; the node load change rate is discretized into low, medium and high and weighted , calculates the load change rate factor , expressed as: ; wherein is the discretized node load change rate of the hth interaction experience corresponding node; the initial power distribution interaction sample priority weight , expressed as: , after normalization, is expressed by ; wherein , and are priority weight coefficients; is the standard time difference error factor; the final power distribution interaction sample priority weight is expressed as: ; wherein is the sampling number, ; is the adjustment factor; is the standard sampling number.

5. The intelligent regulation type low-voltage intelligent power distribution control system according to claim 4, characterized in that: The system training module continuously optimizes decisions based on environmental feedback, thereby realizing the training of the distribution system, and specifically includes the following contents: environment building; load power and energy storage output, as well as bus voltage and branch power flow, are used as input data of the environment; the action space of each edge control unit of the distribution system is determined; the total reward is used to evaluate the pros and cons of each action; historical data is collected as distribution interaction samples, and the distribution interaction sample set is divided into a test set and a training set for distribution system training; initialization; based on DQN, the DQN parameters are initialized, and the action selection strategy is initialized as a greedy strategy; training loop; based on the training set, at the beginning of each training round, the initial state of the distribution system is obtained; according to the current state and the action selection strategy, the distribution system updates the state according to the action to obtain the next state, while the environment calculates and returns the reward according to the reward function; and the current experience tuple is stored in the experience replay buffer, and the distribution interaction samples are replayed based on the distribution interaction sample priority weight; the distribution system updates the parameters according to the current experience tuple and DQN to minimize the loss function; and the priority weight of each node is updated according to the formula of dynamically allocated priority weight; set the reward threshold, determine whether the current round is over, if the preset time step is reached or the reward value fluctuation is less than the reward threshold, the round is over, and the next round of training begins; otherwise, t+1, continue training at the next time step; set the loss threshold, if the average loss of the distribution system on the test set after all rounds are completed is lower than the loss threshold, the distribution system training is completed, otherwise, adjust the DQN initial parameters and retrain.

6. The intelligent regulation type low-voltage intelligent power distribution control system according to claim 5, characterized in that: The intelligent regulation module acquires load power, energy storage output, bus voltage and branch power flow data in real time, each node applies the action selected by the power distribution control unit to the power distribution system environment, thereby realizing power distribution control.

Citation Information

Patent Citations

  • Power distribution network voltage reactive power control method and system based on safety reinforcement learning algorithm

    CN116760047A

  • Multi-time scale voltage regulation method and system based on deep reinforcement learning

    CN119209561A