A multi-objective electrical automation collaborative control method integrating reinforcement learning
By integrating a multi-objective electrical automation collaborative control method with reinforcement learning, the problem that traditional control methods are difficult to balance multiple objectives in complex environments is solved, collaborative control between distribution cabinets is achieved, and the stability and efficiency of the power system are improved.
Patent Information
- Application Number
- CN202510293911.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Traditional electrical automation control methods have difficulty in effectively balancing and achieving multiple conflicting objectives in a complex and changing operating environment, and have poor adaptability to system parameter changes and uncertainties.
A multi-objective electrical automation collaborative control method integrating reinforcement learning is adopted. By collecting the electrical parameters and operating status of the distribution cabinets in real time, a state vector is constructed. Combined with the harmonic and line loss strategies, a hierarchical reward mechanism is set up to achieve collaborative control between distribution cabinets.
It achieves effective coordination between distribution cabinets in the power system, improves operational stability and reliability, reduces line losses, improves load balancing and equipment operating efficiency, and ensures efficient operation of the system.
Smart Images

Figure CN120103766B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electrical automation control technology, and in particular to a multi-objective electrical automation collaborative control method integrating reinforcement learning. Background Art
[0002] With the rapid development of electrical automation technology, the scale and complexity of electrical automation systems are constantly increasing. In actual operation, electrical automation systems must simultaneously meet multiple objectives, such as efficient energy utilization, stable power supply quality, low equipment loss, and good economic benefits. Traditional control methods are typically designed for a single objective, making it difficult to effectively balance and achieve multiple conflicting objectives in complex and changing operating environments. Furthermore, traditional control methods often rely on precise system models, making them less adaptable to system parameter changes, uncertainties, and real-time dynamic adjustments.
[0003] Chinese patent publication number CN111200285B discloses a hybrid coordinated control method for microgrids based on reinforcement learning and multi-agent theory. The method includes: designing a voltage-stratified transition voltage layer control strategy and a dual energy storage role-dividing control strategy. When the energy storage unit operates in voltage regulation mode, the two energy storage units operate independently; when the energy storage unit is required to continuously absorb or supplement power, the two energy storage units switch to coordinated charging / discharging. The method also includes constructing a Q-Learning-based action space and state space. The method also includes designing a multi-agent reinforcement learning control framework, including designing basic update rules for state-action pairs and selecting corresponding value functions. The method also includes designing a basic action selection mechanism and reward value strategy, including designing the selection strategy adopted by the system in the initial state and the reward value in each state. The method also includes designing a reinforcement learning algorithm flow, based on the above strategies, to implement the control strategy.
[0004] It can be seen from this that the technical solution has the following problems: although there is a dual energy storage role-based control strategy, in the entire microgrid system, the dynamic adjustment mechanism for other equipment and the overall system may be relatively concentrated in aspects such as the energy storage working mode conversion, which is not comprehensive and flexible enough, and cannot respond to various changes and uncertainties in the system in a timely and accurate manner. Summary of the Invention
[0005] To this end, the present invention provides a multi-objective electrical automation collaborative control method integrating reinforcement learning to overcome the problem in the prior art that electrical systems cannot achieve device-level and system-level collaborative optimization control.
[0006] To achieve the above objectives, the present invention provides a multi-objective electrical automation collaborative control method integrating reinforcement learning, comprising:
[0007] Step S1: Real-time collection of electrical parameters and operating status of each distribution cabinet in the area to construct its own state vector, and construct neighboring state vectors based on the power transmission direction and line impedance between adjacent distribution cabinets. The distribution cabinet types include power access cabinets, load distribution cabinets, and energy storage control cabinets.
[0008] Step S2: constructing a preset target strategy according to the output current, output voltage, and line loss of each distribution cabinet, and determining the action vector of each distribution cabinet based on its own state vector, neighboring state vectors, and the preset target strategy;
[0009] Step S3, determining the impact direction of the action vector on each distribution cabinet after execution based on the action vector, the own target reward function and the collaborative target reward function, wherein the impact direction includes the individual impact direction and the total impact direction;
[0010] The self-target reward function is determined according to the type of distribution cabinet, including: the self-reward function of the power access cabinet includes a voltage deviation penalty term and a reactive output constraint term; the self-reward function of the load distribution cabinet includes a load balancing term and a tripping number penalty term; the self-reward function of the energy storage control cabinet includes an SOC health term and a charge and discharge efficiency term;
[0011] The collaborative objective reward function is determined based on the total line loss, load balancing rate and communication efficiency between the distribution cabinets;
[0012] Step S4: determining an action adjustment method according to the impact direction.
[0013] Specifically, in step S2, a harmonic target strategy is determined according to the output voltage and the output current, a line loss strategy is determined according to the line loss, and a preset target strategy is determined according to the harmonic target strategy and the line loss strategy.
[0014] Specifically, in step S2, constructing a harmonic target strategy includes:
[0015] Determine the current harmonic content and voltage harmonic content based on the output current and output voltage of each distribution cabinet;
[0016] determining a current influence coefficient based on the current harmonic content and the current harmonic threshold;
[0017] determining a voltage influence coefficient based on the voltage harmonic content and the voltage harmonic threshold;
[0018] The harmonic target strategy is determined based on the current influence coefficient and the voltage influence coefficient.
[0019] Specifically, in step S2, the action vector corresponding to the harmonic target strategy includes:
[0020] The power access cabinet injects reverse harmonics to offset the load harmonics;
[0021] Disconnect the branch switch of the load distribution cabinet with high harmonic load or adjust the three-phase load distribution;
[0022] The energy storage control cabinet offsets harmonics by charging / discharging energy storage.
[0023] Specifically, in step S2, constructing a line loss strategy includes:
[0024] Determine the actual resistance of the line between the distribution cabinets based on the initial resistance of the line, the temperature influence coefficient, and the aging influence coefficient;
[0025] The total line loss is determined according to the actual line resistance and the line current.
[0026] Specifically, in step S2, the temperature influence coefficient is determined based on the historical ambient temperature and the historical line resistance, and the aging influence coefficient is determined based on the line usage time and the historical line resistance.
[0027] Specifically, in step S2, the action vector corresponding to the line loss strategy includes:
[0028] Adjust the reactive output of the power access cabinet to reduce line current;
[0029] Switch the load distribution cabinet switch combination to evenly distribute the load;
[0030] Adjust the charging / discharging power of the energy storage control cabinet to optimize the power flow distribution.
[0031] Specifically, step S3 includes:
[0032] collecting the output current of each load distribution cabinet in the area in real time, determining the average load of the load distribution cabinet based on the output current, and determining the load balancing rate based on the average load;
[0033] Collecting data transmission and data reception of each power distribution cabinet in the area in real time, and determining communication efficiency based on the data transmission and data reception;
[0034] The collaborative objective reward function is determined by weighting according to the load balancing rate and the communication efficiency.
[0035] Specifically, in step S4, the action adjustment method is determined, wherein:
[0036] If the total impact direction is a positive reward, then determining the action adjustment amplitude according to the individual impact direction;
[0037] If the total impact direction is a negative reward, repeat steps S2 to S4 until the total impact direction is a positive reward or the number of iterations is equal to the preset maximum number of iterations.
[0038] Specifically, in step S4, determining the action adjustment range includes:
[0039] If the total impact direction is positive reward and the individual impact direction corresponding to the distribution cabinet is positive incentive, then the action adjustment amplitude of the action vector of the corresponding distribution cabinet is increased;
[0040] If the total impact direction is a positive reward and the individual impact direction corresponding to the distribution cabinet is a negative reward, the action adjustment amplitude of the action vector of the corresponding distribution cabinet is reduced.
[0041] Compared with the prior art, the beneficial effect of the present invention is that the present invention integrates multiple goals such as voltage stability, harmonic control, load balancing, etc. through a reinforcement learning framework to achieve effective collaboration between distribution cabinets in the power system, including using its own state vector and neighboring state vectors to jointly represent the state of the device itself and reflect the network topology relationship, thereby achieving strong coupling of the power system. Constructing preset target strategies and determining action vectors based on multiple factors is more in line with actual needs. At the same time, a hierarchical reward mechanism is set up to distinguish between individual target rewards and collaborative target rewards, effectively evaluate the impact of action vectors on the system, and optimize device-level and system-level control strategies, thereby achieving multi-target collaborative control, improving the operational stability and reliability of the electrical automation system, reducing line losses, improving load balancing, and ensuring efficient system operation.
[0042] Furthermore, the present invention takes harmonic factors into consideration and can perform targeted analysis on the current harmonics and voltage harmonics of the distribution cabinets in the area, determine the current influence coefficient and the voltage influence coefficient based on the analysis results, and thus provide a precise control strategy for harmonics for multi-objective electrical automation collaborative control, facilitate timely discovery and processing of harmonic problems, reduce the risk of failures caused by harmonics, achieve collaborative work between distribution cabinets, improve the operating efficiency and life of equipment, and ensure the stable operation of the electrical system.
[0043] Furthermore, the present invention considers factors such as temperature and aging, determining the actual resistance of each line between distribution cabinets based on line current, temperature influence coefficient, and aging influence coefficient. This determines the total line loss and constructs a line loss strategy, enabling targeted control of line losses. While ensuring safe line operation, this effectively reduces line losses, providing a solid foundation for achieving pre-set target strategies, improving the operating efficiency of the power system, and enhancing the reliability and stability of the entire power system.
[0044] Furthermore, the collaborative reward objective function of the present invention quantitatively weights the load balancing rate and communication efficiency, providing a clear and quantitative evaluation standard for the overall operating status of the power system, which makes it easier for managers to intuitively understand the system performance and facilitates the subsequent value of the collaborative objective reward function to quickly determine whether the system is operating in the optimal state, providing a clear direction for the optimization and adjustment of the system, improving the operating efficiency of the power system, and enhancing the reliability and stability of the entire power system.
[0045] Furthermore, the present invention determines corresponding adjustment methods based on the total impact direction and the individual impact direction, including determining the action adjustment amplitude under the total impact direction as a positive incentive, which can promote the overall performance of the system while taking into account the operating status of individual distribution cabinets. Alternatively, repeated iterations can be performed under the total impact direction as a negative incentive, prompting the system to find the optimal strategy in a complex and changing operating environment, improve its own robustness and adaptability, achieve continuous performance improvement, improve the accuracy and effectiveness of control, realize multi-objective collaborative control, and ensure the efficient operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a schematic diagram of an algorithm for a multi-objective electrical automation collaborative control method integrating reinforcement learning according to an embodiment of the present invention;
[0047] Figure 2 A step diagram of a multi-objective electrical automation collaborative control method integrating reinforcement learning according to an embodiment of the present invention;
[0048] Figure 3 A diagram showing the steps for constructing a harmonic target strategy according to an embodiment of the present invention;
[0049] Figure 4 This is a decision diagram for determining an action adjustment method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0050] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0051] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0052] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0053] See also Figure 1 、 Figure 2 As shown, Figure 1 This is a schematic diagram of an algorithm for a multi-objective electrical automation collaborative control method integrating reinforcement learning according to an embodiment of the present invention. Figure 2 This is a step diagram of a multi-objective electrical automation collaborative control method integrating reinforcement learning according to an embodiment of the present invention.
[0054] Specifically, the present invention provides a multi-objective electrical automation collaborative control method integrating reinforcement learning, comprising:
[0055] Step S1: Real-time collection of electrical parameters and operating status of each distribution cabinet in the area to construct its own state vector, and construct neighboring state vectors based on the power transmission direction and line impedance between adjacent distribution cabinets. The distribution cabinet types include power access cabinets, load distribution cabinets, and energy storage control cabinets.
[0056] Step S2: constructing a preset target strategy according to the output current, output voltage, and line loss of each distribution cabinet, and determining the action vector of each distribution cabinet based on its own state vector, neighboring state vectors, and the preset target strategy;
[0057] Step S3, determining the impact direction of the action vector on each distribution cabinet after execution based on the action vector, the own target reward function and the collaborative target reward function, wherein the impact direction includes the individual impact direction and the total impact direction;
[0058] The self-target reward function is determined according to the type of distribution cabinet, including: the self-reward function of the power access cabinet includes a voltage deviation penalty term and a reactive output constraint term; the self-reward function of the load distribution cabinet includes a load balancing term and a tripping number penalty term; the self-reward function of the energy storage control cabinet includes an SOC health term and a charge and discharge efficiency term;
[0059] The collaborative objective reward function is determined based on the total line loss, load balancing rate and communication efficiency between the distribution cabinets;
[0060] Step S4: determining an action adjustment method according to the impact direction.
[0061] It can be understood that each distribution cabinet within the area is an intelligent agent in reinforcement learning. Step S1 is state perception. This monitors the distribution cabinet's equipment status through its own state vector and neighboring state vectors, reflecting the network topology within the distribution system. The neighboring state vectors accurately describe the dynamics of the distribution network based on power transmission direction and line impedance. Step S2 is policy decision-making. A preset target policy is constructed based on the distribution cabinet's output current, voltage, and line loss. The action vector for each distribution cabinet is determined by combining its own and neighboring state vectors. The action vector determines the operation to be performed by the distribution cabinet. Based on the comparison results, the distribution cabinet is triggered to execute the corresponding action vector, thereby determining whether the current action meets system requirements. Step S3 is reward evaluation. The reward is set for the individual target and the collaborative target. This dual-layer reward determines the impact of the action vector on each distribution cabinet (both individual and overall) after execution. Step S4 is policy optimization. The action adjustment method is determined based on the impact direction, thereby optimizing the distribution cabinet control strategy.
[0062] It's understandable that the preset target strategy is a preliminary decision-making scheme based on rules and experience, providing initial action guidance for the PDC. Its own target reward function and the collaborative target reward function are used to evaluate the results of this behavior. Different reward values are primarily used to provide feedback on whether the system state is moving in the desired direction. If the reward is low, the preset target strategy may need to be adjusted.
[0063] In a specific embodiment, the acquisition frequency of parameters such as electrical parameters, operating status, power transmission direction, and line impedance of each distribution cabinet ranges from 20 times / min to 30 times / min. Preferably, the acquisition frequency is 25 times / min. The acquisition frequency can also be determined based on the rate of change of electrical parameters, and the acquisition frequency is positively correlated with the rate of change of electrical parameters. The self-reward function of the power access cabinet is the weighted sum of the voltage deviation penalty term and the reactive output constraint term of the power access cabinet; the self-reward function of the load distribution cabinet is the weighted sum of the load balancing term and the tripping number penalty term; the self-reward function of the energy storage control cabinet is the weighted sum of the SOC health term and the charge and discharge efficiency term. In implementation, the acquisition frequency value range and preferred value can be determined according to actual conditions, and the self-reward functions corresponding to the power access cabinet, load distribution cabinet, and energy storage control cabinet can also be determined according to actual conditions. No specific limitation is made here and no further details are given.
[0064] The present invention integrates multiple goals such as voltage stability, harmonic control, and load balancing through a reinforcement learning framework to achieve effective collaboration between distribution cabinets in the power system, including using its own state vector and neighboring state vectors to jointly represent the state of the device itself and reflect the network topology relationship, thereby achieving strong coupling of the power system. Constructing preset target strategies and determining action vectors based on multiple factors is more in line with actual needs. At the same time, a hierarchical reward mechanism is set up to distinguish between individual target rewards and collaborative target rewards, effectively evaluate the impact of action vectors on the system, and optimize device-level and system-level control strategies, thereby achieving multi-target collaborative control, improving the operational stability and reliability of the electrical automation system, reducing line losses, improving load balancing, and ensuring efficient system operation.
[0065] Specifically, in step S2, a harmonic target strategy is determined according to the output voltage and the output current, a line loss strategy is determined according to the line loss, and a preset target strategy is determined according to the harmonic target strategy and the line loss strategy.
[0066] It's understandable that using harmonic targets and line loss as preset target strategies can guide the intelligent agent (i.e., the control strategy of each distribution panel) to learn and adjust towards reducing harmonics and line losses. Harmonics can cause equipment overheating, insulation aging, malfunctions, and other problems. They can also increase line losses, voltage fluctuations, and flicker, impacting the stable operation of the entire electrical automation system. The harmonic target strategy helps control harmonic content within a reasonable range, reducing the adverse effects of harmonics on distribution panels and other electrical equipment in the system. The line loss strategy can optimize the operating status and power distribution of each distribution panel, reduce energy losses in the lines, and improve the efficiency of power transmission and utilization.
[0067] See also Figure 3 As shown in FIG, it is a step diagram of constructing a harmonic target strategy according to an embodiment of the present invention. Specifically, in step S2, constructing a harmonic target strategy includes:
[0068] Step S211, determining the current harmonic content and the voltage harmonic content based on the output current and output voltage of each distribution cabinet;
[0069] Step S212, determining a current influence coefficient based on the current harmonic content and the current harmonic threshold;
[0070] Step S213, determining a voltage influence coefficient based on the voltage harmonic content and the voltage harmonic threshold;
[0071] Step S214: determining a harmonic target strategy based on the current influence coefficient and the voltage influence coefficient.
[0072] It's understandable that in electrical systems, both current and voltage harmonics affect equipment operation and power quality. Determining the harmonic content from both output current and voltage perspectives comprehensively reflects the harmonic situation within the distribution panel. Comparing the harmonic content with corresponding thresholds to determine the impact coefficient, and then constructing a harmonic target strategy based on the impact coefficient, quantifies the impact of harmonics on the system, providing a quantitative basis for subsequent control decisions and enabling more precise control.
[0073] In a specific embodiment, a fast Fourier transform is performed on the output voltage and output current of the distribution cabinet to convert the time domain signals into the frequency domain. In the frequency domain, the current harmonic content and voltage harmonic content are determined based on the amplitude and phase information of each frequency component. The current harmonic threshold ranges from 5% to 8%, preferably 5.5%. The voltage harmonic threshold ranges from 3% to 5%, preferably 4%. The current influence coefficient = |current harmonic threshold - current harmonic content| / current harmonic threshold; the closer the current harmonic content is to the current harmonic threshold, the smaller the current influence coefficient; the further the current harmonic threshold is exceeded, the larger the current influence coefficient. The voltage influence coefficient = |voltage harmonic threshold - voltage harmonic content| / voltage harmonic threshold. In implementation, the current influence coefficient and voltage influence coefficient can also be determined using weighted calculation methods, electrical system simulation models, and other methods. The method for determining the current influence coefficient and voltage influence coefficient can be adjusted according to actual conditions and is not specifically limited here and will not be further described.
[0074] In a specific embodiment, the harmonic target strategy is determined based on the current influence coefficient and the voltage influence coefficient, and the harmonic target strategy is characterized based on the harmonic comprehensive index. Specifically, the harmonic comprehensive index = current weight × current influence coefficient + voltage weight × voltage influence coefficient. The sum of the current weight and the voltage weight is 1, and the voltage weight is greater than the current weight. Preferably, the voltage weight takes a value of 0.7, and the current weight takes a value of 0.3. In implementation, the value range and preferred value of the voltage weight and the current weight can be determined according to actual conditions, and are not specifically limited here and will not be repeated.
[0075] Specifically, in step S2, the action vector corresponding to the harmonic target strategy includes:
[0076] The power access cabinet injects reverse harmonics to offset the load harmonics;
[0077] Disconnect the branch switch of the load distribution cabinet with high harmonic load or adjust the three-phase load distribution;
[0078] The energy storage control cabinet offsets harmonics by charging / discharging energy storage.
[0079] It's understandable that nonlinear loads generate a large number of harmonics, which in turn distort current and voltage waveforms and affect power quality. Based on the principle of harmonic cancellation, the power access panel injects harmonics of equal magnitude and opposite direction to the load harmonics, causing them to cancel each other out. This reduces the total harmonic content in the grid and makes the current and voltage waveforms closer to sine waves.
[0080] It's understandable that high-harmonic loads are a major source of harmonics in power systems. Disconnecting the corresponding branch switches directly eliminates these harmonic sources and reduces harmonic content. Three-phase load imbalance can also cause harmonics. Unbalanced loads cause currents in each phase to differ in magnitude and phase, leading to current waveform distortion. Adjusting the three-phase load distribution balances the three-phase loads and reduces the generation of harmonics caused by load imbalance.
[0081] It's understood that the energy storage control cabinet can regulate power output and input during the charging and discharging process. When harmonics are present in the grid, the energy storage control cabinet can control the charging and discharging of the energy storage device based on the harmonics, generating a power component opposite to the harmonics, thereby partially offsetting the harmonics.
[0082] During implementation, active power filters or multi-function inverters are required to achieve dynamic compensation capabilities.
[0083] The present invention takes harmonic factors into consideration and can perform targeted analysis on the current harmonics and voltage harmonics of the power distribution cabinets in the area, determine the current influence coefficient and the voltage influence coefficient based on the analysis results, and thus provide a precise control strategy for harmonics for multi-objective electrical automation collaborative control, facilitate timely discovery and processing of harmonic problems, reduce the risk of failures caused by harmonics, achieve collaborative work between the power distribution cabinets, improve the operating efficiency and life of the equipment, and ensure the stable operation of the electrical system.
[0084] Specifically, in step S2, the temperature influence coefficient is determined based on the historical ambient temperature and the historical line resistance, and the aging influence coefficient is determined based on the line usage time and the historical line resistance.
[0085] It's understandable that electrical system wiring gradually ages over time, deteriorating insulation performance and making leakage more likely. Furthermore, prolonged oxidation increases the resistivity of wiring, leading to increased resistance and, consequently, increased line losses. Furthermore, line resistance generally increases with temperature, with temperature positively correlated to line resistance, and line losses are directly proportional to resistance. Therefore, when determining a line loss strategy, consider both line age and ambient temperature.
[0086] In a specific embodiment, based on the physical properties of the metal conductor, the ambient temperature and the line resistance are in a linear relationship, and the temperature influence coefficient is determined based on a linear regression model. The linear regression model is: Line resistance baseline value = historical line resistance × [1 + temperature influence coefficient × (historical ambient temperature - temperature baseline value)], wherein the line resistance baseline value is the resistance baseline value corresponding to the temperature baseline value, and the temperature baseline value is 25°C. Assuming that the line resistance increases exponentially with time (oxidation, fatigue factors), an exponential decay model is constructed: Historical line resistance = Line resistance baseline value × (1 + ), the aging influence coefficient is represented here as the growth rate of line resistance over time.
[0087] Specifically, in step S2, constructing a line loss strategy includes:
[0088] Determine the actual resistance of the line between the distribution cabinets based on the initial resistance of the line, the temperature influence coefficient, and the aging influence coefficient;
[0089] The total line loss is determined according to the actual line resistance and the line current.
[0090] It's understandable that the line loss strategy is part of the pre-set target strategy. Based on the current state vector of the distribution cabinet and the state vectors of its neighbors, the line loss strategy determines the corresponding action vector for the distribution cabinet, laying the foundation for subsequent data analysis. Line loss primarily results from the heat generated by the resistance of the wires as current flows through them, consuming electrical energy.
[0091] In a specific embodiment, the line resistance of the kth line = the initial line resistance of the kth line × [1 + temperature influence coefficient × temperature difference)] × (1 + aging influence coefficient × line usage time), where the temperature difference is the difference between the ambient temperature and the temperature reference value. Total line loss = The actual line resistance of the kth line.
[0092] Specifically, in step S2, the action vector corresponding to the line loss strategy includes:
[0093] Adjust the reactive output of the power access cabinet to reduce line current;
[0094] Switch the load distribution cabinet switch combination to evenly distribute the load;
[0095] Adjust the charging / discharging power of the energy storage control cabinet to optimize the power flow distribution.
[0096] It is understandable that in the power system, the current consists of active current and reactive current. The unreasonable distribution of reactive power will lead to an increase in the current in the line, thereby increasing the line loss. Therefore, by adjusting the reactive output of the power access cabinet, the reactive power distribution in the system can be changed, so that the power factor can be improved and the line current can be reduced.
[0097] It is understandable that the loads on the lines and distribution cabinets within the power system may be unbalanced, with some lines being overloaded and other lines being lightly loaded. Therefore, by switching the switch combination of the load distribution cabinet, the load distribution can be changed so that the load is evenly distributed, thereby achieving current balance while avoiding excessive losses.
[0098] It's understandable that energy storage devices regulate power in power systems. Flow distribution can be irrational at different times and under different operating conditions, leading to excessive power transmission on some lines and significant losses. By adjusting the charge and discharge power of the energy storage control cabinet, energy can be flexibly stored and released, changing the direction and magnitude of power flow in the system, reducing circuitous and irrational power transmission, and lowering power losses in the lines.
[0099] This invention considers factors such as temperature and aging, and determines the actual resistance of each line between distribution cabinets based on line current, temperature influence coefficient, and aging influence coefficient. This determines the total line loss and constructs a line loss strategy, enabling targeted control of line losses. While ensuring safe line operation, this effectively reduces line losses, providing a solid foundation for achieving pre-set target strategies, improving the operating efficiency of the power system, and enhancing the reliability and stability of the entire power system.
[0100] Specifically, step S3 includes:
[0101] collecting the output current of each load distribution cabinet in the area in real time, determining the average load of the load distribution cabinet based on the output current, and determining the load balancing rate based on the average load;
[0102] Collecting data transmission and data reception of each power distribution cabinet in the area in real time, and determining communication efficiency based on the data transmission and data reception;
[0103] The collaborative objective reward function is determined by weighting according to the load balancing rate and the communication efficiency.
[0104] It is understandable that if the load distribution cabinet cannot achieve load balancing, some server cabinets may be underpowered while other cabinets have excess power, affecting the overall performance and stability of the data center. Therefore, when setting the collaborative objective reward function, the overall load of the load distribution cabinet is prioritized. Secondly, for the power distribution system, data exchange and communication between distribution cabinets are required to achieve coordinated control and status monitoring functions. High communication efficiency means that data can be transmitted quickly between distribution cabinets to ensure responsiveness. By comprehensively considering the load balancing rate and communication efficiency, the overall optimization of the regional power distribution system in terms of load distribution and communication coordination can be achieved, improving the overall performance of the system.
[0105] In a specific embodiment, the average load is the output current average of the load distribution cabinets in the area, and the load balancing rate is calculated using the standard deviation method. The load balancing rate = 1-output current standard deviation / average load. The load balancing rate ranges from 0 to 1, and the closer it is to 1, the more balanced the load. The communication efficiency = data reception amount / data transmission amount. The communication efficiency ranges from 0 to 1, and the higher the value, the higher the communication efficiency. The collaborative target reward function is the weighted sum of the load balancing rate and the communication efficiency, and the sum of the corresponding weights of the load balancing rate and the communication efficiency is 1. Preferably, the corresponding weights of the load balancing rate and the communication efficiency are both 0.5. The larger the value of the collaborative target reward function, the better the overall performance of the system. The value of the collaborative target reward function ranges from 0 to 1. In implementation, the corresponding weights of the load balancing rate and the communication efficiency can be determined according to actual conditions. No specific limitation is made here. As long as the sum of the corresponding weights of the load balancing rate and the communication efficiency is 1, and the corresponding weight of the load balancing rate is less than the corresponding weight of the communication efficiency, it will be sufficient. No further details will be given here.
[0106] The collaborative reward objective function of the present invention quantitatively weights the load balancing rate and communication efficiency, providing a clear and quantitative evaluation standard for the overall operating status of the power system, making it easier for managers to intuitively understand the system performance and facilitate the subsequent value of the collaborative objective reward function to quickly determine whether the system is operating in the optimal state, providing a clear direction for the optimization and adjustment of the system, improving the operating efficiency of the power system, and enhancing the reliability and stability of the entire power system.
[0107] See also Figure 4 As shown, Figure 4 The decision diagram for determining the action adjustment method in the embodiment of the present invention is as follows. Specifically, in step S4, the action adjustment method is determined, wherein:
[0108] If the total impact direction is a positive reward, then determining the action adjustment amplitude according to the individual impact direction;
[0109] If the total impact direction is a negative reward, repeat steps S2 to S4 until the total impact direction is a positive reward or the number of iterations is equal to the preset maximum number of iterations.
[0110] It is understandable that when the total impact direction is a positive reward, it means that the action vector of the distribution cabinet has a positive impact on the overall system. However, the individual impact direction may change with the execution of the distribution cabinet action vector. Therefore, after determining that the total impact direction is a positive reward, the action adjustment amplitude is determined based on the individual impact direction. When the total impact direction is a negative reward, it means that the current action vector and the decision based on the preset target strategy are not moving the system in the desired direction. The system's operating status may become worse or fail to achieve the expected results. At this time, by repeating steps S2 to S4, it is possible to re-evaluate the system status, adjust the preset target strategy, determine a new action vector, and find a strategy that can enable the system to obtain positive rewards.
[0111] Understandably, setting a maximum number of iterations is intended to prevent the algorithm from falling into an infinite loop. In practice, complex situations may make it difficult for the system to find a positive reward strategy in a short period of time. Unlimited iterations can consume significant computing resources and time, potentially leading to system failure.
[0112] In a specific embodiment, the preset maximum number of iterations is in a range of 5000 to 10000 times, and preferably, the preset maximum number of iterations is 6000 times. In implementation, the range and preferred value of the preset maximum number of iterations can be determined according to actual conditions and are not specifically limited here and will not be further described.
[0113] In a specific embodiment, if the action vector taken by the distribution cabinet can reduce the harmonic content or line loss, approaching the preset target strategy, the overall impact direction is positive reward, and the subsequent behavior of the distribution cabinet is determined according to the subsequent two-tier reward strategy. If the action vector taken by the distribution cabinet cannot reduce the harmonic content or line loss, and cannot approach the preset target strategy, the overall impact direction is negative reward.
[0114] Specifically, in step S4, determining the action adjustment range includes:
[0115] If the total impact direction is positive reward and the individual impact direction corresponding to the distribution cabinet is positive incentive, then the action adjustment amplitude of the action vector of the corresponding distribution cabinet is increased;
[0116] If the total impact direction is a positive reward and the individual impact direction corresponding to the distribution cabinet is a negative reward, the action adjustment amplitude of the action vector of the corresponding distribution cabinet is reduced.
[0117] It's understandable that when the total impact direction is positive reward, and the individual impact direction of a distribution cabinet is positive incentive, this indicates that the actions taken by that distribution cabinet not only have a positive impact on the overall system but also have a positive effect on improving its own state. For example, if a load distribution cabinet adjusts its load distribution, the overall load balancing rate will improve (the total impact direction is positive reward), while also making its own current, voltage, and other parameters more stable and reasonable (the individual impact direction is positive incentive), thus increasing the action adjustment amplitude of the action vector corresponding to the distribution cabinet.
[0118] If the overall impact direction is positive, but the individual impact direction of a distribution cabinet is negative, this means that while the actions taken by that distribution cabinet have a positive impact on the overall system, they may have a negative impact on the distribution cabinet itself. For example, an energy storage control cabinet may charge and discharge to optimize system power distribution, reducing overall system losses (positive overall impact direction), but the frequent charging and discharging may affect the life of its own battery (negative individual impact direction). Reducing the action adjustment amplitude of the corresponding distribution cabinet action vector can ensure overall optimization while minimizing the adverse effects on individual distribution cabinets.
[0119] In a specific embodiment, the adjustment percentage of the action adjustment amplitude of the power distribution cabinet corresponding to the action vector ranges from 4% to 8%. Preferably, the adjustment percentage corresponding to the action adjustment amplitude is 5%. In implementation, the range and preferred value of the adjustment percentage can be determined based on actual conditions and are not specifically limited here and will not be elaborated on.
[0120] The present invention determines the corresponding adjustment method based on the total influence direction and the individual influence direction, including determining the action adjustment amplitude under the total influence direction as a positive incentive. This can promote the overall performance of the system while taking into account the operating status of individual distribution cabinets. Alternatively, repeated iterations can be performed under the total influence direction as a negative incentive, prompting the system to find the optimal strategy in a complex and changing operating environment, improve its own robustness and adaptability, achieve continuous performance improvement, improve the accuracy and effectiveness of control, realize multi-objective collaborative control, and ensure the efficient operation of the system.
[0121] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
Claims
1. A multi-objective electrical automation collaborative control method integrating reinforcement learning, characterized in that: include: Step S1: Real-time collection of electrical parameters and operating status of each distribution cabinet in the area to construct its own state vector, and construct neighboring state vectors based on the power transmission direction and line impedance between adjacent distribution cabinets. The distribution cabinet types include power access cabinets, load distribution cabinets, and energy storage control cabinets. Step S2: constructing a preset target strategy according to the output current, output voltage, and line loss of each distribution cabinet, and determining the action vector of each distribution cabinet based on its own state vector, neighboring state vectors, and the preset target strategy; Step S3, determining the impact direction of the action vector on each distribution cabinet after execution based on the action vector, the own target reward function and the collaborative target reward function, wherein the impact direction includes the individual impact direction and the total impact direction; The self-target reward function is determined according to the type of distribution cabinet, including: the self-reward function of the power access cabinet includes a voltage deviation penalty term and a reactive output constraint term; the self-reward function of the load distribution cabinet includes a load balancing term and a tripping number penalty term; the self-reward function of the energy storage control cabinet includes an SOC health term and a charge and discharge efficiency term; collecting the output current of each load distribution cabinet in the area in real time, determining the average load of the load distribution cabinet based on the output current, and determining the load balancing rate based on the average load; Collecting data transmission and data reception of each power distribution cabinet in the area in real time, and determining communication efficiency based on the data transmission and data reception; Weighting is performed according to the load balancing rate and the communication efficiency to determine a collaborative objective reward function; Step S4, determining the action adjustment method according to the impact direction, wherein: If the total influence direction is a positive reward, the action adjustment amplitude is determined according to the individual influence direction, including: if the total influence direction is a positive reward and the individual influence direction corresponding to the distribution cabinet is a positive incentive, then the action adjustment amplitude of the action vector of the corresponding distribution cabinet is increased; if the total influence direction is a positive reward and the individual influence direction corresponding to the distribution cabinet is a negative reward, then the action adjustment amplitude of the action vector of the corresponding distribution cabinet is reduced; If the total impact direction is a negative reward, repeat steps S2 to S4 until the total impact direction is a positive reward or the number of iterations is equal to the preset maximum number of iterations.
2. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 1 is characterized in that: In step S2, a harmonic target strategy is determined according to the output voltage and the output current, a line loss strategy is determined according to the line loss, and a preset target strategy is determined according to the harmonic target strategy and the line loss strategy.
3. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 2 is characterized in that: In step S2, constructing a harmonic target strategy includes: Determine the current harmonic content and voltage harmonic content based on the output current and output voltage of each distribution cabinet; determining a current influence coefficient based on the current harmonic content and the current harmonic threshold; determining a voltage influence coefficient based on the voltage harmonic content and the voltage harmonic threshold; The harmonic target strategy is determined based on the current influence coefficient and the voltage influence coefficient.
4. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 3 is characterized in that: In step S2, the action vector corresponding to the harmonic target strategy includes: The power access cabinet injects reverse harmonics to offset the load harmonics; Disconnect the branch switch of the load distribution cabinet with high harmonic load or adjust the three-phase load distribution; The energy storage control cabinet offsets harmonics by charging / discharging energy storage.
5. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 2 is characterized in that: In step S2, constructing a line loss strategy includes: Determine the actual resistance of the line between the distribution cabinets based on the initial resistance of the line, the temperature influence coefficient, and the aging influence coefficient; The total line loss is determined according to the actual line resistance and the line current.
6. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 5 is characterized in that: In step S2, the temperature influence coefficient is determined based on the historical ambient temperature and the historical line resistance, and the aging influence coefficient is determined based on the line usage time and the historical line resistance.
7. The multi-objective electrical automation collaborative control method integrating reinforcement learning according to claim 5 is characterized in that: In step S2, the action vector corresponding to the line loss strategy includes: Adjust the reactive output of the power access cabinet to reduce line current; Switch the load distribution cabinet switch combination to evenly distribute the load; Adjust the charging / discharging power of the energy storage control cabinet to optimize the power flow distribution.
Citation Information
Patent Citations
A Hybrid Coordinated Control Method for Microgrids Based on Reinforcement Learning and Multi-Agent Theory
CN111200285B
ANFIS-based electrical circuit aging degree online identification system and method
CN113255203A
Double-layer cooperative control method for electricity-heat-gas comprehensive energy system
CN115102158A