Intelligent distributed dc power supply control method and system based on neural network
By combining neural networks and distributed multi-agent reinforcement learning models, a power supply capacity distribution map and a fault risk index are generated to coordinate the power allocation of power nodes. This solves the adaptability and reliability problems of centralized control under uncertain operating conditions and realizes intelligent collaborative control and stable operation of DC microgrids.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SPEYI TECH (BEIJING) CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-06-02
AI Technical Summary
Existing centralized model predictive control technology suffers from insufficient real-time performance and adaptability of control commands when facing uncertain operating conditions such as output fluctuations of parallel power modules, changes in battery pack status, and sudden load changes. This limits system scalability and reliability, and it lacks the ability to detect and defend against local abnormal states within the DC power supply cabinet, thus affecting the long-term stable operation of the system.
A neural network-based intelligent distributed DC power supply control method is adopted. By fusing multi-source data, a power supply capacity distribution map and a potential fault risk index are generated. A distributed multi-agent reinforcement learning model is used to coordinate the power allocation weights of each power supply agent, and voltage correction signals are combined for collaborative control.
It realizes intelligent collaborative control of DC microgrid, improves the system's decision-making intelligence, response agility and operational robustness under fluctuation and fault conditions, and ensures the stable and efficient operation of the system.
Smart Images

Figure CN121863341B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of neural network technology, and in particular to an intelligent distributed DC power supply control method and system based on neural networks. Background Technology
[0002] As smart substations increasingly demand higher reliability and intelligence from DC power systems, problems such as uneven output from multiple parallel-operated DC power modules, differences in battery charge status, bus voltage oscillations caused by sudden load changes, and equipment failures caused by local hot spots within the DC power cabinet are becoming increasingly prominent. Therefore, such application scenarios urgently require a control method that can coordinate the distribution of multiple power modules, suppress voltage fluctuations in real time, and actively identify fault risks to ensure stable and efficient system operation.
[0003] Currently, a representative solution to address the above requirements is a centralized model predictive control technology. This solution collects system operation data through a central controller and, based on a preset optimization objective function and system model, continuously solves for the optimal power distribution command in future time periods. At the same time, it combines voltage feedback control to generate adjustment signals to achieve coordinated control of power output and bus voltage.
[0004] However, this approach relies on accurate system mathematical models and load forecasting results. When faced with uncertain operating conditions such as output fluctuations of parallel power modules, changes in battery pack status, and sudden load changes, the real-time performance and adaptability of control commands are limited. At the same time, the centralized control structure places high demands on communication bandwidth and central processing capabilities, posing challenges in terms of scalability and reliability. Furthermore, this method has limited ability to detect and defend against local abnormal states within the DC power supply cabinet, which is not conducive to the long-term reliable operation of the system. Summary of the Invention
[0005] This application provides a neural network-based intelligent distributed DC power supply control method and system to solve the problems of low operational stability and poor power supply dispatch efficiency of existing intelligent substation DC microgrids under fluctuating conditions.
[0006] To address the aforementioned technical problems, in a first aspect, this application provides an intelligent distributed DC power supply control method based on neural networks, comprising:
[0007] In the distributed DC microgrid of the smart substation, the DC bus voltage data, the output current data of each DC power module, the state of charge data of the battery pack, the power demand timing data of the load end, and the temperature data of key nodes in the DC power cabinet are acquired.
[0008] The output current data, the state of charge data, and the power demand timing data are fused together to generate a power supply capacity distribution map.
[0009] Hotspot distribution information is obtained by inverting the temperature data, and a potential fault risk index is calculated based on the hotspot distribution information.
[0010] The power supply capacity distribution map and the potential fault risk index are input into the distributed multi-agent reinforcement learning model. Through policy coordination and Q-value interaction among multiple power agents in the distributed multi-agent reinforcement learning model, the power allocation weights of each power agent are adjusted to generate power scheduling instructions.
[0011] Based on the DC bus voltage data, a voltage correction signal is generated;
[0012] By utilizing the power scheduling command and the voltage correction signal, the power output and DC bus voltage in the distributed DC power supply are adjusted to achieve intelligent collaborative control of the distributed DC power supply.
[0013] Secondly, this application provides an intelligent distributed DC power supply control system based on neural networks, comprising:
[0014] The acquisition module is used to acquire DC bus voltage data, output current data of photovoltaic panels in distributed DC power sources, state of charge data of wind turbines, power demand time-series data of load terminals, and temperature data of photovoltaic panels in distributed DC microgrids in smart substations.
[0015] The fusion module is used to fuse the output current data, the state of charge data, and the power demand timing data to generate a power supply capacity distribution map.
[0016] The calculation module is used to invert the hotspot distribution information based on the temperature data, and to calculate the potential fault risk index based on the hotspot distribution information;
[0017] The input module is used to input the power supply capacity distribution map and the potential fault risk index into the distributed multi-agent reinforcement learning model. Through policy coordination and Q-value interaction among multiple power agents in the distributed multi-agent reinforcement learning model, the power allocation weight of each power agent is adjusted to generate power scheduling instructions.
[0018] The generation module is used to generate a voltage correction signal based on the DC bus voltage data;
[0019] The adjustment module is used to adjust the power output and DC bus voltage in the distributed DC power supply using the power scheduling command and the voltage correction signal, so as to realize intelligent coordinated control of the distributed DC power supply.
[0020] Thirdly, this application provides an electronic device, comprising:
[0021] Memory, used to store computer programs;
[0022] A processor, configured to execute the computer program to implement the steps of the neural network-based intelligent distributed DC power supply control method as described in the first aspect above.
[0023] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the neural network-based intelligent distributed DC power supply control method described in the first aspect above.
[0024] This application provides an intelligent distributed DC power supply control method based on neural networks, which has the following advantages:
[0025] This application provides a comprehensive and real-time data foundation for system control through the aforementioned process, ensuring that subsequent decisions reflect the actual operating conditions. Furthermore, it transforms multi-source heterogeneous data into a unified and intuitive power supply capacity assessment, providing a forward-looking basis for power allocation. Simultaneously, it enables precise perception of the health status of key components within the DC power supply cabinet and early fault warnings, improving system operational safety. By enabling each power node to collaboratively respond to fluctuations and fault risks, it generates optimized power allocation commands adapted to complex operating conditions. It also quickly identifies and responds to voltage deviations, generating control signals to suppress oscillations and stabilize bus voltage. Ultimately, it achieves coordinated and optimized control of distributed power output and system voltage, ensuring the stable and efficient operation of the microgrid.
[0026] Furthermore, this application also inputs the power supply capacity distribution map and the potential fault risk index into the distributed multi-agent reinforcement learning model, so that each power agent calculates its own Q value with it as a local observation state and receives the Q value information of neighboring agents; then integrates its own and neighboring Q values, uses the policy gradient method to collaboratively update the power allocation strategy, dynamically adjusts the power allocation weight of each node, and finally generates power scheduling instructions based on the updated weights.
[0027] Therefore, the above process enables each power node to autonomously adapt to changes in power supply capacity and system risk through distributed collaborative learning, thereby dynamically optimizing power allocation strategies and improving the system's decision-making intelligence, response agility, and operational robustness under fluctuating and fault conditions.
[0028] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 A flowchart illustrating a neural network-based intelligent distributed DC power supply control method provided in this application embodiment;
[0031] Figure 2 A schematic diagram illustrating a specific implementation of an intelligent distributed DC power supply control method based on a neural network, provided in this application embodiment;
[0032] Figure 3 This is a schematic diagram of a neural network-based intelligent distributed DC power supply control system provided in an embodiment of this application. Detailed Implementation
[0033] In the field of DC microgrid control in smart substations, although existing centralized model predictive control schemes have achieved multi-power source coordination to a certain extent, their control performance is highly dependent on accurate mathematical models and predictive data. Faced with actual operating conditions such as power output fluctuations of parallel modules, changes in battery bank status, and frequent load changes, this scheme is difficult to adjust control commands in a timely manner, resulting in lag in power distribution and voltage regulation response. At the same time, the centralized architecture has high requirements for communication and computing resources, limiting system scalability and reliability, and also lacks the ability to perceive and defend against local abnormal states within the DC power supply cabinet, thus affecting the long-term stable operation of the system.
[0034] To address the aforementioned issues, this application proposes a neural network-based intelligent distributed DC power supply control method. Its core lies in achieving precise perception and intelligent control of the system's operating status through multi-source data fusion and distributed collaborative decision-making. Specifically, this method first generates a power supply capacity distribution map by fusing output current, state of charge, and load data, and then calculates a fault risk index by inverting hotspot distribution based on temperature data. Furthermore, it utilizes a distributed multi-agent reinforcement learning model to enable each power supply agent to dynamically adjust power allocation weights and generate scheduling commands through local information interaction and collaborative strategy updates. Simultaneously, it coordinates the power supply output and bus voltage adjustment in conjunction with voltage correction signals.
[0035] Therefore, this method effectively overcomes the dependence of centralized control on model accuracy and communication resources, improves the system's adaptability, decision-making speed and operational reliability under fluctuating and abnormal operating conditions, and thus realizes intelligent collaborative control of DC microgrids.
[0036] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0037] The core of this application is to provide an intelligent distributed DC power supply control method based on neural networks, and a flowchart of one specific implementation is shown below. Figure 1 As shown, the method includes:
[0038] Step 101: In the distributed DC microgrid of the smart substation, acquire DC bus voltage data, output current data of each DC power module, state of charge data of the battery pack, power demand timing data of the load end, and temperature data of key nodes in the DC power cabinet.
[0039] In step 101, the DC bus voltage refers to the voltage level of the power transmission bus in the distributed DC microgrid of the smart substation, which is a parameter for stable system operation.
[0040] Distributed DC power sources refer to various power units connected to the DC bus, such as parallel DC power modules and battery banks. Changes in their output power directly affect the stability of the DC bus voltage, and it is necessary to control and maintain a constant bus voltage.
[0041] DC bus voltage data refers to the measured value of the voltage on the bus in a DC microgrid; a DC power module is a basic power generation and supply unit in a distributed DC power source; output current data refers to the set of current measurement values at the output terminals of each parallel DC power module.
[0042] State of charge data refers to the percentage of the battery pack's current remaining capacity; power demand time series data refers to a series of recorded values of load power changing over time; and temperature data refers to the temperature measurements of key components such as power devices and bus connection points within the DC power supply cabinet.
[0043] For example, in a smart substation DC microgrid, the DC bus voltage data is collected as 220V by a voltage sensor. At the same time, the output current distribution data is collected by the current sensors of five parallel DC power supply modules, where module A is 50A, module B is 45A, module C is 55A, module D is 48A, and module E is 52A. The state of charge data is collected as 85% by the battery management system. In addition, the current power demand timing data is collected as 15kW by the load monitoring module. The temperature data are collected by five temperature sensors in the DC power supply cabinet as 45℃, 52℃, 48℃, 50℃, and 55℃, respectively.
[0044] Step 102: The output current data, the state of charge data, and the power demand time series data are fused to generate a power supply capacity distribution map.
[0045] In step 102, the power supply capacity distribution map refers to the probabilistic or numerical representation of the available power supply capacity of different power modules and battery packs in the future time period, generated through data fusion.
[0046] In this embodiment, the output current data is first divided into regions and the average current value is calculated. The state of charge data is converted into the available capacity output potential value through the battery discharge characteristic curve. At the same time, the power demand time series data is decomposed into demand components at different time scales. Then, the processed current distribution data, available capacity potential value and demand components are dynamically weighted and fused to form comprehensive power supply capacity data. Finally, the comprehensive situation data is probabilistically transformed based on historical data to generate a power supply capacity distribution map.
[0047] For example, firstly, the average of the five regional values of the output current data (50, 45, 55, 48, 52) is calculated to obtain 50A. Then, the 85% state of charge data is used to calculate the available capacity output potential value of 76.5 using a conversion factor of 0.9. At the same time, the power demand data of 15kW is decomposed into instantaneous demand of 10kW, medium-term demand of 4kW, and long-term demand of 1kW. Next, these data are weighted and fused according to weights of 0.4, 0.3, and 0.3, i.e., 50×0.4+76.5×0.3+10×0.3+4×0.3+1×0.3=47.45. Finally, based on the actual power supply capacity corresponding to similar values of 47.45 in historical data, a power supply capacity distribution map is generated with a module availability probability of 0.7 and a battery availability probability of 0.8.
[0048] Step 103: Obtain hotspot distribution information based on the temperature data, and calculate the potential fault risk index based on the hotspot distribution information.
[0049] In step 103, hot spot distribution information refers to the location and distribution of areas with abnormally high temperatures within the DC power supply cabinet, and potential fault risk refers to the quantitative value of the probability of a fault occurring calculated based on hot spot characteristics.
[0050] In this embodiment, the temperature data is first subjected to spectral analysis to extract characteristic frequency components related to thermal stress; then the temperature field inside the DC power supply cabinet is reconstructed based on the amplitude and phase information of these components; regions exceeding the temperature threshold are identified in the temperature field as hotspot distribution information; finally, the area ratio and temperature rise gradient of the hotspot region are calculated, and a potential fault risk index is obtained by fitting it with historical fault data.
[0051] For example, firstly, the temperature data of 45℃, 52℃, 48℃, 50℃, and 55℃ are decomposed to extract the 0.1Hz component with an amplitude of 2.0 and a phase of 30 degrees and the 0.2Hz component with an amplitude of 1.5 and a phase of 45 degrees. Then, the temperature field is reconstructed based on these components, and it is found that the local temperature reaches 85℃. Next, the area of the region with a temperature exceeding the threshold of 70℃ is identified as 0.05 square meters. Then, the area ratio is calculated as 0.05 / 2 = 0.025, and the temperature rise gradient is (85-65) / 0.1 = 200℃ / m². Finally, the potential failure risk index of 0.6 is obtained by fitting historical data.
[0052] Step 104: Input the power supply capacity distribution map and the potential fault risk index into the distributed multi-agent reinforcement learning model. Through policy coordination and Q-value interaction among multiple power agents in the distributed multi-agent reinforcement learning model, adjust the power allocation weights of each power agent to generate power scheduling instructions.
[0053] In step 104, the power node is the intelligent control entity of each DC power module and battery pack connected to the DC bus in the intelligent substation DC microgrid, and the power allocation weight refers to the proportional coefficient of each power node in the total power allocation.
[0054] The power dispatch command is a control command generated by a distributed multi-agent reinforcement learning model. It contains the target output power value and control parameters of each power node. The command clearly specifies the specific power output value that each DC power module and battery pack needs to achieve within a specific time window. It is used to directly drive the power converter of each power node to adjust its operating state, thereby achieving precise control of the power distribution of the microgrid.
[0055] The explanation, structural design, and parameter design of distributed multi-agent reinforcement learning models can all be found in relevant technologies, and will not be elaborated here.
[0056] In this embodiment, the power supply capacity distribution map and the potential fault risk index are input into the distributed multi-agent reinforcement learning model. Each power agent calculates its own Q value as a local observation state and exchanges Q value information with neighboring agents through a communication network. After integrating its own and neighboring Q values, each agent uses a policy gradient method to collaboratively update the power allocation strategy and adjust the power allocation weights. Finally, a power scheduling instruction is generated based on the adjusted weights.
[0057] For example, power agent 1 is module A, receiving a module availability probability of 0.7 and a risk index of 0.6, calculating its own Q value of 2.5; power agent 2 receives a battery availability probability of 0.8 and a risk index of 0.6, calculating its own Q value of 3.0; after exchanging Q values, power agent 1 calculates a joint Q value of 2.5×0.6+3.0×0.4=2.7 with weights of 0.6 and 0.4, while power agent 2 calculates a joint Q value of 3.0×0.7+2.5×0.3=2.85; after updating the policy using the policy gradient method, the weight of power agent 1 is adjusted to 0.45, and the weight of power agent 2 is adjusted to 0.65; based on the weights, a power scheduling instruction is generated: the target power of module A is 0.45×14.25=6.4125kW, and the target power of the battery pack is 0.65×14.25=9.2625kW.
[0058] Step 105: Generate a voltage correction signal based on the DC bus voltage data.
[0059] In step 105, the voltage correction signal refers to the control signal used to adjust and stabilize the DC bus voltage.
[0060] In this embodiment, the collected DC bus voltage data is first compared with the rated voltage value to obtain the voltage deviation; then, time series analysis is performed on the voltage deviation to extract the voltage fluctuation mode and oscillation frequency components; next, feedforward compensation components and feedback components are generated based on the fluctuation mode and frequency components; finally, these two components are fused to generate a voltage correction signal.
[0061] For example, first, the measured DC bus voltage of 215V is compared with the rated value of 220V to obtain a deviation of -5V; then, the deviation sequence is analyzed to find that the fluctuation amplitude is ±3V and the frequency is 100Hz; then, a feedforward compensation component of +4V and a feedback component of -2V are generated; finally, the voltage correction signal is obtained by fusion as 4+(-2)=+2V.
[0062] Step 106: Using the power scheduling command and the voltage correction signal, adjust the power output and DC bus voltage in the distributed DC power supply to achieve intelligent collaborative control of the distributed DC power supply.
[0063] In step 106, intelligent collaborative control refers to the unified optimization control of multiple system parameters through multi-instruction coordination.
[0064] In this embodiment, a power scheduling command is first sent to the power converters of each DC power module and battery pack to adjust their output power; at the same time, a voltage correction signal is injected into the bus voltage regulator to modulate the voltage level; finally, coordinated control is achieved through the coupling relationship between power and voltage.
[0065] For example, scheduling commands for the target power of module A (6.4125kW) and the target power of the battery pack (9.2625kW) are sent to the corresponding controllers, while a voltage correction signal of +2V is injected into the voltage regulator. The module A controller adjusts its output to 6.4125kW, the battery pack controller adjusts its output to 9.2625kW, and the voltage regulator boosts the bus voltage by 2V to 217V, thereby achieving coordinated control.
[0066] This application achieves optimized power allocation, voltage stability control, and fault defense for intelligent substation DC microgrids under fluctuating operating conditions through multi-source data acquisition and fusion, hotspot fault early warning, and distributed intelligent decision-making and collaborative control, thereby improving the adaptability, reliability, and intelligence level of system operation.
[0067] To address the adaptability and reliability issues of multi-energy coordinated control in distributed DC microgrids of smart substations, in some embodiments, step 104 involves inputting the power supply capacity distribution map and the potential fault risk index into a distributed multi-agent reinforcement learning model. Through policy coordination and Q-value interaction among multiple power agents in the distributed multi-agent reinforcement learning model, the power allocation weights of each power agent are adjusted to generate power scheduling instructions, such as... Figure 2 As shown, it includes:
[0068] Step 201: Input the power supply capacity distribution map and the potential fault risk index into the distributed multi-agent reinforcement learning model. In the distributed multi-agent reinforcement learning model, each power supply agent uses the power supply capacity distribution map and the potential fault risk index as a local observation state.
[0069] In step 201, the power agent is the basic decision-making unit in the distributed multi-agent reinforcement learning model. Each agent represents a specific power node, such as a DC power module or a battery pack.
[0070] Local observation status refers to the set of local environmental information that each power agent can acquire. This set of local environmental information includes a power supply capacity distribution map reflecting the power supply capacity and a potential fault risk index reflecting the health status of the equipment.
[0071] Step 202: Based on the local observation state of each power agent, calculate its own Q value, and receive the Q values of neighboring power agents through the communication network.
[0072] In step 202, the self-Q value refers to the long-term benefit assessment value calculated by the agent based on the local observation state, and the neighbor Q value refers to the long-term benefit assessment value obtained from neighboring agents through the communication network;
[0073] A communication network refers to a pre-deployed communication infrastructure in a distributed DC microgrid of a smart substation, used for data exchange between various power supply agents. This network connects all agents through wired or wireless communication links.
[0074] The communication network is the basic supporting environment for the distributed multi-agent reinforcement learning model. It provides the necessary communication channels for the agents in the model to realize Q-value exchange and policy coordination. The model relies on this network to realize the distributed collaborative learning function.
[0075] In the embodiments of this application, each power agent calculates its own Q value based on the received local observation state using a reinforcement learning algorithm, and exchanges Q value information with directly connected neighbor agents through a communication network, thereby realizing the sharing of local benefit evaluation.
[0076] Step 203: By integrating the Q-values of each power agent and the Q-values of neighboring power agents, the power allocation strategy is updated collaboratively using the policy gradient method to adjust the power allocation weights of each power agent.
[0077] In step 203, the power allocation strategy refers to the control rules that determine the power allocation ratio of each power node.
[0078] In this embodiment, each agent first integrates its own Q value with the received neighbor Q values, then uses the policy gradient method to calculate the policy update direction, and reaches a policy consensus through a distributed coordination mechanism to collaboratively update the power allocation policy, and finally adjusts the power allocation weight of each power node accordingly.
[0079] Step 204: Generate power scheduling instructions for each power node based on the adjusted power allocation weights.
[0080] In this embodiment of the application, based on the adjusted power allocation weight and combined with the current total power demand, the specific target output power value of each power node is calculated and encapsulated into an executable power scheduling instruction.
[0081] Here is a specific example:
[0082] In a smart substation DC microgrid, the module availability probability of 0.7, the battery availability probability of 0.8, and the potential fault risk index of 0.6 in the power supply capacity distribution map are first input into the distributed multi-agent reinforcement learning model. Among them, power agent 1 takes the module availability probability of 0.7 and the risk index of 0.6 as the local observation state, while power agent 2 takes the battery availability probability of 0.8 and the risk index of 0.6 as the local observation state.
[0083] Next, power agent #1 calculates its own Q-value based on the local observation state using the Q-value calculation formula, which is as follows: ,in This indicates the Q value of power supply agent #1. The weighting coefficient is set to 1.5. This indicates the probability that the module is available, for example, a value of 0.7. This is the weighting coefficient, for example, a value of 2.0. This represents the risk index, such as a value of 0.6, which is used in the calculation. ;
[0084] Meanwhile, the No. 2 power supply agent calculates its own Q value based on the local observation state using the Q-value calculation formula, which is as follows: ,in This indicates the Q value of power supply agent #2. This is the weighting coefficient, for example, a value of 2.0. This indicates the probability that the battery is usable, for example, a value of 0.8. This is the weighting coefficient, for example, a value of 1.5. This represents the risk index, such as a value of 0.6, which is used in the calculation. ;
[0085] Subsequently, Power Agent 1 receives the Q-value of 2.5 from its neighbor Power Agent 2 via the communication network, and Power Agent 2 receives the Q-value of 2.25 from its neighbor Power Agent 1. Each agent integrates its own Q-value and the neighbor's Q-value, and Power Agent 1 calculates the joint Q-value using a weighted fusion formula, which is as follows: ,in This represents the combined Q value of power supply agent #1. Set its own weight to 0.6. Let the neighbor weight be 0.4, and substituting it in, we get... ;
[0086] The formula for calculating the joint Q value of power source agent #2 is as follows: ,in This represents the joint Q value of power agent #2. Set its own weight to 0.7. Let the neighbor weight be 0.3, then substituting it into the equation yields... ;
[0087] Then, the power allocation strategy is updated collaboratively using the policy gradient method. After adjustment, the power allocation weight of power agent 1 changes from 0.45 to 0.48, and the power allocation weight of power agent 2 changes from 0.65 to 0.62. Finally, power scheduling instructions for each power node are generated based on the adjusted power allocation weights. The target power of module A is 0.48 × 14.25 = 6.84 kW, and the target power of the battery pack is 0.62 × 14.25 = 8.835 kW, thus completing the instruction generation.
[0088] In this embodiment, the collaborative decision-making and dynamic adjustment of each power node through a distributed multi-agent reinforcement learning framework can adaptively cope with power supply capacity fluctuations and equipment failure risks, thereby generating optimized power scheduling instructions and improving the intelligence level and control accuracy of the DC microgrid operation in the smart substation.
[0089] To address the accuracy and consistency issues in policy collaborative updates and weight adjustments during distributed multi-agent reinforcement learning, in some embodiments, step 203—which involves integrating the Q-values of each power agent and the Q-values of its neighboring power agents, and using a policy gradient method to collaboratively update the power allocation policy to adjust the power allocation weights of each power agent—includes:
[0090] Step 301: For each power agent, the Q value of itself is weighted and fused with the Q values of neighboring power agents to form a target joint Q value.
[0091] In step 301, the target joint Q value refers to the comprehensive benefit evaluation value obtained by each agent after fusing its own Q value and the Q value of its neighbors in a weighted manner. It is used to reflect the overall decision value of local and neighbor information.
[0092] In the embodiments of this application, each power agent calculates its own Q value and the received Q value of the neighboring power agents according to a preset weighting coefficient, thereby forming a target joint Q value, which provides a basis for subsequent policy updates.
[0093] Step 302: Based on the target joint Q value, calculate the gradient update amount of the power allocation strategy using the policy gradient method.
[0094] In step 302, the gradient update amount refers to the adjustment magnitude of the policy parameters calculated using the policy gradient method based on the joint Q value of the objective, which is used to guide the optimization direction of the power allocation policy.
[0095] In the embodiments of this application, each power agent calculates the gradient update amount of the power allocation strategy based on the obtained target joint Q value, thereby determining the adjustment direction and magnitude of the strategy parameters.
[0096] Step 303: Through a distributed coordination mechanism, exchange the gradient update amount with neighboring power agents, adjust the gradient update direction based on the exchange result, and collaboratively update the power allocation strategy based on the updated gradient update direction.
[0097] In step 303, the distributed coordination mechanism refers to the interaction rules that agents reach a consensus on through information exchange, and the gradient update direction refers to the overall orientation of policy parameter adjustment.
[0098] In the embodiments of this application, each power agent exchanges gradient update amounts with neighboring agents through a distributed coordination mechanism, and adjusts its local gradient update direction according to the gradient update amounts of all neighbors, thereby collaboratively updating the power allocation strategy based on the updated direction.
[0099] Step 304: Map the collaboratively updated power allocation strategy to the power allocation weights of each power agent to obtain the adjusted power allocation weights.
[0100] In this embodiment of the application, the collaboratively updated power allocation strategy is converted into the specific power allocation weights of each power agent through a linear mapping relationship, thereby obtaining the adjusted power allocation weights.
[0101] In this embodiment, by weighted fusion of Q-values, distributed coordinated gradient updates, and policy mapping, multi-agent policy collaborative optimization and precise weight adjustment are achieved, thereby improving the rationality of power allocation and system stability.
[0102] To address the issue of translating power allocation weights into specific control commands, in some embodiments, step 204: generating power scheduling commands for each power node based on the adjusted power allocation weights, includes:
[0103] Step 401: Convert the adjusted power allocation weights into the corresponding power settings.
[0104] In step 401, the power setting value refers to the coefficient that can be directly used for power calculation after the power allocation weight has been standardized, and its value ranges from 0 to 1.
[0105] In this embodiment, the adjusted power allocation weights are directly used as power settings because these weight coefficients have been normalized and can directly reflect the proportional relationship of each power node in the total power output.
[0106] Step 402: Calculate the total power reference value based on the current total load demand and bus voltage status of the DC microgrid.
[0107] In step 402, the total power reference value refers to the required total power reference value calculated based on the current total system load demand and bus voltage status, which is used to determine the reference scale of power allocation.
[0108] In this embodiment, the total load demand data and bus voltage data of the DC microgrid are collected in real time. Then, the load demand is converted into a total power demand value based on the bus voltage according to the power calculation formula. Finally, the total power reference value is obtained after considering the system operating status adjustment.
[0109] It should be noted that the embodiments of this application do not specifically limit the form of the power calculation formula, and can be set accordingly according to the actual situation.
[0110] Step 403: Multiply the power setting value with the total power reference value to obtain the target output power of each power node, and encapsulate the target output power into a power scheduling command.
[0111] In step 403, the target output power refers to the specific power value that each power node needs to output.
[0112] In this embodiment, the power setting value is first multiplied by the total power reference value to obtain the target output power of each power node. Then, these power values, together with control information such as node address and timestamp, are encapsulated into structured instructions, which are finally formed into power scheduling instructions that can be sent to the controllers of each power node.
[0113] In this embodiment, through weight conversion, benchmark value calculation and instruction encapsulation, the accurate conversion from power allocation weights to specific control instructions is achieved, thereby ensuring the accuracy and executability of power scheduling instructions, and thus providing a reliable control foundation for the stable operation of the DC microgrid in the smart substation.
[0114] To address the issues of multi-source heterogeneous data fusion and power supply capacity assessment, in some embodiments, step 102: fusing the output current data, the state of charge data, and the power demand time-series data to generate a power supply capacity distribution map includes:
[0115] Step 501: Divide the output current data into regions to form multiple current distribution region data.
[0116] In step 501, the current distribution area data refers to the average or weighted value of the output current in each area after dividing multiple parallel modules in the DC power supply system according to their physical location or electrical grouping, reflecting the output difference of power modules in different areas.
[0117] In this embodiment of the application, the output current data is divided into regions according to the module arrangement in the DC power supply cabinet, and then the average current value of each region is calculated to form multiple current distribution region data.
[0118] Step 502: Convert the state of charge data into available capacity data.
[0119] In step 502, the available capacity data refers to the equivalent power value that can be used to support the load, calculated based on the current state of charge and discharge characteristics of the battery pack, reflecting the power supply potential of the current energy storage system.
[0120] Step 503: Decompose the power demand time series data into demand components at different time scales.
[0121] In step 503, the demand component refers to the power value that reflects the electricity consumption characteristics at different time scales after decomposing the power demand time series data. The demand component includes instantaneous demand, medium-term demand and long-term demand components.
[0122] In this embodiment of the application, a multi-scale decomposition method is used to decompose the power demand time series data into demand components at different time scales in order to capture the time series characteristics of load changes.
[0123] Step 504: Dynamically weight and superimpose multiple current distribution area data, available capacity data, and multiple demand components to form comprehensive power supply capacity data.
[0124] In step 504, the comprehensive power supply capacity data refers to the overall power supply capacity status assessment value obtained by weighted fusion of multiple power supply-related data.
[0125] Step 505: Based on the historical operation data of the distributed DC microgrid of the smart substation, perform probabilistic transformation on the comprehensive power supply capacity data to generate a power supply capacity distribution map.
[0126] In this embodiment of the application, a power supply capacity distribution map is generated by using a probability statistical method based on the actual power supply availability corresponding to similar comprehensive power supply capacity data in historical operation data.
[0127] In this embodiment of the application, by dividing, transforming, decomposing, weighting, and probabilistically processing multi-source data regions, an accurate assessment of power supply capacity is achieved, thereby providing a reliable decision-making basis for subsequent power allocation.
[0128] To address the issue of accurate detection and risk assessment of hotspot faults within DC power supply cabinets, in some embodiments, step 103—the process of obtaining hotspot distribution information based on the temperature data and calculating a potential fault risk index based on the hotspot distribution information—includes:
[0129] Step 601: Perform spectral decomposition on the temperature data to decompose it into multiple frequency components, forming a frequency component set.
[0130] In step 601, the frequency component set refers to a set of multiple frequency components obtained by performing spectral analysis on the temperature data, and each frequency component contains amplitude and phase information.
[0131] Step 602: Select characteristic frequency components from the set of frequency components that match the preset thermal stress variation law.
[0132] In step 602, the thermal stress change law refers to the physical mapping relationship of the material thermal expansion and stress distribution change caused by the Joule heat generated by the current passing through the power device in the DC power supply cabinet. Specifically, it is manifested as the nonlinear mapping law between temperature fluctuations and hot spots within a specific frequency range.
[0133] Characteristic frequency components refer to specific frequency components selected from the set of frequency components that match the pattern of thermal stress changes. These components can reflect the stress characteristics caused by temperature changes.
[0134] Step 603: Reconstruct the temperature distribution field inside the DC power supply cabinet based on the amplitude and phase distribution of the characteristic frequency components.
[0135] In step 603, "inside the DC power supply cabinet" refers to the overall spatial structure composed of key components such as power modules, bus connection points, fuses, and contactors in the DC power supply system; "temperature distribution field" refers to the spatial distribution of temperature inside the cabinet obtained by reconstructing the amplitude and phase information of characteristic frequency components.
[0136] In this embodiment of the application, the temperature values at each location in the DC power supply cabinet are calculated based on the amplitude and phase distribution of the characteristic frequency components, thereby forming a temperature distribution field.
[0137] Step 604: Identify abnormal temperature rise areas in the temperature distribution field where the temperature value exceeds a preset temperature threshold, and use the spatial location, temperature value, and area range information corresponding to the abnormal temperature rise areas as hotspot distribution information.
[0138] In this embodiment of the application, regions with temperature values exceeding a preset threshold are scanned and identified in the temperature distribution field, and the spatial coordinates, specific temperature values, and regional ranges of these abnormal temperature rise regions are recorded, thereby forming hotspot distribution information.
[0139] Step 605: Calculate the area ratio and average temperature gradient value of the hotspot distribution information.
[0140] In step 605, the area ratio refers to the ratio of the area of the abnormal temperature rise zone to the total area of the key monitoring area inside the cabinet, and the average temperature rise gradient value refers to the rate of temperature change per unit distance within the abnormal area.
[0141] In this embodiment of the application, the ratio of the total area of the abnormal region in the hotspot distribution information to the total area of the monitoring area is calculated, and the average gradient value of the temperature change with distance in the abnormal region is also calculated.
[0142] Step 606: Match and fit the area ratio and the average temperature rise gradient value with the preset historical fault database to generate a potential fault risk index.
[0143] In this embodiment of the application, the calculated area ratio and average temperature rise gradient value are matched and fitted with cases in a preset historical fault database, and then a potential fault risk index is generated through a risk assessment model.
[0144] In this embodiment, by using spectrum analysis, temperature field reconstruction, and hotspot feature extraction, accurate detection and risk assessment of hotspot faults in DC power supply cabinets are achieved, thereby providing a reliable basis for preventive maintenance of the system.
[0145] To address the problem of suppressing and stabilizing DC bus voltage fluctuations, in some embodiments, step 105: generating a voltage correction signal based on the DC bus voltage data includes:
[0146] Step 701: Compare the voltage value corresponding to the DC bus voltage data with the preset rated voltage value to obtain the instantaneous voltage deviation.
[0147] In step 701, the instantaneous voltage deviation refers to the difference between the measured DC bus voltage value and the preset rated voltage value, reflecting the instantaneous degree of voltage deviation from the ideal state.
[0148] In this embodiment of the application, the real-time collected DC bus voltage data is compared with the preset rated voltage value, and the difference between the two is calculated to obtain the instantaneous voltage deviation.
[0149] Step 702: Perform time series analysis on the instantaneous voltage deviation to obtain the voltage fluctuation mode and oscillation frequency components.
[0150] In step 702, the voltage fluctuation mode refers to the regular characteristics of the voltage deviation changing over time, and the oscillation frequency component refers to the main frequency components contained in the voltage fluctuation.
[0151] In this embodiment of the application, time series analysis is performed on the instantaneous voltage deviation, and the periodicity and main frequency components of the voltage fluctuation are identified by methods such as spectrum analysis;
[0152] The specific implementation process involves collecting voltage deviation values at continuous time points to form a data sequence, and then identifying the periodic variation patterns and main frequency characteristics through spectrum analysis; for example, if the voltage deviation is detected to fluctuate within a range of ±3V with a period of 0.1 seconds, the dominant oscillation frequency component of 100Hz is extracted.
[0153] Step 703: Generate feedforward compensation component and feedback component based on the voltage fluctuation mode and the oscillation frequency component.
[0154] In step 703, the feedforward compensation component refers to the predictive adjustment amount generated based on the voltage fluctuation mode, and the feedback component refers to the stability adjustment amount generated based on the oscillation frequency component.
[0155] In this embodiment, the feedforward compensation component is calculated based on the amplitude and phase characteristics of the voltage fluctuation mode, and the feedback component is calculated based on the amplitude of the oscillation frequency component.
[0156] The specific implementation process is as follows: first, calculate the forward compensation amount based on the amplitude and phase characteristics of the fluctuation mode; then, calculate the damping suppression amount based on the amplitude of the oscillation frequency component. For example, for the 100Hz oscillation component, generate a feedforward compensation component with an amplitude of +2V and a phase leading of 90 degrees, as well as a feedback component with a gain of -1.5.
[0157] Step 704: Fuse the feedforward compensation component and the feedback component to generate a voltage correction signal.
[0158] In this embodiment, by using voltage deviation detection, fluctuation analysis, and compensation component fusion, precise suppression and rapid stabilization control of DC bus voltage fluctuations are achieved, thereby improving the system voltage quality.
[0159] Here is a specific example:
[0160] In the DC microgrid of the smart substation, Hall current sensors of model AHKC-EKCA deployed in the DC feeder cabinet collect the output current data of each parallel DC power module in real time. At the same time, the battery pack connected to the DC power monitoring unit is equipped with a battery management terminal of model MM912_637 to obtain the state of charge data of the battery pack. On the load side, the power demand timing data of the load side is collected by a smart power monitoring instrument of model DTSD342-5N. Key nodes in the DC power cabinet, such as power device heat sinks, bus connection points and fuse bases, are equipped with surface mount platinum resistance temperature sensors and fiber optic grating temperature sensors of model ATE100M to obtain the temperature data inside the cabinet. On the DC bus side, the DC bus voltage data is collected by a voltage transmitter of model CE-VM02-54MS.
[0161] The various data collected by the aforementioned sensors and monitoring terminals are aggregated to the station protection control unit or edge computing node via an industrial Ethernet switch of model GIS2528 or a process layer network conforming to the IEC 61850 standard. The unit deploys a distributed multi-agent reinforcement learning model and runs on an industrial edge computer of model EIC-2000.
[0162] Furthermore, the power agents corresponding to each DC power module and battery pack interact and coordinate strategies through the station's communication network. The power scheduling commands generated by the model are sent to the digital rectifier controllers of each DC power module and the bidirectional DC / DC converters of the battery packs via IEC 61850 goose messages or Modbus TCP protocol. The voltage correction signal is superimposed on the bus voltage control loop through the automatic voltage regulator, ultimately achieving coordinated regulation of the output power of the DC power module and the DC bus voltage.
[0163] Figure 3 A schematic diagram of a neural network-based intelligent distributed DC power supply control system is provided as an embodiment of this application, as shown below. Figure 3 As shown, the system includes:
[0164] The acquisition module 31 is used to acquire DC bus voltage data, output current data of each DC power module, state of charge data of the battery pack, power demand timing data of the load end, and temperature data of key nodes in the DC power cabinet in the distributed DC microgrid of the smart substation.
[0165] The fusion module 32 is used to fuse the output current data, the state of charge data, and the power demand timing data to generate a power supply capacity distribution map.
[0166] The calculation module 33 is used to invert the hotspot distribution information based on the temperature data and calculate the potential fault risk index based on the hotspot distribution information.
[0167] The input module 34 is used to input the power supply capacity distribution map and the potential fault risk index into the distributed multi-agent reinforcement learning model. Through policy coordination and Q-value interaction among multiple power agents in the distributed multi-agent reinforcement learning model, the power allocation weights of each power agent are adjusted to generate power scheduling instructions.
[0168] The generation module 35 is used to generate a voltage correction signal based on the DC bus voltage data.
[0169] The adjustment module 36 is used to adjust the power output and DC bus voltage in the distributed DC power supply using the power scheduling command and the voltage correction signal, so as to realize intelligent coordinated control of the distributed DC power supply.
[0170] The neural network-based intelligent distributed DC power supply control system of this application embodiment is used to implement the aforementioned neural network-based intelligent distributed DC power supply control method. Therefore, the specific implementation of the neural network-based intelligent distributed DC power supply control system can be found in the embodiment section of the neural network-based intelligent distributed DC power supply control method above. The specific implementation can be referred to the description of the corresponding embodiments, and will not be repeated here.
[0171] This application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described neural network-based intelligent distributed DC power supply control methods.
[0172] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-described neural network-based intelligent distributed DC power supply control methods.
[0173] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory, random access memory, portable hard drives, magnetic disks, or optical disks.
[0174] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the embodiments of the neural network-based intelligent distributed DC power supply control method described above.
[0175] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0176] The above provides a detailed description of the intelligent distributed DC power supply control method and system based on neural networks provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A method for controlling an intelligent distributed DC power supply based on a neural network, characterized in that, include: In the distributed DC microgrid of the smart substation, the DC bus voltage data, the output current data of each DC power module, the state of charge data of the battery pack, the power demand timing data of the load end, and the temperature data of key nodes in the DC power cabinet are acquired. The output current data, the state of charge data, and the power demand timing data are fused together to generate a power supply capacity distribution map. Hotspot distribution information is obtained by inverting the temperature data, and a potential fault risk index is calculated based on the hotspot distribution information. The power supply capacity distribution map and the potential fault risk index are input into the distributed multi-agent reinforcement learning model. Through policy coordination and Q-value interaction among multiple power agents in the distributed multi-agent reinforcement learning model, the power allocation weights of each power agent are adjusted to generate power scheduling instructions, where the Q-value represents the long-term benefit assessment value. Based on the DC bus voltage data, a voltage correction signal is generated; By utilizing the power dispatch command and the voltage correction signal, the power output and DC bus voltage in the distributed DC power supply are adjusted to achieve intelligent collaborative control of the distributed DC power supply. The step of inputting the power supply capacity distribution map and the potential fault risk index into a distributed multi-agent reinforcement learning model, and adjusting the power allocation weights of each power agent through policy coordination and Q-value interaction among multiple power agents in the distributed multi-agent reinforcement learning model to generate power scheduling instructions includes: The power supply capacity distribution map and the potential fault risk index are input into the distributed multi-agent reinforcement learning model. In the distributed multi-agent reinforcement learning model, each power supply agent uses the power supply capacity distribution map and the potential fault risk index as a local observation state. Based on the local observation state of each power agent, calculate its own Q value and receive the Q values of neighboring power agents through the communication network. By integrating the Q-values of each power agent and the Q-values of neighboring power agents, the power allocation strategy is updated collaboratively using the policy gradient method to adjust the power allocation weights of each power agent. Based on the adjusted power allocation weights, power scheduling instructions are generated for each power node. The process involves integrating the Q-values of each power agent and those of its neighboring power agents, and using a policy gradient method to collaboratively update the power allocation strategy to adjust the power allocation weights of each power agent. This includes: For each power agent, its own Q value is weighted and fused with the Q values of its neighboring power agents to form a target joint Q value; Based on the target joint Q value, the gradient update amount of the power allocation strategy is calculated using the policy gradient method; Through a distributed coordination mechanism, the gradient update amount is exchanged with neighboring power agents, and the gradient update direction is adjusted based on the exchange result. Based on the updated gradient update direction, the power allocation strategy is updated collaboratively. The updated power allocation strategy is mapped to the power allocation weights of each power agent to obtain the adjusted power allocation weights.
2. The method according to claim 1, characterized in that, The step of generating power scheduling instructions for each power node based on the adjusted power allocation weights includes: Convert the adjusted power allocation weights into the corresponding power settings; Calculate the total power reference value based on the current total load demand and bus voltage status of the DC microgrid; The target output power of each power node is obtained by multiplying the power setting value with the total power reference value, and the target output power is encapsulated into a power scheduling command.
3. The method according to claim 1, characterized in that, The process of fusing the output current data, the state of charge data, and the power demand time-series data to generate a power supply capacity distribution map includes: The output current data is divided into regions to form multiple current distribution region data; Convert the state of charge data into available capacity data; The power demand time-series data is decomposed into demand components at different time scales; The data of multiple current distribution areas, the available capacity data, and the multiple demand components are dynamically weighted and superimposed to form comprehensive power supply capacity data. Based on the historical operating data of the distributed DC microgrid of the smart substation, the comprehensive power supply capacity data is probabilistically transformed to generate a power supply capacity distribution map.
4. The method according to claim 1, characterized in that, The process of obtaining hotspot distribution information based on the temperature data and calculating a potential fault risk index based on the hotspot distribution information includes: The temperature data is subjected to spectral decomposition to extract multiple frequency components, forming a frequency component set. Select characteristic frequency components that match the preset thermal stress variation law from the set of frequency components; Based on the amplitude and phase distribution of the characteristic frequency components, the temperature distribution field inside the DC power supply cabinet is reconstructed; In the temperature distribution field, an abnormal temperature rise region where the temperature value exceeds a preset temperature threshold is identified, and the spatial location, temperature value and region range information corresponding to the abnormal temperature rise region are used as hotspot distribution information. Calculate the area proportion and average temperature rise gradient value of the hotspot distribution information; The area ratio and the average temperature rise gradient value are matched and fitted with a preset historical fault database to generate a potential fault risk index.
5. The method according to claim 1, characterized in that, The step of generating a voltage correction signal based on the DC bus voltage data includes: The instantaneous voltage deviation is obtained by comparing the corresponding voltage value in the DC bus voltage data with the preset rated voltage value. Time series analysis was performed on the instantaneous voltage deviation to obtain the voltage fluctuation mode and oscillation frequency components; Based on the voltage fluctuation mode and the oscillation frequency components, feedforward compensation components and feedback components are generated; The feedforward compensation component and the feedback component are fused to generate a voltage correction signal.
6. A neural network-based intelligent distributed DC power supply control system, characterized in that, include: The acquisition module is used to acquire DC bus voltage data, output current data of each DC power module, state of charge data of battery packs, power demand timing data of the load end, and temperature data of key nodes in the DC power cabinet in the distributed DC microgrid of the smart substation. The fusion module is used to fuse the output current data, the state of charge data, and the power demand timing data to generate a power supply capacity distribution map. The calculation module is used to invert the hotspot distribution information based on the temperature data, and to calculate the potential fault risk index based on the hotspot distribution information; The input module is used to input the power supply capacity distribution map and the potential fault risk index into the distributed multi-agent reinforcement learning model. Through policy coordination and Q-value interaction among multiple power agents in the distributed multi-agent reinforcement learning model, the power allocation weight of each power agent is adjusted to generate power scheduling instructions, where the Q-value represents the long-term benefit evaluation value. The generation module is used to generate a voltage correction signal based on the DC bus voltage data; The adjustment module is used to adjust the power output and DC bus voltage in the distributed DC power supply using the power scheduling command and the voltage correction signal, so as to realize intelligent coordinated control of the distributed DC power supply. The step of inputting the power supply capacity distribution map and the potential fault risk index into a distributed multi-agent reinforcement learning model, and adjusting the power allocation weights of each power agent through policy coordination and Q-value interaction among multiple power agents in the distributed multi-agent reinforcement learning model to generate power scheduling instructions includes: The power supply capacity distribution map and the potential fault risk index are input into the distributed multi-agent reinforcement learning model. In the distributed multi-agent reinforcement learning model, each power supply agent uses the power supply capacity distribution map and the potential fault risk index as a local observation state. Based on the local observation state of each power agent, calculate its own Q value and receive the Q values of neighboring power agents through the communication network. By integrating the Q-values of each power agent and the Q-values of neighboring power agents, the power allocation strategy is updated collaboratively using the policy gradient method to adjust the power allocation weights of each power agent. Based on the adjusted power allocation weights, power scheduling instructions are generated for each power node. The process involves integrating the Q-values of each power agent and those of its neighboring power agents, and using a policy gradient method to collaboratively update the power allocation strategy to adjust the power allocation weights of each power agent. This includes: For each power agent, its own Q value is weighted and fused with the Q values of its neighboring power agents to form a target joint Q value; Based on the target joint Q value, the gradient update amount of the power allocation strategy is calculated using the policy gradient method; Through a distributed coordination mechanism, the gradient update amount is exchanged with neighboring power agents, and the gradient update direction is adjusted based on the exchange result. Based on the updated gradient update direction, the power allocation strategy is updated collaboratively. The updated power allocation strategy is mapped to the power allocation weights of each power agent to obtain the adjusted power allocation weights.
7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the neural network-based intelligent distributed DC power supply control method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the implementation of the neural network-based intelligent distributed DC power supply control method as described in any one of claims 1 to 5.