A power distribution network source-load coordination control method and device, electronic equipment and storage medium

CN122823631APending Publication Date: 2026-09-25POWER DISPATCHING CONTROL CENT OF GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610981129.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2026-05-28
Filing Date
2026-07-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明实施例提供一种配电网源荷协调控制方法、装置、电子设备及存储介质,能够解决现有技术中单一区域训练样本不足与粗放融合引发跨区域知识冲突以及策略偏移的问题

Benefits of technology

本发明实施例提供一种配电网源荷协调控制方法、装置、电子设备及存储介质。所述方法获取目标配电网区域的光伏节点实时出力、负荷节点实时功率、节点实时电压及线路实时电流,生成实时状态特征;将实时状态特征输入目标策略网络,生成储能功率指令与光伏无功指令,并据此控制物理储能变流器与物理光伏逆变器运行;其中,目标策略网络由对应策略网络经两阶段训练更新得到,第一训练阶段基于多个配电网区域的历史状态特征与状态转移经验样本进行参数更新,形成历史样本集;第二训练阶段基于历史样本集进行性能评估、工况匹配及聚合权重计算,对候选策略网络参数加权聚合并替换更新,生成目标策略网络。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122823631A_ABST
    Figure CN122823631A_ABST
Patent Text Reader

Abstract

The application discloses a power distribution network source-load coordination control method and device, electronic equipment and storage medium, belongs to the power distribution network source-load coordination control technical field, the method includes: obtaining the real-time output of the photovoltaic node, the real-time power of the load node, the node voltage and the line current of the target power distribution network area, generating the real-time state characteristics; the real-time state characteristics are input into the target strategy network updated by two-stage training, the energy storage power instruction and the photovoltaic reactive power instruction are generated, and the energy storage converter and the photovoltaic inverter are controlled to operate;Wherein, two-stage training includes strategy network parameter updating based on historical state transition experience, and strategy network optimization updating based on performance evaluation, working condition matching and weighted aggregation. Therefore, by implementing the application, the problems of insufficient single area training samples, cross-region knowledge conflict and strategy deviation caused by extensive integration in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power distribution network source-load coordination control technology, specifically to a power distribution network source-load coordination control method, device, electronic equipment, and storage medium. Background Technology

[0002] With the large-scale integration of distributed photovoltaic (PV) power and new loads, the operating status of distribution networks exhibits extremely high volatility and uncertainty. Distribution network source-load coordination control relies on physical devices such as energy storage converters and PV inverters to dynamically adjust the active and reactive power within the network, effectively suppressing voltage fluctuations at nodes and maintaining global power balance. This is a core technological support for ensuring the safe and stable operation of new distribution networks.

[0003] Existing data-driven source-load coordinated control methods typically rely on independently training control models using historical data from a single region, or performing indiscriminate global averaging and fusion of model parameters from multiple regions. Independent training in a single region is limited by insufficient local sample diversity, making it difficult to cope with complex and ever-changing extreme operating scenarios. Indiscriminate global averaging and fusion ignores significant differences in topology and source-load fluctuation conditions across different distribution network regions, and fails to incorporate the safety control capabilities of the local models themselves into the fusion considerations. These crude model update and parameter fusion methods are highly susceptible to cross-regional knowledge conflicts and policy deviations, resulting in a final control network that cannot accurately output safe and reliable energy storage and photovoltaic regulation commands when facing specific defective operating conditions. This severely restricts the risk prevention and control capabilities and coordinated control accuracy of the distribution network. Summary of the Invention

[0004] This invention provides a method, device, electronic device, and storage medium for coordinated control of power distribution network sources and loads, which can solve the problems of insufficient training samples in a single region and cross-regional knowledge conflicts and policy deviations caused by extensive fusion in the prior art.

[0005] An embodiment of the present invention provides a source-load coordination control method for a distribution network, comprising: Obtain real-time output of photovoltaic nodes, real-time power of load nodes, real-time voltage of nodes, and real-time current of lines in the target distribution network area; Real-time status characteristics are generated based on the real-time output of photovoltaic nodes, the real-time power of load nodes, the real-time voltage of nodes, and the real-time current of lines. Real-time state characteristics are input into the target policy network so that the target policy network can generate energy storage power commands and photovoltaic reactive power commands based on the real-time state characteristics. Control the operation of physical energy storage converters and physical photovoltaic inverters within the target distribution network area according to energy storage power commands and photovoltaic reactive power commands; The target policy network is generated by updating the policy network corresponding to the target distribution network area through the first training stage and the second training stage. The first training stage generates state transition experience samples based on the historical state characteristics of multiple distribution network areas, and performs parameter updates on the policy networks corresponding to multiple distribution network areas based on the state transition experience samples, thereby generating historical sample sets corresponding to multiple distribution network areas. The second training phase evaluates the performance and matches operating conditions of the strategy networks corresponding to multiple distribution network areas based on historical sample sets, generating candidate strategy networks and aggregate weights for multiple distribution network areas. Weighted aggregation calculations are then performed based on the aggregate weights and network parameters of the candidate strategy networks to generate aggregated strategy network parameters for multiple distribution network areas. Finally, parameter replacements are performed on the strategy networks corresponding to multiple distribution network areas based on the aggregated strategy network parameters to generate updated strategy networks for multiple distribution network areas.

[0006] Furthermore, each distribution network area corresponds to a strategy network; The first training phase includes: Extract the base photovoltaic output and base load power corresponding to multiple distribution network areas within each control cycle; For each distribution network area, the first disturbance superposition calculation is performed based on the basic photovoltaic output corresponding to the current distribution network area and the preset random fluctuation rate of photovoltaic output to generate the historical output of the photovoltaic nodes corresponding to the current distribution network area. The second disturbance superposition calculation is performed based on the base load power corresponding to the current distribution network area and the preset random fluctuation rate of load demand to generate the historical power of the load node corresponding to the current distribution network area. Based on the historical output of photovoltaic nodes, the historical power of load nodes, and preset network topology parameters, power flow calculations are performed to generate the historical node voltage and line current corresponding to the current distribution network area. By integrating the historical output of photovoltaic nodes, the historical power of load nodes, the historical voltage of nodes, and the historical current of lines, the historical state characteristics of the current distribution network area are generated. Input the historical state characteristics of the current distribution network area into the corresponding strategy network for processing, and generate the energy storage power command and photovoltaic reactive power command corresponding to the current distribution network area. Based on the energy storage power command and photovoltaic reactive power command corresponding to the current distribution network area, calculate the operating cost, voltage deviation penalty, and reverse load rate penalty for the current distribution network area. By integrating operating costs, voltage deviation penalties, and reverse load rate penalties, the real-time reward value corresponding to the current distribution network area is calculated and generated. Integrate the historical state characteristics, energy storage power commands, photovoltaic reactive power commands, and real-time reward values ​​corresponding to the current distribution network area to generate state transition experience samples corresponding to the current distribution network area. Store the state transition experience samples corresponding to the current distribution network area into the experience replay cache corresponding to the current distribution network area.

[0007] Furthermore, based on the energy storage power command and photovoltaic reactive power command corresponding to the current distribution network area, the operating cost, voltage deviation penalty, and reverse load rate penalty corresponding to the current distribution network area are calculated, including: Perform amplitude extraction processing on the energy storage power command corresponding to the current distribution network area to generate the energy storage charging and discharging power amplitude corresponding to the current distribution network area; Based on the photovoltaic reactive power command, energy storage power command, historical output of photovoltaic nodes and historical power of load nodes in the current distribution network area, power balance analysis is performed to generate the main grid power purchase and curtailed photovoltaic power in the current distribution network area. The power purchase cost for the current distribution network area is generated by mapping the power purchased from the main grid to the preset time-of-use electricity price of the main grid; the energy storage loss cost for the current distribution network area is generated by mapping the energy storage charging and discharging power amplitude to the preset energy storage depreciation cost coefficient; and the curtailment penalty cost for the current distribution network area is generated by mapping the curtailed solar power to the preset curtailment penalty coefficient. The power purchase cost, energy storage loss cost, and curtailment penalty cost of the current distribution network area are integrated and aggregated to generate the operating cost of the current distribution network area. Based on the historical voltage of nodes corresponding to the current distribution network area, the preset upper limit of node voltage, and the preset lower limit of node voltage, the over-limit deviation is evaluated, and the first over-limit characteristic and the second over-limit characteristic corresponding to the current distribution network area are generated. The penalty is quantified based on the first and second over-limit characteristics of the current distribution network area and the preset node voltage deviation penalty coefficient, and a voltage deviation penalty item corresponding to the current distribution network area is generated. The load rate is calculated based on the historical current of the lines in the current distribution network area and the preset rated current carrying capacity of the lines, and the line load rate in the current distribution network area is generated. Based on the current line load rate of the distribution network area, the preset maximum load rate, and the preset line load rate penalty coefficient, the overload penalty is quantified to generate the reverse load rate penalty item corresponding to the current distribution network area.

[0008] Furthermore, each distribution network area corresponds to a value network; After storing the state transition experience samples corresponding to the current distribution network area into the experience replay cache corresponding to the current distribution network area, the following is also included: Extract state transition experience samples from the experience replay cache corresponding to the current distribution network area; Based on state transition experience samples, update the network parameters of the value network and the network parameters of the strategy network corresponding to the current distribution network area. Integrate all state transition experience samples in the experience playback cache corresponding to the current distribution network area, and splice them according to the time series to generate the historical sample set corresponding to the current distribution network area.

[0009] Furthermore, the second training phase includes: Historical voltage of nodes, historical current of lines, historical output of photovoltaic nodes, and historical power of load nodes were extracted from the historical sample set for multiple distribution network areas. Based on the historical node voltage and line current of multiple distribution network areas, the performance of the strategy network corresponding to multiple distribution network areas is evaluated, and the security evaluation index corresponding to multiple distribution network areas is generated. For each distribution network area, operating condition matching is performed based on the historical output of photovoltaic nodes, the historical power of load nodes, and the preset network topology parameters to generate operating condition feature labels for the current distribution network area. Based on the security assessment indicators and preset security thresholds corresponding to the current distribution network area, the network optimization category corresponding to the current distribution network area is determined. Based on the network optimization category and operating condition feature label of the current distribution network area, extract the defect distribution features of the current distribution network area; Based on the defect distribution characteristics of the current distribution network area, from the strategy networks corresponding to multiple distribution network areas, the strategy networks that meet the safety assessment indicators and match the defect distribution characteristics are selected to generate the candidate strategy network corresponding to the current distribution network area. The weighting coefficients are calculated based on the security assessment indicators corresponding to the candidate strategy network and the security assessment indicators corresponding to the current distribution network area to generate the aggregate weight corresponding to the current distribution network area. Based on the aggregation weights corresponding to the current distribution network area and the network parameters of the candidate strategy network, a weighted aggregation calculation is performed to generate the aggregation strategy network parameters corresponding to the current distribution network area. The parameters of the strategy network corresponding to the current distribution network area are replaced according to the aggregation strategy network parameters of the current distribution network area, and the updated strategy network corresponding to the current distribution network area is generated.

[0010] Furthermore, based on the historical node voltages and line currents corresponding to multiple distribution network areas, performance evaluations are performed on the strategy networks corresponding to each of the multiple distribution network areas, generating security evaluation indicators for each distribution network area, including: Based on the historical voltage of nodes corresponding to multiple distribution network areas, preset voltage safety boundary conditions, and preset first safety penalty coefficient, the risk of exceeding the limit is quantified, and a first risk feature sequence corresponding to multiple distribution network areas is generated. Overload risk is quantified based on the historical line currents of multiple distribution network areas, the preset line load safety boundary conditions, and the preset second safety penalty coefficient, and a second risk feature sequence corresponding to multiple distribution network areas is generated. The first risk feature sequence and the second risk feature sequence corresponding to multiple distribution network areas are integrated and superimposed to construct a comprehensive risk sequence corresponding to multiple distribution network areas. The comprehensive risk sequence corresponding to multiple distribution network areas is calculated by time averaging to generate safety assessment indicators for multiple distribution network areas.

[0011] Furthermore, based on the security assessment indicators corresponding to the candidate strategy network and the security assessment indicators corresponding to the current distribution network area, weighting coefficients are calculated to generate the aggregate weight corresponding to the current distribution network area, including: Based on the security assessment indicators corresponding to the current distribution network area and the preset zero constant for prevention and elimination, smoothing processing is performed to generate smoothed security indicators corresponding to the current distribution network area. Perform inverse proportional mapping on the smoothed safety index to generate the safety performance defect degree corresponding to the current distribution network area; Perform inverse proportional mapping on the security evaluation index corresponding to the candidate policy network to generate the individual security contribution of the candidate policy network. The individual security contribution and security performance defect of the candidate strategy network are integrated and correlated for evaluation to generate the initial fusion weight of the candidate strategy network corresponding to the current distribution network area. The initial fusion weights are aggregated to generate a global weight baseline; The aggregate weights corresponding to the current distribution network area are generated by normalizing the initial fusion weights and the global weight benchmark.

[0012] Based on the above method embodiments, the present invention provides corresponding apparatus embodiments.

[0013] One embodiment of the present invention provides a power distribution network source-load coordination control device, comprising: a data acquisition module, a feature generation module, an instruction generation module, and a coordination control module; The data acquisition module is used to acquire the real-time output of photovoltaic nodes, the real-time power of load nodes, the real-time voltage of nodes, and the real-time current of lines in the target distribution network area. The feature generation module is used to generate real-time status features based on the real-time output of photovoltaic nodes, the real-time power of load nodes, the real-time voltage of nodes, and the real-time current of lines. The instruction generation module is used to input real-time state characteristics into the target strategy network, so that the target strategy network can generate energy storage power instructions and photovoltaic reactive power instructions based on the real-time state characteristics. The target strategy network is generated by updating the strategy network corresponding to the target distribution network area through a first training phase and a second training phase. The first training phase generates state transition experience samples based on historical state characteristics of multiple distribution network areas, and updates the parameters of the strategy networks corresponding to multiple distribution network areas based on the state transition experience samples, generating a historical sample set for multiple distribution network areas. The second training phase performs performance evaluation and operating condition matching on the strategy networks corresponding to multiple distribution network areas based on the historical sample set, generating candidate strategy networks and aggregate weights for multiple distribution network areas. Weighted aggregation calculations are performed based on the aggregate weights and the network parameters of the candidate strategy networks to generate aggregated strategy network parameters for multiple distribution network areas. Parameter replacement is performed on the strategy networks corresponding to multiple distribution network areas based on the aggregated strategy network parameters to generate the updated strategy networks for multiple distribution network areas. The coordination control module is used to control the operation of physical energy storage converters and physical photovoltaic inverters within the target distribution network area according to energy storage power commands and photovoltaic reactive power commands.

[0014] Based on the above method embodiments, the present invention provides corresponding electronic device embodiments.

[0015] An embodiment of the present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements any one of the power distribution network source-load coordination control methods described in the above-described method embodiments.

[0016] Based on the above method embodiments, the present invention provides corresponding storage medium embodiments.

[0017] One embodiment of the present invention provides a storage medium storing a computer program thereon, wherein, when the computer program is running, it controls the device where the storage medium is located to execute any of the power distribution network source-load coordination control methods described in the above-described method embodiments.

[0018] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a method, apparatus, electronic device, and storage medium for coordinated control of power distribution network sources and loads. The method acquires real-time output of photovoltaic nodes, real-time power of load nodes, real-time node voltage, and real-time line current in a target power distribution network area to generate real-time state characteristics. These real-time state characteristics are then input into a target strategy network to generate energy storage power commands and photovoltaic reactive power commands, which are used to control the operation of physical energy storage converters and physical photovoltaic inverters. The target strategy network is obtained by training and updating a corresponding strategy network in two stages. The first training stage updates parameters based on historical state characteristics and state transition experience samples from multiple power distribution network areas, forming a historical sample set. The second training stage performs performance evaluation, operating condition matching, and aggregate weight calculation based on the historical sample set, weighting and aggregating the parameters of candidate strategy networks and replacing and updating them to generate the target strategy network.

[0019] In the second training phase, this invention utilizes historical sample sets to perform performance evaluation and operating condition matching on the local policy network parameter set, accurately screening candidate policy networks and generating aggregation weights. This directly overcomes the shortcomings of traditional indiscriminate global fusion, which ignores safety control capabilities and differences in specific operating conditions. Furthermore, based on the aggregation weights and the corresponding local policy network parameter sets of the candidate policy networks, weighted aggregation calculations and parameter replacements are performed, completely eliminating cross-regional knowledge conflicts and policy biases caused by insufficient training samples in a single region and coarse fusion. The updated target policy network can accurately generate energy storage power commands and photovoltaic reactive power commands for extreme defective operating conditions, significantly improving the risk prevention and control capabilities and global control accuracy of the distribution network. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a power distribution network source-load coordination control method according to an embodiment of the present invention.

[0021] Figure 2 This is a schematic diagram of the structure of a power distribution network source-load coordination control device provided in an embodiment of the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] like Figure 1 As shown, to address the problems of insufficient training samples in a single region and cross-regional knowledge conflicts and policy shifts caused by extensive fusion in existing technologies, an embodiment of the present invention provides a power distribution network source-load coordinated control method, which includes at least the following steps: Step S1: Obtain the real-time output of photovoltaic nodes, real-time power of load nodes, real-time voltage of nodes, and real-time current of lines in the target distribution network area; Specifically, for the source-load coordination control process of the distribution network, the primary preliminary action is to comprehensively perceive the physical operating status of the target distribution network area. This involves real-time collection of active power output data from each distributed photovoltaic (PV) node within the target distribution network area from the network's operational site, forming the real-time PV node output. Simultaneously, active power data from each load node within the target distribution network area is collected, forming the real-time load node power. To comprehensively understand the health status and load capacity of the transmission and distribution network, the voltage amplitude of each voltage node within the target distribution network area is further collected as the real-time node voltage, and the load current of each transmission line within the target distribution network area is obtained as the real-time line current.

[0024] Real-time output of photovoltaic nodes objectively reflects the immediate active power supply level of renewable energy generation units within the target distribution network area. Real-time power of load nodes accurately characterizes the immediate active power consumption demand of electricity users within the target distribution network area. Real-time node voltage is used to subsequently quantify the voltage exceedance risk boundary of the target distribution network area. Real-time line current is used to assess the overload operation degree of transmission lines in the target distribution network area.

[0025] The target distribution network area comprises a variety of physical nodes and transmission lines, and the collected multi-dimensional data entities constitute the underlying spatial vector structure. Specifically, the real-time power output of photovoltaic nodes satisfies the photovoltaic dimensional distribution relationship: In the formula, This provides the real-time power output vector for the photovoltaic nodes. It is the set of real numbers; This represents the total number of distributed photovoltaic nodes within the target distribution network area.

[0026] The real-time power of the load node conforms to the load dimension mapping condition: In the formula, This represents the real-time power vector of the load node. This represents the total number of load nodes within the target distribution network area.

[0027] The formula for the vectorized representation of node voltage corresponding to the real-time node voltage is as follows: In the formula, For the real-time voltage vector of the node; This represents the total number of voltage nodes within the target distribution network area.

[0028] The real-time line current follows the line current dimension constraint formula: In the formula, This represents the real-time current vector of the line. This refers to the total number of transmission lines within the target distribution network area.

[0029] All the acquired multidimensional state data entities together constitute the underlying physical observation foundation for the source-load coordinated control of the distribution network. By fusion and acquisition of photovoltaic output, load demand, node voltage, and line current within the target distribution network area, a high-fidelity data support base can be provided for source-load coordinated optimization control, ensuring that the subsequent strategy network can accurately perceive the dynamic fluctuation characteristics and topological carrying capacity of both sides of the distribution network.

[0030] Step S2: Generate real-time status characteristics based on the real-time output of photovoltaic nodes, the real-time power of load nodes, the real-time voltage of nodes, and the real-time current of lines; Specifically, after acquiring multi-dimensional physical observation data of the target distribution network area, the discrete physical parameters need to be transformed into standardized state expressions that can be directly processed by the strategy network. At the execution level, real-time output of photovoltaic nodes, real-time power of load nodes, real-time voltage of nodes, and real-time current of lines are spatially vectorized and dimensionally fused to generate real-time state characteristics of the target distribution network area. Real-time output of photovoltaic nodes and real-time power of load nodes jointly map the active power supply and demand balance within the target distribution network area, while real-time voltage of nodes and real-time current of lines jointly characterize the operational safety margin and network carrying capacity of the physical topology of the target distribution network area.

[0031] The mathematical construction logic of real-time state features follows the feature concatenation formula: In the formula, This is the real-time state feature vector.

[0032] By performing the aforementioned spatial dimension vector aggregation operation, source-side output data, load-side demand data, and grid state parameters scattered across different topological locations within the target distribution network area are integrated into a complete high-dimensional feature space. This high-dimensional feature space, composed of active power boundary features and safety boundary features, fully encompasses the underlying physical information required for the interaction process in the reinforcement learning environment. Through feature integration of multi-dimensional physical measurement data, a complete and standardized decision-making basis is provided for the subsequent generation of globally optimal source-load coordination control commands by the target policy network.

[0033] Step S3: Input the real-time state characteristics into the target strategy network so that the target strategy network can generate energy storage power commands and photovoltaic reactive power commands based on the real-time state characteristics; The target policy network is generated by updating the policy network corresponding to the target distribution network area through the first training stage and the second training stage. The first training stage generates state transition experience samples based on the historical state characteristics of multiple distribution network areas, and performs parameter updates on the policy networks corresponding to multiple distribution network areas based on the state transition experience samples, thereby generating historical sample sets corresponding to multiple distribution network areas. The second training phase evaluates the performance and matches operating conditions of the strategy networks corresponding to multiple distribution network areas based on historical sample sets, generating candidate strategy networks and aggregate weights for multiple distribution network areas. Weighted aggregation calculations are then performed based on the aggregate weights and network parameters of the candidate strategy networks to generate aggregated strategy network parameters for multiple distribution network areas. Finally, parameter replacements are performed on the strategy networks corresponding to multiple distribution network areas based on the aggregated strategy network parameters to generate updated strategy networks for multiple distribution network areas.

[0034] In a preferred embodiment, each distribution network area corresponds to a strategy network; The first training phase includes: Extract the base photovoltaic output and base load power corresponding to multiple distribution network areas within each control cycle; For each distribution network area, the first disturbance superposition calculation is performed based on the basic photovoltaic output corresponding to the current distribution network area and the preset random fluctuation rate of photovoltaic output to generate the historical output of the photovoltaic nodes corresponding to the current distribution network area. The second disturbance superposition calculation is performed based on the base load power corresponding to the current distribution network area and the preset random fluctuation rate of load demand to generate the historical power of the load node corresponding to the current distribution network area. Based on the historical output of photovoltaic nodes, the historical power of load nodes, and preset network topology parameters, power flow calculations are performed to generate the historical node voltage and line current corresponding to the current distribution network area. By integrating the historical output of photovoltaic nodes, the historical power of load nodes, the historical voltage of nodes, and the historical current of lines, the historical state characteristics of the current distribution network area are generated. Input the historical state characteristics of the current distribution network area into the corresponding strategy network for processing, and generate the energy storage power command and photovoltaic reactive power command corresponding to the current distribution network area. Based on the energy storage power command and photovoltaic reactive power command corresponding to the current distribution network area, calculate the operating cost, voltage deviation penalty, and reverse load rate penalty for the current distribution network area. By integrating operating costs, voltage deviation penalties, and reverse load rate penalties, the real-time reward value corresponding to the current distribution network area is calculated and generated. Integrate the historical state characteristics, energy storage power commands, photovoltaic reactive power commands, and real-time reward values ​​corresponding to the current distribution network area to generate state transition experience samples corresponding to the current distribution network area. Store the state transition experience samples corresponding to the current distribution network area into the experience replay cache corresponding to the current distribution network area.

[0035] In a preferred embodiment, based on the energy storage power command and photovoltaic reactive power command corresponding to the current distribution network area, the operating cost, voltage deviation penalty, and reverse load rate penalty corresponding to the current distribution network area are calculated, including: Perform amplitude extraction processing on the energy storage power command corresponding to the current distribution network area to generate the energy storage charging and discharging power amplitude corresponding to the current distribution network area; Based on the photovoltaic reactive power command, energy storage power command, historical output of photovoltaic nodes and historical power of load nodes in the current distribution network area, power balance analysis is performed to generate the main grid power purchase and curtailed photovoltaic power in the current distribution network area. The power purchase cost for the current distribution network area is generated by mapping the power purchased from the main grid to the preset time-of-use electricity price of the main grid; the energy storage loss cost for the current distribution network area is generated by mapping the energy storage charging and discharging power amplitude to the preset energy storage depreciation cost coefficient; and the curtailment penalty cost for the current distribution network area is generated by mapping the curtailed solar power to the preset curtailment penalty coefficient. The power purchase cost, energy storage loss cost, and curtailment penalty cost of the current distribution network area are integrated and aggregated to generate the operating cost of the current distribution network area. Based on the historical voltage of nodes corresponding to the current distribution network area, the preset upper limit of node voltage, and the preset lower limit of node voltage, the over-limit deviation is evaluated, and the first over-limit characteristic and the second over-limit characteristic corresponding to the current distribution network area are generated. The penalty is quantified based on the first and second over-limit characteristics of the current distribution network area and the preset node voltage deviation penalty coefficient, and a voltage deviation penalty item corresponding to the current distribution network area is generated. The load rate is calculated based on the historical current of the lines in the current distribution network area and the preset rated current carrying capacity of the lines, and the line load rate in the current distribution network area is generated. Based on the current line load rate of the distribution network area, the preset maximum load rate, and the preset line load rate penalty coefficient, the overload penalty is quantified to generate the reverse load rate penalty item corresponding to the current distribution network area.

[0036] In a preferred embodiment, each distribution network area corresponds to a value network; After storing the state transition experience samples corresponding to the current distribution network area into the experience replay cache corresponding to the current distribution network area, the following is also included: Extract state transition experience samples from the experience replay cache corresponding to the current distribution network area; Based on state transition experience samples, update the network parameters of the value network and the network parameters of the strategy network corresponding to the current distribution network area. Integrate all state transition experience samples in the experience playback cache corresponding to the current distribution network area, and splice them according to the time series to generate the historical sample set corresponding to the current distribution network area.

[0037] In a preferred embodiment, the second training phase includes: Historical voltage of nodes, historical current of lines, historical output of photovoltaic nodes, and historical power of load nodes were extracted from the historical sample set for multiple distribution network areas. Based on the historical node voltage and line current of multiple distribution network areas, the performance of the strategy network corresponding to multiple distribution network areas is evaluated, and the security evaluation index corresponding to multiple distribution network areas is generated. For each distribution network area, operating condition matching is performed based on the historical output of photovoltaic nodes, the historical power of load nodes, and the preset network topology parameters to generate operating condition feature labels for the current distribution network area. Based on the security assessment indicators and preset security thresholds corresponding to the current distribution network area, the network optimization category corresponding to the current distribution network area is determined. Based on the network optimization category and operating condition feature label of the current distribution network area, extract the defect distribution features of the current distribution network area; Based on the defect distribution characteristics of the current distribution network area, from the strategy networks corresponding to multiple distribution network areas, the strategy networks that meet the safety assessment indicators and match the defect distribution characteristics are selected to generate the candidate strategy network corresponding to the current distribution network area. The weighting coefficients are calculated based on the security assessment indicators corresponding to the candidate strategy network and the security assessment indicators corresponding to the current distribution network area to generate the aggregate weight corresponding to the current distribution network area. Based on the aggregation weights corresponding to the current distribution network area and the network parameters of the candidate strategy network, a weighted aggregation calculation is performed to generate the aggregation strategy network parameters corresponding to the current distribution network area. The parameters of the strategy network corresponding to the current distribution network area are replaced according to the aggregation strategy network parameters of the current distribution network area, and the updated strategy network corresponding to the current distribution network area is generated.

[0038] In a preferred embodiment, based on the historical node voltages and historical line currents corresponding to multiple distribution network areas, the performance of the strategy networks corresponding to the multiple distribution network areas is evaluated, generating security evaluation indicators for the multiple distribution network areas, including: Based on the historical voltage of nodes corresponding to multiple distribution network areas, preset voltage safety boundary conditions, and preset first safety penalty coefficient, the risk of exceeding the limit is quantified, and a first risk feature sequence corresponding to multiple distribution network areas is generated. Overload risk is quantified based on the historical line currents of multiple distribution network areas, the preset line load safety boundary conditions, and the preset second safety penalty coefficient, and a second risk feature sequence corresponding to multiple distribution network areas is generated. The first risk feature sequence and the second risk feature sequence corresponding to multiple distribution network areas are integrated and superimposed to construct a comprehensive risk sequence corresponding to multiple distribution network areas. The comprehensive risk sequence corresponding to multiple distribution network areas is calculated by time averaging to generate safety assessment indicators for multiple distribution network areas.

[0039] In a preferred embodiment, a weighted coefficient is calculated based on the security assessment index corresponding to the candidate policy network and the security assessment index corresponding to the current distribution network area to generate the aggregate weight corresponding to the current distribution network area, including: Based on the security assessment indicators corresponding to the current distribution network area and the preset zero constant for prevention and elimination, smoothing processing is performed to generate smoothed security indicators corresponding to the current distribution network area. Perform inverse proportional mapping on the smoothed safety index to generate the safety performance defect degree corresponding to the current distribution network area; Perform inverse proportional mapping on the security evaluation index corresponding to the candidate policy network to generate the individual security contribution of the candidate policy network. The individual security contribution and security performance defect of the candidate strategy network are integrated and correlated for evaluation to generate the initial fusion weight of the candidate strategy network corresponding to the current distribution network area. The initial fusion weights are aggregated to generate a global weight baseline; The aggregate weights corresponding to the current distribution network area are generated by normalizing the initial fusion weights and the global weight benchmark.

[0040] Specifically, after acquiring real-time state features, these features need to be input into the target policy network to drive it to perform forward propagation operations, thereby generating energy storage power commands and photovoltaic reactive power commands with physical security constraints. The target policy network is not trained using data from a single scenario, but rather iteratively updated from the policy networks corresponding to the target distribution network regions through a first training phase and a second training phase. The first training phase focuses on extensive exploration of the underlying operating environment, using historical state features from multiple distribution network regions to generate state transition experience samples, and then using these samples to update the internal parameters of the policy networks for multiple distribution network regions, generating historical sample sets for each region. The second training phase focuses on the targeted transfer of global knowledge from the cloud, using the historical sample sets to perform performance evaluation and operating condition matching on the policy networks for multiple distribution network regions, generating candidate policy networks and aggregated weights for each region. By combining the aggregation weights and the network parameters of the candidate policy networks, a weighted aggregation calculation is performed to generate aggregated policy network parameters for multiple distribution network areas. The aggregated policy network parameters are then used to replace the parameters of the policy networks for multiple distribution network areas, ultimately generating updated policy networks for multiple distribution network areas.

[0041] In the specific execution process of the first training phase, each distribution network area corresponds to an independent strategy network. Within each control cycle, the base photovoltaic output and base load power corresponding to multiple distribution network areas are comprehensively extracted. For each distribution network area, a first perturbation superposition calculation is performed by combining the base photovoltaic output corresponding to the current distribution network area with the preset random fluctuation rate of photovoltaic output to generate the historical output of the photovoltaic nodes corresponding to the current distribution network area. A second perturbation superposition calculation is performed using the base load power corresponding to the current distribution network area with the preset random fluctuation rate of load demand to obtain the historical power of the load nodes corresponding to the current distribution network area.

[0042] The preset random volatility of photovoltaic output and the preset random volatility of load demand characterize the extreme uncertainty boundaries on both the source and load sides of the distribution network. The values ​​of the preset random volatility of photovoltaic output and the preset random volatility of load demand are pre-calibrated based on historical meteorological statistics and regional electricity consumption behavior distribution patterns. They are used to artificially construct low-probability extreme over-limit operating conditions in the local training environment, forcing the strategy network to fully explore the physical safety boundaries of the distribution network.

[0043] Based on the historical output of photovoltaic (PV) nodes, the historical power of load nodes, and preset network topology parameters, power flow calculations are performed on the distribution network to deduce the historical node voltages and line currents corresponding to the current distribution network area. The historical output of PV nodes, the historical power of load nodes, the historical node voltages, and the historical line currents are spatially spliced ​​to construct the historical state features corresponding to the current distribution network area. Subsequently, the historical state features corresponding to the current distribution network area are input into the corresponding policy network for feature extraction and action mapping to predict and generate the energy storage power command and PV reactive power command corresponding to the current distribution network area.

[0044] To quantify the overall performance of the power generation and load coordination strategy in the distribution network, amplitude extraction processing is performed on the energy storage power command corresponding to the current distribution network area to obtain the energy storage charging and discharging power amplitude corresponding to the current distribution network area. Power balance analysis is then conducted by combining the photovoltaic reactive power command, energy storage power command, historical output of photovoltaic nodes, and historical power of load nodes corresponding to the current distribution network area to estimate the main grid power purchase and curtailed photovoltaic power corresponding to the current distribution network area. The power balance analysis process conforms to the law of conservation of energy and satisfies the node power balance mapping formula: In the formula, This represents the curtailed solar power in the current distribution network area. Select an operator for the maximum value; Historical power output of the photovoltaic nodes corresponding to the current distribution network area; This represents the energy storage charging and discharging power amplitude corresponding to the current distribution network area; This refers to the historical power of the load nodes corresponding to the current distribution network area. This represents the power purchased by the main grid for the current distribution network area.

[0045] After obtaining the main grid's purchased power and curtailed solar power, a cost mapping is performed using the main grid's purchased power and a preset main grid time-of-use tariff to generate the power purchase cost for the current distribution network area. A cost mapping is then performed based on the energy storage charging and discharging power amplitude and a preset energy storage depreciation cost coefficient to generate the energy storage loss cost for the current distribution network area. Finally, a cost mapping is performed using the curtailed solar power and a preset curtailment penalty coefficient to generate the curtailment penalty cost for the current distribution network area. The power purchase cost, energy storage loss cost, and curtailment penalty cost for the current distribution network area are then aggregated and merged into the operating cost for the current distribution network area.

[0046] The preset time-of-use electricity price for the main grid is pre-configured according to the actual transaction electricity price catalog of the power grid enterprise to which the target distribution network area belongs. The preset energy storage depreciation cost coefficient is pre-quantified according to the charge-discharge life attenuation model of the battery cluster matched with the physical energy storage converter. The preset curtailment penalty coefficient represents the economic penalty intensity of the grid assessment standard for renewable energy accommodation failure. By injecting the above real economic prior parameters, it is ensured that the generated operating cost fully conforms to the actual grid operation rules.

[0047] Simultaneously, the historical node voltage, preset upper limit of node voltage and preset lower limit of node voltage corresponding to the current distribution network area are used to perform out-of-limit deviation assessment, and the first out-of-limit feature and the second out-of-limit feature corresponding to the current distribution network area are extracted. According to the first out-of-limit feature, the second out-of-limit feature and the preset node voltage deviation penalty coefficient corresponding to the current distribution network area, penalty quantification is performed, and the voltage deviation penalty term corresponding to the current distribution network area is obtained through conversion. The historical line current and the preset rated current carrying capacity of the line corresponding to the current distribution network area are used to convert the load factor, and the line load factor corresponding to the current distribution network area is obtained. Combined with the line load factor, the preset maximum load factor and the preset line load factor penalty coefficient corresponding to the current distribution network area, overload penalty quantification processing is performed, and the reverse load rate penalty term corresponding to the current distribution network area is obtained.

[0048] As dimensionless hyperparameters, the preset node voltage deviation penalty coefficient and the preset line load factor penalty coefficient determine the penalty steepness of the mapping from out-of-limit physical quantities to economic penalties. By increasing the preset node voltage deviation penalty coefficient and the preset line load factor penalty coefficient, the policy network can be guided to always take avoiding voltage out-of-limit and line overload as the highest priority task in the local optimization process, thereby constructing a strict security defense line.

[0049] By integrating the operating cost, the voltage deviation penalty term and the reverse load rate penalty term, the immediate reward value corresponding to the current distribution network area is calculated and generated, and the core calculation logic of the immediate reward value satisfies the reward evaluation expression: In the formula, is the immediate reward value; is the operating cost; is the preset first safety penalty coefficient; is the voltage deviation penalty term; is the preset second safety penalty coefficient; is the reverse load rate penalty term.

[0050] The historical state characteristics, energy storage power commands, photovoltaic reactive power commands, and real-time reward values ​​corresponding to the current distribution network area are packaged according to the time sequence to generate state transition experience samples for the current distribution network area, and stored in the experience replay cache corresponding to the current distribution network area. Each distribution network area, in addition to the policy network, also independently corresponds to a value network. State transition experience samples are extracted from the experience replay cache corresponding to the current distribution network area. Based on the environmental feedback information covered by the state transition experience samples, the network parameters of the value network and the network parameters of the policy network corresponding to the current distribution network area are iteratively updated. All state transition experience samples in the experience replay cache corresponding to the current distribution network area are integrated and concatenated according to the time sequence to form the historical sample set corresponding to the current distribution network area.

[0051] After entering the second training phase, historical node voltages, line currents, photovoltaic node outputs, and load node power for multiple distribution network regions are extracted from the historical sample set. Over-limit risks are quantified using the historical node voltages, preset voltage safety boundary conditions, and preset first safety penalty coefficients for multiple distribution network regions, generating a first risk feature sequence for each region. Overload risks are quantified by combining the historical line currents, preset line load safety boundary conditions, and preset second safety penalty coefficients for multiple distribution network regions, generating a second risk feature sequence. The first and second risk feature sequences are then superimposed and aggregated to construct a comprehensive risk sequence for each distribution network region. A time-averaged calculation is performed on this comprehensive risk sequence to derive the safety assessment indices for each region. The dimensionality reduction and collapse process of the safety assessment indices satisfies the time-series statistical evaluation formula: In the formula, As a safety assessment indicator; This represents the total number of time steps included in the comprehensive risk sequence. For discrete time step variables; This represents a single point value of the first risk feature sequence at a specific time step. This represents a single point value of the second risk feature sequence at a specific time step.

[0052] For each distribution network area, in-depth operating condition matching is conducted by combining the historical output of photovoltaic nodes, the historical power of load nodes, and preset network topology parameters to extract the corresponding operating condition feature labels for the current distribution network area. The core of operating condition matching lies in quantifying the relative relationship between the level of renewable energy penetration and the load side, satisfying the source-load output ratio evaluation formula: In the formula, The proportion of power output from the source load; The total contribution of photovoltaic nodes throughout history; This represents the global sum of historical power at each load node. A zero-prevention constant is set as a preset value, which is defined as a very small positive real number that prevents singular calculations caused by a denominator of zero.

[0053] The network is categorized and classified using the security assessment indicators and preset security thresholds corresponding to the current distribution network area, and the network optimization category corresponding to the current distribution network area is anchored.

[0054] The preset security threshold represents a hard evaluation criterion for determining whether the network security performance of a policy network meets the standards in the cloud. Policy networks with security assessment indicators lower than or equal to the preset security threshold are considered to have excellent security generalization capabilities, while policy networks with security assessment indicators higher than the preset security threshold are considered to have control deficiencies and must undergo targeted knowledge transfer in the cloud. The preset security threshold acts as a digital barrier in the cloud to isolate inferior policy networks from superior policy networks.

[0055] Based on the network optimization category and operating condition feature labels corresponding to the current distribution network area, the defect distribution characteristics of the current distribution network area are extracted. Based on the defect distribution characteristics of the current distribution network area, from the policy networks corresponding to multiple distribution network areas, policy networks that meet the safety assessment indicators and completely match the defect distribution characteristics are accurately selected to form the candidate policy network corresponding to the current distribution network area.

[0056] The security assessment indicators corresponding to the current distribution network area are extracted and smoothed using a preset zero constant to generate a smoothed security indicator for the current distribution network area. An inverse proportional mapping is then performed on the smoothed security indicator to convert it into a security performance defect degree for the current distribution network area. Simultaneously, an inverse proportional mapping is performed on the security assessment indicators corresponding to the candidate strategy networks to generate individual security contribution degrees for each candidate strategy network. Finally, the individual security contribution degrees and security performance defect degrees of the candidate strategy networks are integrated and a correlation evaluation process is performed to calculate the initial fusion weights of the candidate strategy networks for the current distribution network area. The weight evaluation process satisfies the correlation mapping formula: In the formula, For safety performance defects; These are the initial fusion weights; These are the security evaluation metrics corresponding to the candidate policy network.

[0057] The initial fusion weights are aggregated and summed to generate a global weight benchmark. The initial fusion weights are then normalized by dividing the global weight benchmark, generating the aggregated weights corresponding to the current distribution network region. Finally, a weighted aggregation calculation is performed using the aggregated weights corresponding to the current distribution network region and the network parameters of the candidate strategy network to generate the aggregated strategy network parameters corresponding to the current distribution network region. The physical fusion logic of the aggregation calculation satisfies the parameter weighting formula: In the formula, For aggregation strategy network parameters; For summation operators; For aggregate weights; These are the network parameters for the candidate policy network.

[0058] Based on the aggregation strategy network parameters corresponding to the current distribution network area, the internal network parameters of the strategy network corresponding to the current distribution network area are directly replaced, and finally the updated strategy network corresponding to the current distribution network area is generated.

[0059] By introducing a dual collaborative training architecture that includes local underlying environment exploration and cloud-based targeted knowledge transfer, and by strictly standardizing the underlying mathematical mapping rules, the problem of insufficient safety generalization ability caused by the limited exploration space of a single model is effectively avoided. This enables the updated policy network to accurately output source-load coordination control commands that take into account both global economic benefits and stringent physical safety boundaries in the complex and ever-changing high-proportion renewable energy distribution network operation environment, significantly enhancing the robust defense capability of the distribution network dispatch control strategy in the face of extreme operating condition changes.

[0060] Step S4: Control the operation of physical energy storage converters and physical photovoltaic inverters within the target distribution network area according to the energy storage power command and photovoltaic reactive power command; Specifically, after the target policy network completes forward propagation and outputs control signals, it needs to transmit the digital control intent down to the underlying hardware infrastructure within the target distribution network area. Specifically, this involves controlling the operation of physical energy storage converters and physical photovoltaic inverters within the target distribution network area based on energy storage power commands and photovoltaic reactive power commands. After receiving the energy storage power command, the physical energy storage converter adjusts the duty cycle of its internal power electronic switching devices to change the active power exchange state between the energy storage battery cluster and the distribution network. Limited by the hardware manufacturing capacity of the underlying equipment, the actual physical charging and discharging power output of the physical energy storage converter must adhere to the energy storage physical operating boundary formula: In the formula, The actual physical charging and discharging power output by the physical energy storage converter; This represents the maximum allowable active power limit for a physical energy storage converter.

[0061] When the value corresponding to the energy storage power command is greater than zero, the physical energy storage converter is controlled to enter the discharge mode and inject active power into the target distribution network area; when the value corresponding to the energy storage power command is less than zero, the physical energy storage converter is controlled to enter the charging mode and absorb active power from the target distribution network area; when the value corresponding to the energy storage power command is equal to zero, the physical energy storage converter is controlled to maintain the standby mode and cut off the active power interaction.

[0062] Synchronously, after receiving the photovoltaic reactive power command, the physical photovoltaic inverter adjusts the output voltage phase and amplitude of its internal inverter bridge, changing the reactive power interaction level between the distributed photovoltaic nodes and the distribution network. The actual physical reactive power output by the physical photovoltaic inverter satisfies the photovoltaic physical operating boundary formula: In the formula, The actual physical reactive power output of the physical photovoltaic inverter; This represents the absolute value of the actual physical reactive power. This represents the maximum allowable reactive power limit for a physical photovoltaic inverter.

[0063] When the value corresponding to the photovoltaic reactive power command is greater than zero, the physical photovoltaic inverter is controlled to send inductive reactive power to the target distribution network area to increase the voltage amplitude of the voltage drop node; when the value corresponding to the photovoltaic reactive power command is less than zero, the physical photovoltaic inverter is controlled to absorb capacitive reactive power from the target distribution network area to suppress the node overvoltage phenomenon.

[0064] During the operation of the underlying physical equipment, the combined action of the physical energy storage converter and the physical photovoltaic inverter restructures the global power flow distribution of the target distribution network area. The source-load coordination control process is always strictly constrained by the physical laws of the power grid, and the actual physical operating voltage of all voltage nodes within the target distribution network area satisfies the voltage safety constraint formula: In the formula, This refers to the actual physical operating voltage of the voltage node; The preset node voltage lower limit; This is the preset upper limit of the node voltage.

[0065] The actual physical load current of all transmission lines within the target distribution network area satisfies the line load safety constraint formula: In the formula, This refers to the actual physical load current of the transmission line; The rated physical current carrying capacity of the transmission line; This is the preset maximum load rate.

[0066] By accurately translating the high-dimensional intelligent control strategy generated through cloud-based updates and iterations into actual action commands for underlying physical devices, and under the premise of strictly adhering to the hard physical operating boundaries of the power grid, real-time collaborative scheduling of massive distributed source-load-storage resources is achieved, ensuring voltage stability and topology load safety of the distribution network under complex operating conditions and extreme disturbances.

[0067] Based on the above method embodiments, the present invention provides corresponding apparatus embodiments.

[0068] like Figure 2 As shown, an embodiment of the present invention provides a power distribution network source-load coordination control device, including: a data acquisition module, a feature generation module, an instruction generation module, and a coordination control module; The data acquisition module is used to acquire the real-time output of photovoltaic nodes, the real-time power of load nodes, the real-time voltage of nodes, and the real-time current of lines in the target distribution network area. The feature generation module is used to generate real-time status features based on the real-time output of photovoltaic nodes, the real-time power of load nodes, the real-time voltage of nodes, and the real-time current of lines. The instruction generation module is used to input real-time state characteristics into the target strategy network, so that the target strategy network can generate energy storage power instructions and photovoltaic reactive power instructions based on the real-time state characteristics. The target strategy network is generated by updating the strategy network corresponding to the target distribution network area through a first training phase and a second training phase. The first training phase generates state transition experience samples based on historical state characteristics of multiple distribution network areas, and updates the parameters of the strategy networks corresponding to multiple distribution network areas based on the state transition experience samples, generating a historical sample set for multiple distribution network areas. The second training phase performs performance evaluation and operating condition matching on the strategy networks corresponding to multiple distribution network areas based on the historical sample set, generating candidate strategy networks and aggregate weights for multiple distribution network areas. Weighted aggregation calculations are performed based on the aggregate weights and the network parameters of the candidate strategy networks to generate aggregated strategy network parameters for multiple distribution network areas. Parameter replacement is performed on the strategy networks corresponding to multiple distribution network areas based on the aggregated strategy network parameters to generate the updated strategy networks for multiple distribution network areas. The coordination control module is used to control the operation of physical energy storage converters and physical photovoltaic inverters within the target distribution network area according to energy storage power commands and photovoltaic reactive power commands.

[0069] It should be noted that the embodiments of the device described above correspond to the embodiments of the present invention described above, and can realize the power distribution network source-load coordination control method described above in any one of the present invention. Furthermore, the embodiments of the device described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without creative effort.

[0070] Based on the above-described method embodiments of the present invention, a corresponding embodiment of an electronic device is provided.

[0071] An embodiment of the present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the power distribution network source-load coordination control method according to any one of the present invention, or the processor executes the computer program to implement the functions of each module in the above-described device embodiments.

[0072] For example, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the terminal device.

[0073] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0074] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.

[0075] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc.; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital card (SD card), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0076] Based on the above method embodiments, the present invention provides corresponding storage medium embodiments; Another embodiment of the present invention provides a storage medium including a stored computer program, wherein, when the computer program is running, it controls the device where the storage medium is located to execute any of the above-described power distribution network source-load coordination control methods of the present invention.

[0077] The aforementioned storage medium is a computer-readable storage medium, and the computer program includes computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0078] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0079] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for coordinated control of power distribution network sources and loads, characterized in that, include: Obtain real-time output of photovoltaic nodes, real-time power of load nodes, real-time voltage of nodes, and real-time current of lines in the target distribution network area; Real-time status characteristics are generated based on the real-time output of photovoltaic nodes, the real-time power of load nodes, the real-time voltage of nodes, and the real-time current of lines. Real-time state characteristics are input into the target policy network so that the target policy network can generate energy storage power commands and photovoltaic reactive power commands based on the real-time state characteristics. Control the operation of physical energy storage converters and physical photovoltaic inverters within the target distribution network area according to energy storage power commands and photovoltaic reactive power commands; The target policy network is generated by updating the policy network corresponding to the target distribution network area through the first training stage and the second training stage. The first training stage generates state transition experience samples based on the historical state characteristics of multiple distribution network areas, and performs parameter updates on the policy networks corresponding to multiple distribution network areas based on the state transition experience samples, thereby generating historical sample sets corresponding to multiple distribution network areas. The second training phase evaluates the performance and matches operating conditions of the strategy networks corresponding to multiple distribution network areas based on historical sample sets, generating candidate strategy networks and aggregate weights for multiple distribution network areas. Weighted aggregation calculations are then performed based on the aggregate weights and network parameters of the candidate strategy networks to generate aggregated strategy network parameters for multiple distribution network areas. Finally, parameter replacements are performed on the strategy networks corresponding to multiple distribution network areas based on the aggregated strategy network parameters to generate updated strategy networks for multiple distribution network areas.

2. The power distribution network source-load coordination control method as described in claim 1, characterized in that, Each distribution network area corresponds to a strategy network; The first training phase includes: Extract the base photovoltaic output and base load power corresponding to multiple distribution network areas within each control cycle; For each distribution network area, the first disturbance superposition calculation is performed based on the basic photovoltaic output corresponding to the current distribution network area and the preset random fluctuation rate of photovoltaic output to generate the historical output of the photovoltaic nodes corresponding to the current distribution network area. The second disturbance superposition calculation is performed based on the base load power corresponding to the current distribution network area and the preset random fluctuation rate of load demand to generate the historical power of the load node corresponding to the current distribution network area. Based on the historical output of photovoltaic nodes, the historical power of load nodes, and preset network topology parameters, power flow calculations are performed to generate the historical node voltage and line current corresponding to the current distribution network area. By integrating the historical output of photovoltaic nodes, the historical power of load nodes, the historical voltage of nodes, and the historical current of lines, the historical state characteristics of the current distribution network area are generated. Input the historical state characteristics of the current distribution network area into the corresponding strategy network for processing, and generate the energy storage power command and photovoltaic reactive power command corresponding to the current distribution network area. Based on the energy storage power command and photovoltaic reactive power command corresponding to the current distribution network area, calculate the operating cost, voltage deviation penalty, and reverse load rate penalty for the current distribution network area. By integrating operating costs, voltage deviation penalties, and reverse load rate penalties, the real-time reward value corresponding to the current distribution network area is calculated and generated. Integrate the historical state characteristics, energy storage power commands, photovoltaic reactive power commands, and real-time reward values ​​corresponding to the current distribution network area to generate state transition experience samples corresponding to the current distribution network area. Store the state transition experience samples corresponding to the current distribution network area into the experience replay cache corresponding to the current distribution network area.

3. The power distribution network source-load coordination control method as described in claim 2, characterized in that, Based on the energy storage power command and photovoltaic reactive power command corresponding to the current distribution network area, calculate the operating cost, voltage deviation penalty, and reverse load rate penalty for the current distribution network area, including: Perform amplitude extraction processing on the energy storage power command corresponding to the current distribution network area to generate the energy storage charging and discharging power amplitude corresponding to the current distribution network area; Based on the photovoltaic reactive power command, energy storage power command, historical output of photovoltaic nodes and historical power of load nodes in the current distribution network area, power balance analysis is performed to generate the main grid power purchase and curtailed photovoltaic power in the current distribution network area. The power purchase cost for the current distribution network area is generated by mapping the power purchased from the main grid to the preset time-of-use electricity price of the main grid; the energy storage loss cost for the current distribution network area is generated by mapping the energy storage charging and discharging power amplitude to the preset energy storage depreciation cost coefficient; and the curtailment penalty cost for the current distribution network area is generated by mapping the curtailed solar power to the preset curtailment penalty coefficient. The power purchase cost, energy storage loss cost, and curtailment penalty cost of the current distribution network area are integrated and aggregated to generate the operating cost of the current distribution network area. Based on the historical voltage of nodes corresponding to the current distribution network area, the preset upper limit of node voltage, and the preset lower limit of node voltage, the over-limit deviation is evaluated, and the first over-limit characteristic and the second over-limit characteristic corresponding to the current distribution network area are generated. The penalty is quantified based on the first and second over-limit characteristics of the current distribution network area and the preset node voltage deviation penalty coefficient, and a voltage deviation penalty item corresponding to the current distribution network area is generated. The load rate is calculated based on the historical current of the lines in the current distribution network area and the preset rated current carrying capacity of the lines, and the line load rate in the current distribution network area is generated. Based on the current line load rate of the distribution network area, the preset maximum load rate, and the preset line load rate penalty coefficient, the overload penalty is quantified to generate the reverse load rate penalty item corresponding to the current distribution network area.

4. The power distribution network source-load coordination control method as described in claim 3, characterized in that, Each distribution network area corresponds to a value network; After storing the state transition experience samples corresponding to the current distribution network area into the experience replay cache corresponding to the current distribution network area, the following is also included: Extract state transition experience samples from the experience replay cache corresponding to the current distribution network area; Based on state transition experience samples, update the network parameters of the value network and the network parameters of the strategy network corresponding to the current distribution network area. Integrate all state transition experience samples in the experience playback cache corresponding to the current distribution network area, and splice them according to the time series to generate the historical sample set corresponding to the current distribution network area.

5. The power distribution network source-load coordination control method as described in claim 4, characterized in that, The second training phase includes: Historical voltage of nodes, historical current of lines, historical output of photovoltaic nodes, and historical power of load nodes were extracted from the historical sample set for multiple distribution network areas. Based on the historical voltage of nodes and the historical current of lines in multiple distribution network areas, the performance of the strategy network in multiple distribution network areas is evaluated, and the security evaluation indicators for multiple distribution network areas are generated. For each distribution network area, operating condition matching is performed based on the historical output of photovoltaic nodes, the historical power of load nodes, and the preset network topology parameters to generate operating condition feature labels for the current distribution network area. Based on the security assessment indicators and preset security thresholds corresponding to the current distribution network area, the network optimization category corresponding to the current distribution network area is determined. Based on the network optimization category and operating condition feature label of the current distribution network area, extract the defect distribution features of the current distribution network area; Based on the defect distribution characteristics of the current distribution network area, from the strategy networks corresponding to multiple distribution network areas, the strategy networks that meet the safety threshold and match the defect distribution characteristics are selected to generate the candidate strategy network corresponding to the current distribution network area. The weighting coefficients are calculated based on the security assessment indicators corresponding to the candidate strategy network and the security assessment indicators corresponding to the current distribution network area to generate the aggregate weight corresponding to the current distribution network area. Based on the aggregation weights corresponding to the current distribution network area and the network parameters of the candidate strategy network, a weighted aggregation calculation is performed to generate the aggregation strategy network parameters corresponding to the current distribution network area. The parameters of the strategy network corresponding to the current distribution network area are replaced according to the aggregation strategy network parameters of the current distribution network area, and the updated strategy network corresponding to the current distribution network area is generated.

6. The power distribution network source-load coordination control method as described in claim 5, characterized in that, Based on the historical node voltages and line currents corresponding to multiple distribution network areas, the performance of the strategy networks corresponding to these multiple distribution network areas is evaluated, generating security evaluation indicators for each distribution network area, including: Based on the historical voltage of nodes corresponding to multiple distribution network areas, preset voltage safety boundary conditions, and preset first safety penalty coefficient, the risk of exceeding the limit is quantified, and a first risk feature sequence corresponding to multiple distribution network areas is generated. Overload risk is quantified based on the historical line currents of multiple distribution network areas, the preset line load safety boundary conditions, and the preset second safety penalty coefficient, and a second risk feature sequence corresponding to multiple distribution network areas is generated. The first risk feature sequence and the second risk feature sequence corresponding to multiple distribution network areas are integrated and superimposed to construct a comprehensive risk sequence corresponding to multiple distribution network areas. The comprehensive risk sequence corresponding to multiple distribution network areas is calculated by time averaging to generate safety assessment indicators for multiple distribution network areas.

7. The power distribution network source-load coordination control method as described in claim 6, characterized in that, Based on the security assessment indicators corresponding to the candidate strategy network and the security assessment indicators corresponding to the current distribution network area, weighting coefficients are calculated to generate the aggregate weight corresponding to the current distribution network area, including: Based on the security assessment indicators corresponding to the current distribution network area and the preset zero constant for prevention and elimination, smoothing processing is performed to generate smoothed security indicators corresponding to the current distribution network area. Perform inverse proportional mapping on the smoothed safety index to generate the safety performance defect degree corresponding to the current distribution network area; Perform inverse proportional mapping on the security evaluation index corresponding to the candidate policy network to generate the individual security contribution of the candidate policy network. The individual security contribution and security performance defect of the candidate strategy network are integrated and correlated for evaluation to generate the initial fusion weight of the candidate strategy network corresponding to the current distribution network area. The initial fusion weights are aggregated to generate a global weight baseline; The aggregate weights corresponding to the current distribution network area are generated by normalizing the initial fusion weights and the global weight benchmark.

8. A power distribution network source-load coordination control device, characterized in that, include: Data acquisition module, feature generation module, instruction generation module, and coordination and control module; The data acquisition module is used to acquire the real-time output of photovoltaic nodes, the real-time power of load nodes, the real-time voltage of nodes, and the real-time current of lines in the target distribution network area. The feature generation module is used to generate real-time status features based on the real-time output of photovoltaic nodes, the real-time power of load nodes, the real-time voltage of nodes, and the real-time current of lines. The instruction generation module is used to input real-time state characteristics into the target strategy network, so that the target strategy network can generate energy storage power instructions and photovoltaic reactive power instructions based on the real-time state characteristics. The target strategy network is generated by updating the strategy network corresponding to the target distribution network area through a first training phase and a second training phase. The first training phase generates state transition experience samples based on historical state characteristics of multiple distribution network areas, and updates the parameters of the strategy networks corresponding to multiple distribution network areas based on the state transition experience samples, generating a historical sample set for multiple distribution network areas. The second training phase performs performance evaluation and operating condition matching on the strategy networks corresponding to multiple distribution network areas based on the historical sample set, generating candidate strategy networks and aggregate weights for multiple distribution network areas. Weighted aggregation calculations are performed based on the aggregate weights and the network parameters of the candidate strategy networks to generate aggregated strategy network parameters for multiple distribution network areas. Parameter replacement is performed on the strategy networks corresponding to multiple distribution network areas based on the aggregated strategy network parameters to generate the updated strategy networks for multiple distribution network areas. The coordination control module is used to control the operation of physical energy storage converters and physical photovoltaic inverters within the target distribution network area according to energy storage power commands and photovoltaic reactive power commands.

9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the power distribution network source-load coordination control method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the storage medium is located to perform the power distribution network source-load coordination control method as described in any one of claims 1 to 7.