Adjustment method and device of power distribution network, computer program product and electronic equipment

By reducing the dimensionality of the distribution network historical data and optimizing the intelligent model, energy storage and photovoltaic installation strategies are generated, and the line overload and stability problems when photovoltaic is connected to the distribution network are solved, and the robustness and economicality of the distribution network are improved.

CN120237678APending Publication Date: 2025-07-01STATE GRID BEIJING ELECTRIC POWER CO +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510400371.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

There are problems of line overload and poor stability when photovoltaics are connected to the distribution network. Traditional line expansion and transformation methods lead to high investment costs and waste of resources, and traditional trend calculations are insufficient in accuracy under high proportion of photovoltaic penetration.

Method used

By obtaining historical status data of the distribution network for dimensionality reduction processing, using intelligent models to generate energy storage and photovoltaic installation strategies, and adjusting the charging and discharging behavior of the energy storage equipment of the distribution network based on control instructions, combining stacked encoder and deep learning model to optimize decisions.

Benefits of technology

It enhances the robustness of the distribution network, slows down line overload, improves equipment utilization and economy, and achieves more efficient grid operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120237678A_ABST
    Figure CN120237678A_ABST
Patent Text Reader

Abstract

The invention discloses an adjustment method and device of a power distribution network, a computer program product and electronic equipment. Relates to the field of power distribution network planning or other related fields, and the method comprises the steps: obtaining M pieces of state data of a power distribution network in a historical time period, carrying out the dimension reduction processing of the M pieces of state data, and obtaining M pieces of processed state data, the M pieces of state data at least comprising node load data, energy storage data and photovoltaic installation data of the power distribution network; inputting the M processed state data into a first intelligent agent model, and processing to obtain a power distribution network planning strategy; the power distribution network planning strategy is input into the second intelligent agent model, a power distribution network control instruction is obtained through processing, the power distribution network is adjusted based on the power distribution network control instruction, and the power distribution network control instruction represents an instruction for controlling charging and discharging behaviors of energy storage equipment of the power distribution network. According to the invention, the problems of line overload and poor stability during photovoltaic access to the power distribution network in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of distribution network planning or other related fields. Specifically, it relates to a method, device, computer program product, and electronic device for adjusting a distribution network. Background Art

[0002] With the accelerating promotion of the clean energy transformation, distributed photovoltaic power generation has been widely applied and widely connected due to its green and environmental protection characteristics. However, this transformation has brought significant changes to the operating characteristics of the distribution network. Since the output of photovoltaic power is affected by the intensity of sunlight, the output of photovoltaic power reaches its peak at noon, while the user load is relatively low at this time, which is likely to cause a special phenomenon of reverse power flow in the distribution network, that is, power is fed back from the distribution network to the main grid, forming the so-called "reverse overload". In contrast, during the peak load period at night, when the output of photovoltaic power drops significantly, if the user load surges, "forward overload" will occur, that is, the power transmitted from the main grid to the distribution network is too large, putting pressure on the line carrying capacity. In addition, seasonal factors such as the increase in winter electric heating load will superimpose additional peak loads on the original peak load, exacerbating the risk of line overload.

[0003] The distribution network planning strategy in the related art solves the above problems by means of line capacity expansion and transformation, that is, by increasing the line capacity to improve the carrying capacity of the distribution network. Although this method can alleviate the overload problem to a certain extent, line capacity expansion often means high investment costs. However, due to the volatility of photovoltaic output and the uncertainty of user load, the line may not be fully utilized most of the time, resulting in waste of resources and low economic efficiency. In addition, the power flow calculation of the distribution network is an indispensable part of the planning process, which involves complex non-linear optimization problems. Especially in the context of high penetration of photovoltaic power, the traditional second-order cone relaxation method may not accurately reflect the actual operating state of the distribution network, reducing the reliability and accuracy of planning decisions.

[0004] Aiming at the problems of line overload and poor stability when photovoltaic power is connected to the distribution network in the related art, no effective solution has been proposed yet. Summary of the Invention

[0005] The main purpose of the present application is to provide a method, device, computer program product, and electronic device for adjusting a distribution network to solve the problems of line overload and poor stability when photovoltaic power is connected to the distribution network in the related art.

[0006] To achieve the above object, according to one aspect of the present application, a method for adjusting a distribution network is provided. The method includes: obtaining M state data of the distribution network in a historical time period, performing dimensionality reduction processing on the M state data to obtain M processed state data, where the M state data at least includes: node load data of the distribution network, energy storage data, and photovoltaic installation data, and M is a positive integer; inputting the M processed state data into a first intelligent agent model to process and obtain a distribution network planning strategy, where the distribution network planning strategy is used to indicate the installation strategies of energy storage and photovoltaic in the distribution network; inputting the distribution network planning strategy into a second intelligent agent model to process and obtain a distribution network control instruction, and adjusting the distribution network based on the distribution network control instruction, where the distribution network control instruction represents an instruction for controlling the charge and discharge behavior of the energy storage device of the distribution network.

[0007] Further, the first intelligent agent model includes a policy network, a value network, and a soft value network. Inputting the M processed state data into the first intelligent agent model to process and obtain a distribution network planning strategy includes: obtaining a candidate action space, inputting the M processed state data and the candidate action space into the policy network to process and obtain an action probability distribution, and sampling an initial action vector from the action probability distribution, where the candidate action space includes an action vector composed of N candidate planning strategies, and the action probability distribution refers to the probability distribution of the N candidate planning strategies, and N is a positive integer; inputting the M processed state data and the initial action vector into the value network to process and obtain a first action value parameter; inputting the M processed state data into the soft value network to process and obtain a state value parameter; when the state value parameter and the action value parameter meet the preset requirements, determining the candidate planning strategy corresponding to the initial action vector as the distribution network planning strategy.

[0008] Further, the second intelligent agent model includes multiple sub-intelligent agents, and each sub-intelligent agent includes a first network and a second network. Inputting the distribution network planning strategy into the second intelligent agent model, the processing to obtain the distribution network control instruction includes: operating the distribution network according to the distribution network planning strategy, and after the distribution network operates, obtaining the node status data of K distribution network nodes of the distribution network to obtain K sets of node status data, where each set of node status data includes at least one of the following: node energy storage installation situation, energy storage status, battery health status, node photovoltaic output, and node load information, and K is a positive integer; inputting each set of node status data into the first network, and processing to obtain K local control instructions, where each local control instruction is used to adjust the charge and discharge power of the energy storage device; operating the distribution network based on the K local control instructions, and after the distribution network operates, obtaining the status data of each sub-intelligent agent, where each sub-intelligent agent controls one or more distribution network nodes of the distribution network; inputting the K local control instructions and the status data of each sub-intelligent agent into the second network, processing to obtain the second action value parameter, and updating the K local control instructions based on the second action value parameter to obtain the distribution network control instruction.

[0009] Further, performing dimensionality reduction processing on M state data to obtain M processed state data includes: obtaining a stacked encoder, where the stacked encoder is obtained by stacking Y autoencoders, and each autoencoder includes an encoder and a decoder, and Y is a positive integer; obtaining preset encoding parameters and preset decoding parameters, processing the Y encoders in the stacked encoder based on the preset encoding parameters to obtain Y processed encoders, and processing the Y decoders in the stacked encoder based on the preset decoding parameters to obtain Y processed decoders; performing dimensionality reduction mapping on the M state data by the Y processed encoders to obtain low-dimensional mapping features; and performing feature reconstruction processing on the low-dimensional mapping features by the Y processed decoders to obtain M processed state data.

[0010] Further, the first agent model is trained as follows: Obtain a preset first agent model and an initialized distribution network, where the preset first agent model includes a preset policy network, a preset value network, and a preset soft value network; obtain G groups of initial state data of the initialized distribution network, and output G training adjustment policies by the preset first agent model according to the G groups of initial state data, where G is a positive integer; when the initialized distribution network is operated based on each training adjustment policy, obtain the operation interaction parameters of the initialized distribution network, and adjust the network parameters of the networks in the preset first agent model according to each operation interaction parameter to obtain adjusted network parameters, where the adjusted network parameters include the network parameters of the preset policy network, the network parameters of the preset value network, and the network parameters of the preset soft value network; calculate a reward function according to the adjusted network parameters to obtain a reward function value, and when the reward function value converges, determine the adjusted first agent model corresponding to the adjusted network parameters as the first agent model.

[0011] Further, the second agent model is trained as follows: Obtain a preset second agent model and an initialized distribution network, where the preset second agent model includes a plurality of preset sub-agents; obtain P initial local state data of the initialized distribution network, and output P local training control instructions by the preset second agent model according to the P initial local state data, where P is a positive integer; when the initialized distribution network is operated based on the P local training control instructions, obtain the operation state data of the operated distribution network, and determine the distribution network power loss based on the operation state data; calculate a reward function value according to the distribution network power loss, and when the reward function value does not converge, adjust the preset second agent model to obtain an adjusted second agent model until the reward function value converges, and determine the adjusted second agent model as the second agent model.

[0012] Further, determining the distribution network power loss based on the operation state data includes: obtaining the topological structure of the initialized distribution network, and constructing a topological graph model according to the topological structure, where the nodes in the topological graph model represent the nodes of the initialized distribution network, and the edges in the topological graph model represent the transmission lines of the initialized distribution network; determining the power supply node of the initialized distribution network as the root node, and using the forward algorithm to calculate the node voltage and node current in the topological graph model from the root node, and determining whether the topological graph model satisfies the constraint conditions, where the constraint conditions at least include: the node voltage is within a preset fluctuation range, and the node current is less than or equal to a preset current; when the topological graph model satisfies the constraint conditions, perform a power flow calculation on the operation state data to obtain the distribution network power loss.

[0013] To achieve the above object, according to another aspect of the present application, there is provided an adjustment device for a distribution network. The device includes: an acquisition unit configured to acquire M state data of the distribution network in a historical time period, perform dimensionality reduction processing on the M state data to obtain M processed state data, where the M state data at least includes: node load data, energy storage data, and photovoltaic installation data of the distribution network, and M is a positive integer; a first input unit configured to input the M processed state data into a first intelligent agent model to process and obtain a distribution network planning strategy, where the distribution network planning strategy is used to indicate the installation strategies of energy storage and photovoltaic of the distribution network; a second input unit configured to input the distribution network planning strategy into a second intelligent agent model to process and obtain a distribution network control instruction, and adjust the distribution network based on the distribution network control instruction, where the distribution network control instruction represents an instruction for controlling the charge and discharge behavior of the energy storage device of the distribution network.

[0014] According to another aspect of the embodiments of the present invention, there is also provided a computer storage medium for storing a program, where the program, when running, controls a device where the computer storage medium is located to execute an adjustment method for a distribution network.

[0015] According to another aspect of the embodiments of the present invention, there is also provided an electronic device including one or more processors and a memory; the memory stores computer-readable instructions, and the processor is configured to run the computer-readable instructions, where the computer-readable instructions, when running, execute an adjustment method for a distribution network.

[0016] According to another aspect of the embodiments of the present invention, there is also provided a computer program product including a computer program, where the computer program, when executed by a processor, executes an adjustment method for a distribution network.

[0017] Through this application, the following steps are adopted: obtaining M state data of the distribution network in a historical time period, performing dimensionality reduction processing on the M state data to obtain M processed state data, where the M state data at least includes: node load data, energy storage data, and photovoltaic installation data of the distribution network, and M is a positive integer; inputting the M processed state data into a first intelligent agent model to process and obtain a distribution network planning strategy, where the distribution network planning strategy is used to indicate the installation strategies of energy storage and photovoltaic in the distribution network; inputting the distribution network planning strategy into a second intelligent agent model to process and obtain a distribution network control instruction, and adjusting the distribution network based on the distribution network control instruction, where the distribution network control instruction represents an instruction for controlling the charging and discharging behavior of the energy storage device of the distribution network, solving the problems of line overload and poor stability when photovoltaic is connected to the distribution network in the related art. By obtaining the state data of the distribution network, processing the data by the first intelligent agent model to obtain a distribution network planning strategy, and then processing it by the second intelligent agent model to obtain a distribution network control instruction, and adjusting the distribution network based on the distribution network control instruction, the robustness of the distribution network is enhanced, line overload is alleviated, and the utilization rate and economy of equipment are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0019] Figure 1 is a flowchart of a method for adjusting a distribution network according to an embodiment of this application;

[0020] Figure 2 is a schematic diagram of a method for calculating the power loss of a distribution network according to an embodiment of this application;

[0021] Figure 3 is a flowchart of an optional distribution network adjustment system according to an embodiment of this application;

[0022] Figure 4 is a schematic diagram of a distribution network adjustment device according to an embodiment of this application;

[0023] Figure 5 is a schematic diagram of an electronic device according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0025] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set between this system and relevant users or institutions. Before obtaining relevant information, a request for acquisition needs to be sent to the aforementioned user or institution through the interface, and after receiving the consent information feedback from the aforementioned user or institution, the relevant information can be obtained.

[0028] It should be noted that the information collected in this application is information and data authorized by the user or fully authorized by all parties, and the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application complies with relevant laws, regulations and standards in the relevant regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides a corresponding operation entry for users to choose to authorize the use or refuse to use.

[0029] The present invention will be described below in conjunction with preferred implementation steps. Figure 1 is a flowchart of a method for adjusting a distribution network according to an embodiment of this application, as Figure 1 shown, the method includes the following steps:

[0030] Step S101: Obtain M state data of the distribution network in a historical time period, and perform dimensionality reduction processing on the M state data to obtain M processed state data. Among them, the M state data at least includes: node load data of the distribution network, energy storage data, and photovoltaic installation data, and M is a positive integer.

[0031] To provide more accurate and efficient distribution network adjustment decisions and achieve long-term intelligent planning of the distribution network, thereby improving energy utilization efficiency, first, multiple state data of the distribution network can be obtained. Among them, the state data can include node load data of the distribution network in the past month, energy storage data representing the current energy storage installation status (i.e., the installation location and capacity data of energy storage devices), photovoltaic installation data representing the current photovoltaic installation situation (i.e., the installation location and capacity data of photovoltaic devices), time information, energy storage unit price, photovoltaic unit price and other data. The node load data can reflect the magnitude of power consumption of each node at a specific time point; the energy storage data not only includes the current state of the energy storage device (such as charge and discharge state, remaining power, etc.), but also can include historical charge and discharge records and the health status of the device; the photovoltaic installation data includes the distribution information of photovoltaic devices in the distribution network, including the installation location and rated power of each photovoltaic unit, and the historical record of photovoltaic output.

[0032] Furthermore, since traditional distribution network state data often has high dimensions, it will not only increase the pressure of data storage and transmission, but also make it difficult for subsequent deep learning models to process efficiently. Therefore, a Stacked Autoencoder (SAE) can be used to perform dimensionality reduction processing on the above-obtained original state data to obtain processed state data. Among them, as a deep learning model, SAE can extract key features in the data through unsupervised learning and map the original data to a low-dimensional feature space. That is, SAE can process multi-dimensional data such as node load data, energy storage data, and photovoltaic installation data. Through multiple stackings of autoencoders, deep abstract features in the data can be learned, such as the periodic change of load, the charge and discharge mode of energy storage devices, and the seasonal trend of photovoltaic output. The dimensionality-reduced data can not only retain the important information of the original data, but also reduce the data dimension.

[0033] Step S102: Input the M processed state data into the first intelligent agent model, and process to obtain a distribution network planning strategy, where the distribution network planning strategy is used to indicate the installation strategies of energy storage and photovoltaic in the distribution network.

[0034] Specifically, after obtaining multiple processed state data through the stacked encoder, they can be passed as input to the upper-layer single agent (i.e., the first agent model). This model uses an integrated deep learning model to analyze this information and then generate an energy storage and photovoltaic equipment installation strategy for the distribution network (i.e., obtain a distribution network planning strategy). Among them, the upper-layer single agent adopts the Soft Actor-Critic (SAC) algorithm. By dynamically balancing the selection of actions (i.e., charge and discharge strategies) and the exploration of actions (i.e., trying different charge and discharge schemes in different scenarios), it can make energy storage and photovoltaic installation decisions over a long period (e.g., monthly).

[0035] Step S103: Input the distribution network planning strategy into the second agent model, process it to obtain a distribution network control instruction, and adjust the distribution network based on the distribution network control instruction. Among them, the distribution network control instruction represents an instruction for controlling the charge and discharge behavior of the energy storage device of the distribution network.

[0036] Specifically, after the first agent model outputs the distribution network planning strategy, the lower-layer multi-agent (i.e., the second agent model) can convert the distribution network planning strategy into specific control instructions and dynamically adjust the distribution network accordingly. Among them, the distribution network control instruction can indicate the charge and discharge behavior of the energy storage device at different time points and different nodes, as well as the configuration and working mode of the photovoltaic device, such as an instruction to charge during the peak photovoltaic output and discharge during the peak load.

[0037] While executing the control instruction, the real-time operating state of the distribution network can be continuously monitored and transmitted as feedback information to the second agent model. At this time, the agent can use this feedback information to evaluate the effect of the control instruction and adjust its decision-making strategy according to the actual operating state.

[0038] It should be noted that the decision-making process of this agent can include a policy network (Actor) and a value network (Critic). The policy network is used to generate specific actions, that is, suggestions for the installation location and capacity of energy storage and photovoltaic equipment in the distribution network; while the value network can evaluate the potential value of these actions based on immediate rewards (such as reduced network losses, avoided overloads, and overlimit penalties) and expected long-term benefits (such as equipment investment recovery, operating cost savings), ensuring that the selected strategy can not only meet the immediate grid operating requirements but also improve the economy and reliability of the grid in the long term.

[0039] The adjustment method of the distribution network provided by the embodiment of the present application obtains M state data of the distribution network in a historical time period, performs dimensionality reduction processing on the M state data to obtain M processed state data, where the M state data at least includes: node load data of the distribution network, energy storage data, and photovoltaic installation data, and M is a positive integer; inputs the M processed state data into a first intelligent agent model, and processes to obtain a distribution network planning strategy, where the distribution network planning strategy is used to indicate the installation strategies of energy storage and photovoltaic in the distribution network; inputs the distribution network planning strategy into a second intelligent agent model, processes to obtain a distribution network control instruction, and adjusts the distribution network based on the distribution network control instruction, where the distribution network control instruction represents an instruction for controlling the charging and discharging behavior of the energy storage device of the distribution network, solves the problems of line overload and poor stability when photovoltaic is connected to the distribution network in the related art, obtains the state data of the distribution network, processes the data by the first intelligent agent model to obtain a distribution network planning strategy, then processes it by the second intelligent agent model to obtain a distribution network control instruction, and adjusts the distribution network based on the distribution network control instruction, thereby achieving the effects of enhancing the robustness of the distribution network, alleviating line overload, and improving equipment utilization rate and economy.

[0040] Optionally, in the adjustment method of the distribution network provided by the embodiment of the present application, the first intelligent agent model includes a policy network, a value network, and a soft value network. Inputting the M processed state data into the first intelligent agent model and processing to obtain a distribution network planning strategy includes: obtaining a candidate action space, inputting the M processed state data and the candidate action space into the policy network, processing to obtain an action probability distribution, and sampling an initial action vector from the action probability distribution, where the candidate action space includes an action vector composed of N candidate planning strategies, and the action probability distribution refers to the probability distribution of the N candidate planning strategies, and N is a positive integer; inputting the M processed state data and the initial action vector into the value network, processing to obtain a first action value parameter; inputting the M processed state data into the soft value network, processing to obtain a state value parameter; when the state value parameter and the action value parameter meet the preset requirements, determining the candidate planning strategy corresponding to the initial action vector as the distribution network planning strategy.

[0041] It should be noted that the first intelligent agent model can generate strategies through three networks. The policy network is responsible for generating the distribution of actions, and the intelligent agent selects actions according to this distribution. The soft Q network is used to evaluate the value of actions and helps update the parameters of the policy network and the soft value network by comparing with the prediction of the target soft value network. The soft value network provides an internal value estimate of a state, that is, it is used to evaluate the impact of policy entropy.

[0042] Among them, the policy network π φ(a|s) with parameter φ, associated with a candidate action space. That is, after inputting the processed state data s into this network, it can output the action probability distribution of the action vector a composed of each candidate planning strategy (i.e., the policy network maps the state s to the probability distribution of the action a). Then, an initial action vector a is extracted from the probability distribution by adding a noise term sampled from the standard normal distribution. This action vector a can represent specific installation decisions for energy storage and photovoltaic, such as at which nodes to install equipment and the installed capacity.

[0043] The value network, i.e., the soft Q-network Q θ (s,a) with parameter θ. After inputting the above-mentioned processed state data s and the initial action vector a into this network, it can output the first action value parameter (i.e., output the soft Q value), so as to be able to determine the expected return of the candidate planning strategy in the current state based on this parameter; the soft value network V ψ (s) with parameter ψ. After inputting the processed state data s into this network, it can output the state value parameter, which represents the cumulative sum of the expected future rewards when acting according to a certain strategy under the state data s. It should be noted that the soft Q-network (Soft Q-Network) is mainly used to evaluate the soft Q value of taking the initial action vector a in a given state, and the soft Q value represents the entropy of the strategy, that is, the randomness or uncertainty of taking the action corresponding to the initial action vector a. That is, the soft Q value not only reflects the immediate and future expected returns of taking the initial action vector a under the state data s, but also includes the consideration of uncertainty (entropy). The soft value network can help the intelligent agent understand the total expected return of acting according to the current strategy in each state, taking into account the randomness and exploratory nature of action selection.

[0044] It should be noted that when inputting the processed state data s into different networks, it can be arranged by the following formula: where the state data includes the load data of each node in the past month, the current energy storage installation status, the current photovoltaic installation situation, time information, the unit price of energy storage, and the unit price of photovoltaic, that is, f L is the load feature representation (which can be the dimension-reduced state representation), c ES is the current energy storage installation situation, c PV is the current photovoltaic installation situation, year is the year, month is the month, C ES,unit is the unit price of energy storage, C PV,unit is the unit price of photovoltaic, F L is the dimension of the load feature representation vector, and N is the number of nodes. Each candidate planning strategy in the candidate action space can be represented as an action vector a upper,t ∈A upper which is characterized as a continuous energy storage and photovoltaic installation plan and can be represented by the following formula: a upper,t= [a ES,t ; a PV,t ∈ R 2N , where a ES,t and a PV,t respectively represent the installation schemes of energy storage and photovoltaic. Each element corresponds to a node in the distribution network, and can indicate whether energy storage or photovoltaic equipment is installed at that node and the specific capacity of the equipment.

[0045] Furthermore, since the value network and the soft value network can respectively provide the optimization objective and feedback information for the policy network from the perspectives of action value and state value (i.e., the average long-term return that the agent can obtain in this state), helping it learn the action probability policy that can maximize the long-term return in a given state. Therefore, after obtaining the initial action vector, the value network and the soft value network can further guide the parameter optimization of the policy network by evaluating the value of the initial action vector, enabling the agent to make better behavior decisions, that is, by comparing the action value parameter and the state value parameter, determining whether the candidate planning policy corresponding to the initial action vector meets the preset requirements. When the state value parameter and the action value parameter meet the preset requirements, that is, if a higher long-term benefit is expected by adopting this policy, the candidate planning policy corresponding to the initial action vector is determined as the distribution network planning policy. Through the interaction of the policy network, the value network and the soft value network in this embodiment, the upper-layer single agent can output economic and robust energy storage and photovoltaic installation decisions according to the current state of the distribution network, so as to improve the operation efficiency and reliability of the distribution network, significantly improve the decision-making efficiency, and improve the robustness and adaptability of the planning scheme in actual operation.

[0046] Optionally, in the distribution network adjustment method provided in the embodiments of the present application, the second agent model includes multiple sub-agents, and each sub-agent includes a first network and a second network. Inputting the distribution network planning policy into the second agent model and processing to obtain the distribution network control instruction includes: operating the distribution network according to the distribution network planning policy, and after the distribution network runs, obtaining the node state data of K distribution network nodes of the distribution network to obtain K sets of node state data, where each set of node state data includes at least one of the following: node energy storage installation situation, energy storage state, battery health status, node photovoltaic output, and node load information, and K is a positive integer; inputting each set of node state data into the first network and processing to obtain K local control instructions, where each local control instruction is used to adjust the charging and discharging power of the energy storage device; operating the distribution network based on the K local control instructions, and after the distribution network runs, obtaining the state data of each sub-agent, where each sub-agent controls one or more distribution network nodes of the distribution network; inputting the K local control instructions and the state data of each sub-agent into the second network, processing to obtain the second action value parameter, and updating the K local control instructions based on the second action value parameter to obtain the distribution network control instruction.

[0047] Specifically, after the distribution network planning strategy is output by the first agent model, the distribution network environment can be updated by the second agent model, that is, the distribution network planning strategy generated by the upper layer is converted into an immediate power grid control instruction to regulate the charging and discharging behavior of the energy storage device. First, the distribution network is operated based on this strategy, and then the node state data of all distribution network nodes in the distribution network after operation is obtained. Among them, each group of node state data may include the energy storage installation situation of the node, the state of charge (SOC) of the energy storage, the health status of the battery, the photovoltaic output of the node, and the load information of the node.

[0048] Since this model includes multiple sub-agents, each sub-agent includes a first network (i.e., the Actor network) and a second network (i.e., the Critic network), and a sub-agent is deployed for each node in the distribution network, or multiple nodes are controlled by one sub-agent, so that each sub-agent can control the charging and discharging behavior of the node energy storage of one or more nodes. The Actor network has parameters φ i , and the node state data of one or more distribution network nodes responsible for each sub-agent (i.e., the local state s i of sub-agent i) can be input into this network, so as to output the corresponding local control instruction a i of the distribution network node. Among them, each local control instruction includes a suggestion for adjusting the charging and discharging power of the energy storage device on the corresponding node, and then the energy storage device is regulated immediately to cut peaks and fill valleys and balance the node power to ensure that the distribution network operates in an optimal state. It should be noted that the above local state of sub-agent i may include the energy storage installation situation of the node, the energy storage SOC, the battery health status, the node photovoltaic output, the node load, and the time information, and can be represented by the following formula:

[0049] s lower,i,t = [c ES,i ; SOC i,t ; P PV,i,t ; P D,i,t ; hour; SOH i,t ∈ R 6 ;

[0050] Among them, c ES,i represents the energy storage installation situation of the node, SOC i,t represents the energy storage SOC, P PV,i,t represents the battery health status, P D,i,t represents the node photovoltaic output, SOH i,t represents the node load, and hour represents the time information (hour). The local control instruction a i characterizes the charging and discharging power of the energy storage and can be expressed as: alower,i,t = P ES,i,t ∈ [-P dis,max,i , P ch,max,i , where P ES,i,t is the charging and discharging power of energy storage i. A negative value indicates discharging, and a positive value indicates charging. P dis,max,i and P ch,max,i are the maximum discharging power and maximum charging power of energy storage i, respectively.

[0051] The Critic network Q θi (s, a), with parameters θ i , operates the distribution network based on local control instructions in the distribution network. For example, during peak load, the energy storage device is commanded to discharge to relieve the burden on the main cable; when the photovoltaic output is excessive, the energy storage device is charged with the excess electrical energy to prevent reverse power overload. Thus, the state data of the distribution network nodes corresponding to all sub - agents after operating the distribution network can be obtained again. Then, all the above - mentioned state data s = [s1,..., s Nagent and all local control instructions a = [a1,..., a Nagent are input into this network, and then the second action value parameter (i.e., the Q - value of the action) is output. This parameter can evaluate the impact of local control instructions on the operation effect of the distribution network, including network loss reduction, battery loss cost, voltage deviation, and line overload penalty, etc. Based on the second action value parameter, the sub - agent can adjust its local control instructions to further optimize the charging and discharging strategy of the energy storage device, and then obtain the distribution network control instructions, where Nagent represents the number of sub - agents. In this embodiment, through multiple rounds of distribution network operation and control instruction optimization, the sub - agents gradually learn more efficient and economical energy storage regulation strategies, can effectively cope with the random fluctuations of photovoltaic output and load, improve the power balance ability of the distribution network at different times, reduce network loss and equipment loss, and enhance the overall economy and operation reliability of the power grid.

[0052] Optionally, in the distribution network adjustment method provided in the embodiments of the present application, the M state data are processed by dimensionality reduction to obtain M processed state data, including: obtaining a stacked encoder, where the stacked encoder is composed of Y auto - encoders stacked together, and each auto - encoder includes an encoder and a decoder, and Y is a positive integer; obtaining preset encoding parameters and preset decoding parameters, processing the Y encoders in the stacked encoder based on the preset encoding parameters to obtain Y processed encoders, processing the Y decoders in the stacked encoder based on the preset decoding parameters to obtain Y processed decoders; mapping the M state data to low - dimensional features by the Y processed encoders, and performing feature reconstruction processing on the low - dimensional features by the Y processed decoders to obtain M processed state data.

[0053] Since the state data of the upper-layer agent contains load data for the past month and has a high dimension, to improve the processing efficiency of the model, Stacked Autoencoders (SAE) can be used to reduce the dimension of the historical data, so as to reduce the data dimension and extract key features, providing a more efficient and informative input for the agent model. Since the stacked autoencoder is composed of multiple autoencoders (AE), each autoencoder consists of an encoder and a decoder. The encoder can map high-dimensional input data to low-dimensional feature representations, while the decoder reconstructs the original high-dimensional data from the low-dimensional feature representations. Finally, through unsupervised learning and minimizing the reconstruction error, the effective feature representations of the data are learned and the dimension is reduced.

[0054] For example, when the SAE is an SAE including multiple hidden layers, for example, the input layer can receive the original state data Lpast; the hidden layer 1 (encoder) is expressed as: h (1) = σ(W (1) L past + b (1) ); the hidden layer 2 (encoder) is expressed as: h (2) = σ(W (2) h (1) + b (2) ); the hidden layer 3 (encoder) is expressed as: f L = σ(W (3) h (2) + b (3) ), f L is the load feature after dimension reduction, that is, the low-dimensional mapping feature is obtained. This feature representation not only has a smaller data size but also contains the core information of the distribution network state data, such as load trends, periodicity of photovoltaic output, and changes in the state of charge of energy storage devices. The hidden layer 4 (decoder) is expressed as: h (4) = σ(W (4) f L + b (4) ); the hidden layer 5 (decoder) is expressed as: h (5) = σ(W (5) h (4) + b (5) ); the output layer (decoder) is expressed as: is the reconstructed load data, that is, the processed state data. Among them, h (i) , W (i) , b (i)They are respectively the output vector, weight matrix, and bias vector of the i-th layer (i.e., the preset encoding parameters and preset decoding parameters). It should be noted that although the reconstructed data may have slight differences from the original data, it significantly reduces the dimension of the data while retaining the key features, which is beneficial to the subsequent decision-making of the model.

[0055] When training the above stacked encoder, the layer-by-layer training method, backpropagation algorithm, and Adam optimizer can be used for training. That is, first train the encoder and decoder of the first autoencoder, then fix its parameters, and then train the encoder and decoder of the second autoencoder, and so on until all autoencoders are trained. By training the encoder layer by layer, the deep features of the data can be gradually extracted to achieve more effective dimensionality reduction. In addition, the mean squared error can also be used as the loss function to minimize the reconstruction error, which can be expressed as: where N month is the number of months for single encoding. In this embodiment, by training the stacked encoder, the mapping process from low-dimensional features to high-dimensional data is optimized, and the accuracy of feature reconstruction is improved, so as to ensure that the data after dimensionality reduction can still retain the key information of the original data, significantly reduce the dimension of the input data of the agent model, improve the efficiency and speed of model training, provide more accurate and decision-making valuable information for the agent model, and improve the adaptability and generalization ability of the model to new data. Even when facing the uncertainty of photovoltaic output and load, reasonable decisions can be made.

[0056] Optionally, in the method for adjusting the distribution network provided in the embodiment of the present application, the first agent model is trained as follows: obtain a preset first agent model and obtain the initialized distribution network, where the preset first agent model includes a preset policy network, a preset value network, and a preset soft value network; obtain G groups of initial state data of the initialized distribution network, and output G training adjustment strategies by the preset first agent model according to the G groups of initial state data, where G is a positive integer; when the initialized distribution network runs based on each training adjustment strategy, obtain the operation interaction parameters of the initialized distribution network, and adjust the network parameters of the networks in the preset first agent model according to each operation interaction parameter to obtain the adjusted network parameters, where the adjusted network parameters include the network parameters of the preset policy network, the network parameters of the preset value network, and the network parameters of the preset soft value network; calculate the reward function according to the adjusted network parameters to obtain the reward function value, and when the reward function value converges, determine the adjusted first agent model corresponding to the adjusted network parameters as the first agent model.

[0057] Specifically, before processing data based on the first agent model, the model needs to be trained so that it can generate a planning strategy that maximizes the economic benefits and operational stability of the distribution network while considering the randomness of photovoltaic power output and load. First, a preset first agent model needs to be obtained. This model consists of a preset policy network, a preset value network, and a preset soft value network. These three networks are responsible for generating action strategies, evaluating policy values, and calculating state values respectively. At the same time, initialize the network parameters of these three networks, that is, randomly initialize the policy network parameter φ, the value network parameter θ, the soft value network parameter ψ, and the target value network parameter And initialize the experience replay buffer D.

[0058] Furthermore, in each complete cycle of interaction between the agent and the environment (that is, the cycle from one environmental state to reaching a certain termination state), it is also necessary to initialize the distribution network environment, and then obtain the initial state data s upper,0 And the training adjustment strategy output by the preset first agent model according to the initial state data, and each strategy corresponds to a set of initial state data.

[0059] After obtaining the training adjustment strategy, the agent model can interact with the distribution network based on these strategies, that is, apply each training adjustment strategy in the simulation environment, observe the reaction of the distribution network, and obtain operation interaction parameters such as network loss, equipment cost, voltage deviation, line overload penalty, etc. Then use these parameters as feedback for immediate rewards or long-term values to evaluate the strategy effect generated by the agent model and guide the parameter adjustment of the model.

[0060] Specifically, at each time step t, select an action through the action distribution generated by the policy network, and interact with the environment to obtain a reward and the next state, and then store this data in the experience replay buffer D. Furthermore, randomly sample a batch of samples from the experience replay buffer D, and update the policy network, value network, and soft value network based on this batch. Among them, the network parameters of the policy network can be updated by minimizing the KL divergence:

[0061] In the formula, DKL is the divergence, and Z θ (s t ) is the partition function used for normalization. At this time, the output of the policy network is expressed as: a t =f φ (ε t ; s t ), in the formula, ε t is the noise vector sampled from the standard normal distribution, and f φ represents a differentiable function; the network parameters of the value network can be updated using the Bellman equation: Wherein, is the soft value, α is the temperature coefficient that can control the importance of the policy entropy, is the parameter of the target policy value network, and the target updated policy network can be expressed as: The parameters of the target value network can be updated by using soft update (i.e., exponential moving average): Wherein, τ is a soft update coefficient. The network parameters of the soft value network can also be updated by minimizing the loss function: Wherein, D is the experience replay buffer, π φ (a t |s t ) is the policy network.

[0062] After obtaining the adjusted network parameters, the reward function reflecting the performance of the agent model in the distribution network planning can be calculated according to the adjusted network parameters, and the reward function value can be obtained. Among them, the reward function can be expressed as: r upper,t = -(C install + C penalty ), C install is the installation cost, which can be calculated by the following formula: Among them, C penalty is the penalty term, which can be calculated by the following formula: C penalty = w voltage × V penalty + w overload × I overload , wherein, a ES,i,t , a PV,i,t are the capacities of installing energy storage and photovoltaic at node i respectively, V is the node set, w voltage , w overload are the penalty weights of voltage deviation and line overload respectively, V penalty , I overload are the voltage deviation penalty and current overload penalty respectively, V i,t is the voltage of node i, V ref is the voltage set value, E is the edge set, I j,t is the current of edge j, I co,j is the rated current of edge j.

[0063] During the training process, if it is detected that the reward function value converges, it indicates that the agent model can continuously produce high-quality decisions in multiple different operating scenarios of the distribution network while maintaining the stability and consistency of the policy. At this time, the adjusted first agent model corresponding to the adjusted network parameters can be determined as the first agent model. Through the adaptive training mechanism of deep reinforcement learning in this embodiment, the first agent model can continuously adjust its policy network and learn more efficient and economical distribution network planning strategies, so as to still maintain the stable operation of the power grid, improve the generalization performance, and ensure the wide applicability and robustness of the planning strategy under the condition of high uncertainty of photovoltaic output and load.

[0064] Optionally, in the distribution network adjustment method provided in the embodiments of the present application, the second agent model is trained in the following manner: obtain a preset second agent model and an initialized distribution network, where the preset second agent model includes multiple preset sub-agents; obtain P initial local state data of the initialized distribution network, and the preset second agent model outputs P local training control instructions according to the P initial local state data, where P is a positive integer; when the initialized distribution network runs based on the P local training control instructions, obtain the operation state data of the running distribution network, and determine the distribution network power loss based on the operation state data; calculate the reward function value according to the distribution network power loss, and when the reward function value does not converge, adjust the preset second agent model to obtain the adjusted second agent model until the reward function value converges, and determine the adjusted second agent model as the second agent model.

[0065] Specifically, in order to enable the second agent model to generate optimal control instructions for the local state of the power grid and maintain the stable operation and economy of the power grid under the condition of frequent fluctuations in photovoltaic power generation and load demand, the above model also needs to be trained. At this time, a centralized training-distributed execution framework can be adopted, that is, in the training stage, the Critic network of each sub-agent can access the state and action information of all agents (that is, allowing agents to share information), for example, evaluate the joint actions of all agents through a centralized Critic network to learn a more comprehensive Q-value estimate. In the deployment and execution stage, each sub-agent only needs to make decisions based on its own local state and select actions using the Actor network without communicating with other agents.

[0066] First, it is necessary to initialize the Actor network parameters and Critic network parameters of each sub-agent i, and initialize the target network parameters, the experience replay buffer Di, and the distribution network environment. Then, obtain multiple initial local state data s from the initialized distribution network lower,0, this data can characterize multiple different local operating conditions. This data is passed as input to the sub-agent in the preset second intelligent agent model. Each sub-agent, based on its own local information, uses the Actor network to generate a local training control instruction, that is, a charging and discharging power adjustment recommendation for the energy storage device.

[0067] Furthermore, each local training control instruction is executed in a simulated or actual distribution network. At each time step, the sub-agent adjusts its control instruction according to the latest local state data and interacts with the environment. Then, the operating state data of the distribution network after operation is collected, and based on this data, the network loss and energy efficiency index of the distribution network are calculated.

[0068] That is, at each time step, each sub-agent selects an action and adds noise, jointly interacts with the environment with the action, obtains a reward and the next state, and stores the experience (s lower,t , a lower,t , r lower,t , s lower,t+1 ) into the experience replay buffer. By randomly sampling a batch of samples from the experience replay buffer Di, and then updating the Actor network and Critic network of each agent based on the samples: The reward function value is calculated based on the network loss and possible battery loss, voltage violation, line overload penalty, etc., and can comprehensively reflect the overall operating performance of the second intelligent agent model under the current parameter settings. At this time, the network loss and battery loss of the distribution network can be calculated, and the reward function can be calculated by adding the upper-layer violation penalty: r lower,i,t = -(w loss ×P loss,t + w deg ×C deg,i,t ) + r upper , where w loss , w deg are the weight coefficients of the network loss and battery loss respectively, P loss,t is the active network loss, C deg,i,t is the battery loss cost, which is proportional to the charge and discharge amount.

[0069] Finally, the model is adjusted according to the reward function value. If the reward function value does not meet the convergence condition, it indicates that there is still room for optimization in the current parameters of the second agent model. For example, a batch of samples can be continuously drawn from the experience replay buffer Di as training data. Based on these data, the current and target network parameters are compared, and optimization algorithms such as gradient descent or Adam are used to update the parameters of the Actor network and the Critic network until the reward function value converges, that is, the agent model can generate control instructions that meet the expected goals under multiple different local states and operating conditions. Finally, the adjusted second agent model corresponding to the convergence of the reward function value is determined as the second agent model.

[0070] It should be noted that when training the first agent model and the second agent model, the lower-layer multi-agent (i.e., the second agent model) can be pre-trained first. Pre-training can enable the second agent model to learn basic energy storage control strategies, such as discharging during peak load, charging during valley load, and charging when photovoltaic output is excessive. In actual operation, when the distribution network environment in which the second agent model is located changes, for example, the strategy output by the first agent model is to install new energy storage or photovoltaic, the second agent model can be trained incrementally at this time. Incremental training can enable the second agent model to adapt to the new environment and optimize its control strategy. Incremental training can use the pre-trained model parameters as the initial values and fine-tune them with new environmental data.

[0071] In this embodiment, by training the second agent model, the second agent model continuously learns and adapts to the uncertainties of photovoltaic power generation and load demand, enabling the sub-agent to quickly respond to changes in the grid operating state, accurately adjust the charging and discharging behavior of energy storage devices, achieve peak shaving and valley filling, improve the robustness of the model and its adaptability to complex grid environments, improve the energy efficiency and stability of the grid. In addition, through the centralized training-distributed execution mechanism, the second agent model can achieve the optimal allocation of global resources, thereby reducing the long-term operating cost and improving the economic benefits.

[0072] Optionally, in the method for adjusting a distribution network provided in the embodiments of the present application, determining the distribution network power loss based on the operating state data includes: obtaining the topological structure of the initialized distribution network, and constructing a topological graph model according to the topological structure, where the nodes in the topological graph model represent the nodes of the initialized distribution network, and the edges in the topological graph model represent the transmission lines of the initialized distribution network; determining the power source node of the initialized distribution network as the root node, and using the forward sweep algorithm to calculate the node voltages and node currents in the topological graph model from the root node, and determining whether the topological graph model satisfies the constraint conditions, where the constraint conditions at least include: the node voltage is within a preset fluctuation range, and the node current is less than or equal to a preset current; in the case where the topological graph model satisfies the constraint conditions, performing a power flow calculation on the operating state data to obtain the distribution network power loss.

[0073] Specifically, the distribution network power loss refers to the energy loss caused by factors such as resistance and reactance during the power transmission process in the power grid. By calculating the power loss in the operating state data, the efficiency of the control command in reducing energy loss and improving system utilization can be evaluated. When calculating the distribution network power loss, in order to quickly evaluate the operating state (voltage, current, power, etc.) of the distribution network and calculate the reward function, the DFS (Depth-First Search) power flow calculation method can be used at this time. First, the topological structure of the initialized distribution network can be obtained, that is, the connection relationship between the nodes (such as substations, load points, distributed power sources, etc.) in the distribution network and the transmission lines therebetween can be obtained, and then a topological graph model is constructed based on the obtained topological structure, where the nodes represent the actual nodes in the distribution network, and the edges represent the transmission lines. The topological graph model not only intuitively shows the structure of the distribution network but also provides a basic framework for the subsequent forward sweep algorithm.

[0074] Furthermore, since the power source node is the starting point for obtaining electric energy from the main power grid and provides the necessary input power for the distribution network, the power source node in the model can be determined as the root node at this time, and then the forward sweep algorithm (i.e., the forward-backward substitution method) is used to start from the root node and calculate the voltage and current of each node in the topological graph model along the topological structure of the distribution network, that is, based on the radial structure of the distribution network, starting from the power source node, advancing node by node along the transmission lines, and calculating the node voltage and line current level by level. According to the input power, output power, line parameters of the node, and the voltage and current information of the previous node, the state of each node is dynamically updated.

[0075] After calculating the node voltages and currents, it is also necessary to determine whether the topological graph model meets the predetermined constraint conditions. For example, it is necessary to determine whether the node voltages are maintained within a preset fluctuation range and whether the node currents do not exceed the preset maximum current limit. When the topological graph model meets the constraint conditions, power flow calculation can be further performed to evaluate the actual operating state of the distribution network. That is, the distribution network topology structure (node connection relationship, line parameters), load data of each node, PV output data of each node, charge and discharge power of the energy storage at each node (provided by the lower-level agent), and the voltage of the power source node are used as input parameters for the power flow calculation program, and then the voltage amplitude and phase angle of each node, the current and power of each line, and the total active power loss of the distribution network are output.

[0076] Figure 2 is a schematic diagram of the calculation method for the power loss of the distribution network provided by the embodiments of the present application. As Figure 2 shown, first, a topological graph model of the distribution network is established, and the grid-connected point voltage is set. Among them, the grid-connected point represents the connection point between the distribution network and the upper-level power grid. The DFS algorithm starts from the grid-connected point, sets the voltage of this point as the reference value (which can be set to 1), and uses it as the initialization condition for the algorithm execution.

[0077] Before starting the DFS search, it is necessary to initialize the information, that is, initialize all information related to the search, and update the load of each street at time t, that is, update the load conditions of each node in the distribution network to the current time step t according to the time series. Further, by default, u = 1 is the grid-connected point, that is, set the variable u = grid-connected point, which is the starting node for the DFS algorithm search. Start the DFS algorithm, starting from the grid-connected point u, and perform a depth-first traversal along the topological structure of the power grid. At this time, each line can be followed to sequentially visit its downstream nodes until the end of the power grid or the search termination condition is met.

[0078] When accessing each node, the forward-backward substitution method is used to calculate the node voltage and the power flow of the line. That is, it is judged whether the u node still has successor nodes. If there is still a successor node v, the voltage of the v point can be updated at this time, and the function is called itself to perform a power flow calculation on the v point. Then, data such as line power and node power are updated, and then the load power is superimposed on the injection power of this node. If the u node has no successor nodes, the load power is directly superimposed on the injection power of this node, and then the power flow result is output. During the calculation process, it is checked whether the power grid state converges, that is, whether the voltage and power flow stabilize within an acceptable range, and the network loss situation of the distribution network is recorded and updated. Since the distribution network operates periodically, a 24-hour period can be used at this time. When the time step t exceeds 24 hours, it indicates that the DFS algorithm has completed the traversal and calculation of all states of the distribution network in one day. At this time, the cumulative total network loss data can be output to obtain the distribution network loss. In this embodiment, by calculating the network loss, the optimization effect of the distribution network operation control strategy can be intuitively fed back, guiding the intelligent agent model to make more economical and efficient decisions to reduce energy waste and improve system energy efficiency, so as to maintain the operation of the distribution network within a safe voltage and current range, avoid the occurrence of faults such as overload and short circuit, and ensure the stable operation of the power grid and the quality of power supply.

[0079] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0080] The embodiment of the present application also provides an adjustment system for a distribution network. Figure 3 It is a flowchart of an optional adjustment system for a distribution network provided according to the embodiment of the present application. As Figure 3 shown, the system includes: a strategic decision-making layer, a tactical execution layer, a physical coupling layer, and a physical system layer.

[0081] Among them, the strategic decision-making layer is used to obtain the load data of each node in the past month, the current installation status of energy storage and photovoltaic, time information (such as year, month), and the unit prices of energy storage and photovoltaic by the first intelligent agent model. Then, the stacked encoder is used to reduce the dimension of the load adjustment and extract the spatio-temporal adjustment at the same time. Then, the first intelligent agent model outputs a long-term planning decision, that is, plans the installation location and capacity of energy storage and photovoltaic on a relatively long time scale (such as year / month), so as to achieve annual / monthly scheduling and energy storage photovoltaic site selection.

[0082] The tactical execution layer is used for the processing of the second agent model, which may include an energy storage control unit and a photovoltaic coordination unit. The energy storage control unit can achieve SOC optimization and charge-discharge decision-making, and the photovoltaic coordination unit can achieve power output tracking and fluctuation suppression, that is, execute the long-term planning decisions output by the first agent model, obtain the data after operation, and input the data into the second agent model. The physical coupling layer is used for training the second agent model, that is, calculating the power grid loss of the distribution network through the power flow calculation engine and the forward-backward substitution method, and then using multi-time scale simulation and the power grid loss of the distribution network for model training, so as to help the second agent model achieve predictions from hourly to annual simulations. The physical system layer can run the distribution network control instructions output by the second agent model, that is, perform the update and adjustment of the urban distribution network environment, such as adjusting the node voltage and current, while paying attention to the topological dynamic changes, and can also use the real-time monitoring system for equipment health monitoring to achieve data closed-loop feedback.

[0083] Through the learning of the first agent model and the second agent model in this embodiment, it is possible to achieve forward-looking planning for future distribution network requirements, optimize the installation location and capacity of energy storage and photovoltaic, ensure that the power grid operates in the best state, and lay a foundation for the sustainable development of the distribution network.

[0084] The embodiment of the present application also provides an adjustment device for a distribution network. It should be noted that the adjustment device for the distribution network in the embodiment of the present application can be used to execute the adjustment method for the distribution network provided in the embodiment of the present application. The following introduces the adjustment device for the distribution network provided in the embodiment of the present application.

[0085] Figure 4 is a schematic diagram of the adjustment device for the distribution network provided in the embodiment of the present application, as Figure 4 shown, the device includes: an acquisition unit 40, a first input unit 41, and a second input unit 42.

[0086] The acquisition unit 40 is configured to acquire M state data of the distribution network in a historical time period, perform dimensionality reduction processing on the M state data to obtain M processed state data, where the M state data at least includes: node load data of the distribution network, energy storage data, and photovoltaic installation data, and M is a positive integer;

[0087] The first input unit 41 is configured to input the M processed state data into the first agent model, and process to obtain a distribution network planning strategy, where the distribution network planning strategy is used to indicate the installation strategy of the energy storage and photovoltaic of the distribution network;

[0088] The second input unit 42 is configured to input the distribution network planning strategy into the second agent model, process to obtain a distribution network control instruction, and adjust the distribution network based on the distribution network control instruction, where the distribution network control instruction represents an instruction for controlling the charge and discharge behavior of the energy storage device of the distribution network.

[0089] The adjustment device for the distribution network provided by the embodiment of the present application obtains M state data of the distribution network in a historical time period through an acquisition unit 40, performs dimensionality reduction processing on the M state data to obtain M processed state data, where the M state data at least includes: node load data of the distribution network, energy storage data, and photovoltaic installation data, and M is a positive integer; a first input unit 41 inputs the M processed state data into a first intelligent agent model, and processes to obtain a distribution network planning strategy, where the distribution network planning strategy is used to indicate the installation strategies of the energy storage and photovoltaic of the distribution network; a second input unit 42 inputs the distribution network planning strategy into a second intelligent agent model, processes to obtain a distribution network control instruction, and adjusts the distribution network based on the distribution network control instruction, where the distribution network control instruction represents an instruction for controlling the charge and discharge behavior of the energy storage device of the distribution network, solves the problems of line overload and poor stability when photovoltaic is connected to the distribution network in the related art, obtains the state data of the distribution network, processes the data by the first intelligent agent model to obtain a distribution network planning strategy, then processes it by the second intelligent agent model to obtain a distribution network control instruction, and adjusts the distribution network based on the distribution network control instruction, thereby achieving the effects of enhancing the robustness of the distribution network, alleviating line overload, and improving equipment utilization rate and economy.

[0090] Optionally, in the adjustment device for the distribution network provided by the embodiment of the present application, the first input unit 41 includes: a first acquisition module, configured to acquire a candidate action space, input the M processed state data and the candidate action space into a policy network, process to obtain an action probability distribution, and sample an initial action vector from the action probability distribution, where the candidate action space includes an action vector composed of N candidate planning strategies, and the action probability distribution refers to the probability distribution of the N candidate planning strategies, and N is a positive integer; a first input module, configured to input the M processed state data and the initial action vector into a value network, and process to obtain a first action value parameter; a second input module, configured to input the M processed state data into a soft value network, and process to obtain a state value parameter; a first determination module, configured to determine the candidate planning strategy corresponding to the initial action vector as the distribution network planning strategy when the state value parameter and the action value parameter meet preset requirements.

[0091] Optionally, in the adjustment device for a distribution network provided in the embodiments of the present application, the second input unit 42 includes: a first operation module, configured to operate the distribution network according to a distribution network planning strategy, and after the distribution network operates, obtain node status data of K distribution network nodes of the distribution network, to obtain K sets of node status data, where each set of node status data includes at least one of the following: node energy storage installation situation, energy storage status, battery health status, node photovoltaic power output, and node load information, and K is a positive integer; a third input module, configured to input each set of node status data into a first network, and process to obtain K local control instructions, where each local control instruction is used to adjust the charge and discharge power of an energy storage device; a second operation module, configured to operate the distribution network based on the K local control instructions, and after the distribution network operates, obtain the status data of each sub-agent, where each sub-agent controls one or more distribution network nodes of the distribution network; a fourth input module, configured to input the K local control instructions and the status data of each sub-agent into a second network, process to obtain a second action value parameter, and update the K local control instructions based on the second action value parameter to obtain a distribution network control instruction.

[0092] Optionally, in the adjustment device for a distribution network provided in the embodiments of the present application, the acquisition unit 40 includes: a second acquisition module, configured to acquire a stacked encoder, where the stacked encoder is obtained by stacking Y autoencoders, and each autoencoder includes an encoder and a decoder, and Y is a positive integer; a third acquisition module, configured to acquire a preset encoding parameter and a preset decoding parameter, process the Y encoders in the stacked encoder based on the preset encoding parameter to obtain Y processed encoders, and process the Y decoders in the stacked encoder based on the preset decoding parameter to obtain Y processed decoders; a mapping module, configured to perform dimensionality reduction mapping on M status data by the Y processed encoders to obtain low-dimensional mapping features; a processing module, configured to perform feature reconstruction processing on the low-dimensional mapping features by the Y processed decoders to obtain M processed status data.

[0093] Optionally, in the adjustment device of the distribution network provided in the embodiments of the present application, the first input unit 41 includes: a fourth acquisition module, configured to acquire a preset first agent model and an initialized distribution network, where the preset first agent model includes a preset policy network, a preset value network, and a preset soft value network; a fifth acquisition module, configured to acquire G sets of initial state data of the initialized distribution network, and output G training adjustment policies by the preset first agent model according to the G sets of initial state data, where G is a positive integer; a sixth acquisition module, configured to acquire the operation interaction parameters of the initialized distribution network when the initialized distribution network is operated based on each training adjustment policy, and adjust the network parameters of the network in the preset first agent model according to each operation interaction parameter to obtain adjusted network parameters, where the adjusted network parameters include the network parameters of the preset policy network, the network parameters of the preset value network, and the network parameters of the preset soft value network; a first calculation module, configured to calculate a reward function according to the adjusted network parameters to obtain a reward function value, and when the reward function value converges, determine the adjusted first agent model corresponding to the adjusted network parameters as the first agent model.

[0094] Optionally, in the adjustment device of the distribution network provided in the embodiments of the present application, the second input unit 42 includes: a seventh acquisition module, configured to acquire a preset second agent model and an initialized distribution network, where the preset second agent model includes a plurality of preset sub-agent models; an eighth acquisition module, configured to acquire P sets of initial local state data of the initialized distribution network, and output P local training control instructions by the preset second agent model according to the P sets of initial local state data, where P is a positive integer; a ninth acquisition module, configured to acquire the operation state data of the operated distribution network when the initialized distribution network is operated based on the P local training control instructions, and determine the distribution network power loss based on the operation state data; a second calculation module, configured to calculate a reward function value according to the distribution network power loss, and when the reward function value does not converge, adjust the preset second agent model to obtain an adjusted second agent model until the reward function value converges, and determine the adjusted second agent model as the second agent model.

[0095] Optionally, in the adjustment device of the distribution network provided in the embodiments of the present application, the second input unit 42 includes: a tenth acquisition module, configured to acquire the topological structure of the initialized distribution network, and construct a topological graph model according to the topological structure, where the nodes in the topological graph model represent the nodes of the initialized distribution network, and the nodes in the topological graph model represent the transmission lines of the initialized distribution network; a second determination module, configured to determine the power supply node of the initialized distribution network as the root node, and use the forward algorithm to calculate the node voltage and node current in the topological graph model, and determine whether the topological graph model meets the constraint conditions, where the constraint conditions at least include: the node voltage is within a preset fluctuation range, and the node current is less than or equal to a preset current; a third calculation module, configured to perform a power flow calculation on the operation state data when the topological graph model meets the constraint conditions, so as to obtain the distribution network power loss.

[0096] The above-mentioned adjustment device of the distribution network includes a processor and a memory. The above-mentioned acquisition unit 40, the first input unit 41, the second input unit 42, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions.

[0097] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problems of line overload and poor stability in the related art when photovoltaic power is connected to the distribution network can be solved.

[0098] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0099] The embodiments of the present invention provide a computer storage medium for storing a program, where when the program runs, it controls a device where the computer storage medium is located to execute an adjustment method for a distribution network.

[0100] Figure 5 is a schematic diagram of an electronic device provided according to an embodiment of the present application. As Figure 5 shown, the embodiments of the present invention provide an electronic device. The electronic device 50 includes a processor, a memory, and a program stored on the memory and executable on the processor. The processor is used to run computer-readable instructions, where when the computer-readable instructions run, they execute an adjustment method for a distribution network. The devices herein can be servers, PCs, PADs, mobile phones, etc.

[0101] The present application also provides a computer program product, including a computer program, where when the computer program is executed by a processor, it implements the steps of an adjustment method for a distribution network in various embodiments of the present application.

[0102] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0103] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0104] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0106] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0107] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0108] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0109] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0110] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for adjusting a distribution network, characterized in that: include: Acquire M state data of the distribution network in a historical time period, perform dimensionality reduction processing on the M state data, and obtain M processed state data, wherein the M state data at least include: node load data, energy storage data, and photovoltaic installation data of the distribution network, and M is a positive integer; Inputting the M processed state data into a first intelligent agent model to obtain a distribution network planning strategy, wherein the distribution network planning strategy is used to indicate an energy storage and photovoltaic installation strategy of the distribution network; The distribution network planning strategy is input into a second intelligent agent model, and a distribution network control instruction is obtained through processing, and the distribution network is adjusted based on the distribution network control instruction, wherein the distribution network control instruction represents an instruction for controlling the charging and discharging behavior of an energy storage device of the distribution network.

2. The method according to claim 1, characterized in that: The first agent model includes a strategy network, a value network and a soft value network. The M processed state data are input into the first agent model, and the distribution network planning strategy obtained by processing includes: Acquire a candidate action space, input the M processed state data and the candidate action space into the strategy network, process to obtain an action probability distribution, and sample an initial action vector from the action probability distribution, wherein the candidate action space includes action vectors composed of N candidate planning strategies, the action probability distribution refers to the probability distribution of the N candidate planning strategies, and N is a positive integer; Inputting the M processed state data and the initial action vector into the value-taking network to obtain a first action value parameter; Inputting the M processed state data into the soft value network to obtain state value parameters; When the state value parameter and the action value parameter meet preset requirements, the candidate planning strategy corresponding to the initial action vector is determined as the distribution network planning strategy.

3. The method according to claim 1, characterized in that The second agent model includes a plurality of sub-agents, each of which includes a first network and a second network. The distribution network planning strategy is input into the second agent model, and the distribution network control instructions obtained by processing include: The distribution network is operated according to the distribution network planning strategy, and after the distribution network is operated, node status data of K distribution network nodes of the distribution network are obtained to obtain K groups of node status data, wherein each group of node status data includes at least one of the following: node energy storage installation status, energy storage status, battery health status, node photovoltaic output and node load information, and K is a positive integer; Inputting each group of node status data into the first network, and processing to obtain K local control instructions, wherein each local control instruction is used to adjust the charging and discharging power of the energy storage device; Running the distribution network based on the K local control instructions, and obtaining status data of each sub-agent after the distribution network is running, wherein each sub-agent controls one or more distribution network nodes of the distribution network; The K local control instructions and the status data of each sub-agent are input into the second network, processed to obtain a second action value parameter, and the K local control instructions are updated based on the second action value parameter to obtain the distribution network control instruction.

4. The method according to claim 1, characterized in that: The M state data are subjected to dimensionality reduction processing to obtain M processed state data including: Obtain a stacked encoder, wherein the stacked encoder is obtained by stacking Y autoencoders, each autoencoder includes an encoder and a decoder, and Y is a positive integer; Obtaining preset encoding parameters and preset decoding parameters, processing Y encoders in the stacked encoder based on the preset encoding parameters to obtain Y processed encoders, and processing Y decoders in the stacked encoder based on the preset decoding parameters to obtain Y processed decoders; The Y processed encoders perform dimensionality reduction mapping on the M state data to obtain low-dimensional mapping features; The Y processed decoders perform feature reconstruction processing on the low-dimensional mapping features to obtain the M processed state data.

5. The method according to claim 1, characterized in that: The first agent model is trained in the following way: Obtaining a preset first agent model and obtaining an initialized distribution network, wherein the preset first agent model includes a preset strategy network, a preset value network, and a preset soft value network; Acquire G groups of initial state data of the initialized distribution network, and output G training adjustment strategies according to the G groups of initial state data by the preset first agent model, where G is a positive integer; In the case of operating the initialized distribution network based on each training adjustment strategy, obtaining the operation interaction parameters of the initialized distribution network, and adjusting the network parameters of the network in the preset first agent model according to each operation interaction parameter to obtain the adjusted network parameters, wherein the adjusted network parameters include the network parameters of the preset strategy network, the network parameters of the preset value network, and the network parameters of the preset soft value network; A reward function is calculated according to the adjusted network parameters to obtain a reward function value, and when the reward function value converges, an adjusted first agent model corresponding to the adjusted network parameters is determined as the first agent model.

6. The method according to claim 1, characterized in that The second agent model is trained by: Acquire a preset second agent model and acquire an initialized power distribution network, wherein the preset second agent model includes a plurality of preset sub-agents; Acquire P initial local state data of the initialized distribution network, and output P local training control instructions according to the P initial local state data by the preset second intelligent agent model, where P is a positive integer; In the case of operating the initialized distribution network based on the P local training control instructions, obtaining operation status data of the distribution network after operation, and determining the network loss of the distribution network based on the operation status data; The reward function value is calculated according to the distribution network loss. When the reward function value does not converge, the preset second intelligent agent model is adjusted to obtain an adjusted second intelligent agent model until the reward function value converges, and the adjusted second intelligent agent model is determined as the second intelligent agent model.

7. The method according to claim 6, characterized in that Determining the power distribution network loss based on the operating status data includes: Acquire the topological structure of the initialized distribution network, and construct a topological graph model according to the topological structure, wherein the nodes in the topological graph model represent the nodes of the initialized distribution network, and the nodes in the topological graph model represent the transmission lines of the initialized distribution network; Determine the power supply node of the initialized distribution network as the root node, calculate the node voltage and node current in the topology graph model from the root node using a forward algorithm, and determine whether the topology graph model meets the constraint conditions, wherein the constraint conditions at least include: the node voltage is within a preset fluctuation range, and the node current is less than or equal to a preset current; When the topology model satisfies the constraint conditions, power flow calculation is performed on the operating status data to obtain the power distribution network loss.

8. A distribution network adjustment device, characterized in that: include: An acquisition unit is used to acquire M state data of the distribution network in a historical time period, and perform dimensionality reduction processing on the M state data to obtain M processed state data, wherein the M state data at least include: node load data, energy storage data and photovoltaic installation data of the distribution network, and M is a positive integer; A first input unit is used to input the M processed state data into a first intelligent agent model to obtain a distribution network planning strategy, wherein the distribution network planning strategy is used to indicate an energy storage and photovoltaic installation strategy of the distribution network; The second input unit is used to input the distribution network planning strategy into a second intelligent agent model, process it to obtain a distribution network control instruction, and adjust the distribution network based on the distribution network control instruction, wherein the distribution network control instruction represents an instruction for controlling the charging and discharging behavior of the energy storage device of the distribution network.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for adjusting the distribution network according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the distribution network adjustment method described in any one of claims 1 to 7.