Scheduling method and device for power distribution network resources and electronic equipment

Through the adversarial game model generation of energy storage charging and discharging strategies, the problem of low robust scheduling capabilities of high-proportion photovoltaic distribution networks is solved, and the stability and economicality of the power grid is improved, and it is suitable for distribution network resource scheduling.

CN120237679APending Publication Date: 2025-07-01STATE GRID BEIJING ELECTRIC POWER CO +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510401301.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

High-proportion photovoltaic distribution networks have problems of low robust scheduling capabilities and poor stability. The existing scheduling methods are difficult to effectively deal with the uncertainty of photovoltaic output and load fluctuations, resulting in power grid imbalance, voltage fluctuations and low resource utilization efficiency.

Method used

Adversarial game model is adopted, by obtaining the operating data and meteorological data of the distribution network, and using disturbance generators and energy storage optimization discriminators to generate energy storage charge and discharge power adjustment strategies, coordinate distributed resources, and realize training and scheduling decisions of the adversarial game model.

Benefits of technology

It significantly improves the operating stability and robustness of a high proportion of photovoltaic distribution network, ensures that the power grid can operate safely, stably and economically in the face of uncertainty, and improves the utilization rate of distributed resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120237679A_ABST
    Figure CN120237679A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution network resource scheduling method and apparatus, and an electronic device. Relates to the power grid dispatching field or other related fields, and the method comprises the following steps: obtaining operation data and meteorological data of a power distribution network in a preset time period, the operation data at least comprising one of photovoltaic data, load value data and charge state data; the operation data and the meteorological data are input into the confrontation game model, scheduling action data are output, and the scheduling action data represent an energy storage charging and discharging power adjustment strategy of the power distribution network; and scheduling adjustment is carried out on the resources of the power distribution network based on the scheduling action data, and the scheduling adjustment mode at least comprises one of the following modes: adjusting the energy storage charging and discharging power of the power distribution network, and coordinating and scheduling the distributed resources of the power distribution network. According to the invention, the technical problems of low robust scheduling capability and poor stability of a high-proportion photovoltaic power distribution network in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of power grid dispatching or other related fields. Specifically, it relates to a method and device for dispatching distribution network resources and an electronic device. Background Art

[0002] Under the background of the current global energy transformation, renewable energy is widely connected. For example, the large-scale deployment of distributed photovoltaic power sources is reshaping the operation mode of the distribution network. With the continuous impact of climate change and the pursuit of clean and low-carbon energy, the grid connection of a high proportion of renewable energy has become an inevitable trend. The popularization of distributed photovoltaic power sources not only helps with energy self-sufficiency but also promotes the clean transformation of the energy structure. However, how to effectively manage the intermittency on the source side and the dynamics on the load side to ensure the stability and efficiency of the distribution network has become a new challenge.

[0003] The uncertainty of the photovoltaic power output mainly stems from its high dependence on light conditions, such as cloud occlusion and weather changes. The randomness of these natural factors makes it impossible to accurately predict the photovoltaic output power. Similarly, the prediction of the user-side load has become increasingly complex. Due to the diversity of user behavior, the suddenness of special events, especially the popularization of smart appliances and electric vehicles, the load demand shows obvious volatility and unpredictability. The superposition of uncertainties at both the source and load ends poses a major challenge to the safe operation of the distribution network and the effectiveness of the dispatching strategy. Traditional distribution network dispatching methods, such as strategies based on deterministic optimization, often rely on accurate prediction models, especially the estimation of photovoltaic output and load demand. However, the existing dispatching methods have insufficient analysis of the fluctuation characteristics of photovoltaic output and load, resulting in defects in the robustness and adaptability of the dispatching strategy, exacerbating the difficulty of dispatching and the complexity of uncertainty management. For example, the dispatching plan cannot effectively balance supply and demand due to prediction deviations, leading to problems such as power imbalance, voltage fluctuation, and line overload in the power grid, affecting the stability of the power grid and the consumption of renewable energy. At the same time, overly conservative dispatching strategies limit the full utilization of renewable energy, thereby increasing the operating cost.

[0004] To address the source-load uncertainty, reinforcement learning (RL), as a method of learning decision-making strategies by interacting with the environment, has been applied to distribution network dispatching in order to obtain an optimal dispatching strategy through the learning process. However, the existing methods have limited capabilities in dealing with the rapid fluctuations and mutations of the source and load in the distribution network, and have low training efficiency and insufficient generalization ability in a high-dimensional state space, making it difficult to ensure the performance of the dispatching strategy in the face of diverse and extreme uncertainties.

[0005] Aiming at the technical problems of low robust dispatching ability and poor stability in the high-proportion photovoltaic distribution network in the related art, no effective solution has been proposed yet. Summary of the Invention

[0006] The main object of the present application is to provide a dispatching method, device and electronic device for distribution network resources, so as to solve the technical problems of low robust dispatching ability and poor stability in the related art of high-proportion photovoltaic distribution networks.

[0007] To achieve the above object, according to one aspect of the present application, a dispatching method for distribution network resources is provided. The method includes: obtaining the operation data and meteorological data of the distribution network within a preset time period, where the operation data at least includes one of the following: photovoltaic data, load value data, and state of charge data; inputting the operation data and meteorological data into an adversarial game model, and outputting dispatching action data, where the dispatching action data represents the energy storage charge and discharge power adjustment strategy of the distribution network; and performing dispatching adjustment on the resources of the distribution network based on the dispatching action data, where the dispatching adjustment methods at least include one of the following: adjusting the energy storage charge and discharge power of the distribution network, and coordinating the dispatching of distributed resources of the distribution network.

[0008] Further, the adversarial game model includes a perturbation generator and an energy storage optimization discriminator. The perturbation generator includes a feature extraction layer, a time series modeling layer, and a fully connected output layer. Inputting the operation data and meteorological data into the adversarial game model and outputting the dispatching action data includes: determining a state data matrix according to the operation data and meteorological data; extracting spatial features in the state data matrix by the feature extraction layer and outputting spatial features, where the feature extraction layer includes N sub-feature extraction layers, and each sub-feature extraction layer is associated with a first weight matrix and a first bias parameter vector, and N is a positive integer; performing time series modeling on the spatial features by the time series modeling layer and outputting time series features, where the time series modeling layer includes Y sub-time series modeling layers, and each sub-time series modeling layer is associated with a second weight matrix and a second bias parameter vector, and Y is a positive integer; calculating the perturbation parameter by the fully connected output layer according to a preset activation function, inputting the perturbation parameter into the energy storage optimization discriminator, and outputting the dispatching action data.

[0009] Further, the energy storage optimization discriminator includes a hidden layer and an output layer. Inputting the perturbation parameter into the energy storage optimization discriminator and outputting the dispatching action data includes: splicing the state data matrix and the perturbation parameter to obtain an input vector; calculating the input vector based on the network parameters of the hidden layer to obtain a hidden layer data matrix; obtaining K candidate dispatching action data within a preset time period, and calculating the action value function vectors of the K candidate dispatching action data by the output layer according to the hidden layer data matrix, where K is a positive integer; screening the K candidate dispatching action data, and obtaining the candidate dispatching action data with the largest action value function vector, and determining the candidate dispatching action data with the largest action value function vector as the dispatching action data.

[0010] Further, the adversarial game model is trained in the following manner: Obtain M types of historical operation data and historical meteorological data of the distribution network in a historical time period, and obtain the historical state data of the distribution network in the historical time period, where at least one of the M types of historical operation data of the distribution network includes: distribution network topology data, historical load data, historical photovoltaic processing data, and energy storage system data, and the historical state data is used for the state of the distribution network during operation in the historical time period, and M is a positive integer; preprocess each type of historical operation data of the distribution network respectively to obtain M types of processed operation data, and preprocess the historical meteorological data to obtain processed historical meteorological data; perform per-unit value conversion on the M types of processed operation data to obtain M types of converted operation data, and use the historical state data, the M types of processed operation data, and the processed meteorological data to form a sample set, where the sample set includes sample input data and sample output data, the sample input data is composed of the M types of processed operation data and the processed meteorological data, and the sample output data is composed of the historical state data; use the sample set to train a preset adversarial game model to obtain the adversarial game model, where the preset adversarial game model includes a preset perturbation generator and a preset energy storage optimization discriminator.

[0011] Further, using the historical state data, the M types of processed operation data, and the processed meteorological data to form a sample set includes: Extract the processed historical load data and the processed photovoltaic processing data from the M types of processed operation data to obtain first operation data; perform discrete wavelet transform on the first operation data using a wavelet basis function to obtain component vectors, where the wavelet basis function is used to capture the time-frequency characteristics in the first operation data, the component vectors include a first component vector and a second component vector, the frequency of the first component vector is less than the frequency of the second component vector, the first component vector corresponds to a trend characteristic, and the second component vector corresponds to a volatility characteristic; obtain a sparse autoencoder, and use the sparse autoencoder to extract the component characteristics of the component vectors to obtain schedulable component characteristics and uncontrollable fluctuation component characteristics; construct sample input data according to the schedulable component characteristics, the uncontrollable fluctuation component characteristics, the processed historical load data, the processed photovoltaic processing data, and the processed meteorological data, and construct sample output data according to the historical state data.

[0012] Further, training the preset adversarial game model using the sample set to obtain the adversarial game model includes: the preset perturbation generator outputs a perturbation distribution according to the sample input data in the sample set, samples a historical perturbation distribution from the historical state data, and determines an interpolation perturbation distribution according to the perturbation distribution and the historical perturbation distribution; obtaining the penalty parameter of the distribution network, and calculating the first loss function of the preset energy storage optimization discriminator according to the perturbation distribution, the interpolation perturbation distribution, the historical state data, and the penalty parameter to obtain the first loss function; obtaining the discount factor and network parameters, calculating the gradient target value according to the M types of converted operation data, the discount factor, and the network parameters in the sample set, and determining the second loss function of the preset energy storage optimization discriminator according to the discount factor to obtain the second loss function; calculating the sum of the first loss function and the second loss function to obtain the discriminator loss function; obtaining the preset optimizer, and using the preset optimizer and the gradient ascent algorithm to determine the minimum value of the discriminator loss function to obtain the first minimum loss function value; determining the empirical parameters corresponding to the first minimum loss function value, determining the discriminator network parameters according to the empirical parameters corresponding to the first minimum loss function value, and adjusting the model parameters of the preset energy storage optimization discriminator based on the discriminator network parameters to obtain the energy storage optimization discriminator in the adversarial game model.

[0013] Further, training the preset adversarial game model using the sample set to obtain the adversarial game model includes: obtaining the perturbation distribution output by the preset perturbation generator; the preset perturbation generator calculates a perturbation score according to the sample set, and determines the loss function of the preset perturbation generator based on the perturbation score and the perturbation distribution to obtain the generator loss function; obtaining the preset optimizer, and using the preset optimizer and the gradient descent algorithm to determine the minimum value of the generator loss function to obtain the second minimum loss function value; determining the empirical parameters corresponding to the second minimum loss function value, determining the generator network parameters according to the empirical parameters corresponding to the second minimum loss function value, and adjusting the model parameters of the preset perturbation generator based on the generator network parameters to obtain the perturbation generator in the adversarial game model.

[0014] To achieve the above object, according to another aspect of the present application, there is provided a scheduling device for distribution network resources. The device includes: an acquisition unit, configured to acquire the operation data and meteorological data of the distribution network within a preset time period, where the operation data includes at least one of the following: photovoltaic data, load value data, and state of charge data; an input unit, configured to input the operation data and meteorological data into the adversarial game model and output scheduling action data, where the scheduling action data represents the energy storage charge and discharge power adjustment strategy of the distribution network; an adjustment unit, configured to perform scheduling adjustment on the resources of the distribution network based on the scheduling action data, where the scheduling adjustment method includes at least one of the following: adjusting the energy storage charge and discharge power of the distribution network, and coordinately scheduling the distributed resources of the distribution network.

[0015] According to another aspect of the embodiments of the present invention, there is also provided a computer storage medium for storing a program, wherein when the program runs, it controls a device where the computer storage medium is located to execute a scheduling method for distribution network resources.

[0016] According to another aspect of the embodiments of the present invention, there is also provided an electronic device including one or more processors and a memory; computer-readable instructions are stored in the memory, and the processor is used to run the computer-readable instructions, wherein when the computer-readable instructions run, they execute a scheduling method for distribution network resources.

[0017] According to another aspect of the embodiments of the present invention, there is also provided a computer program product including a computer program, and when the computer program is executed by a processor, it executes a scheduling method for distribution network resources.

[0018] Through the present application, the following steps are adopted: obtaining operation data and meteorological data of the distribution network within a preset time period, wherein the operation data at least includes one of the following: photovoltaic data, load value data, and state of charge data; inputting the operation data and meteorological data into an adversarial game model to output scheduling action data, wherein the scheduling action data represents an energy storage charge and discharge power adjustment strategy of the distribution network; based on the scheduling action data, scheduling and adjusting the resources of the distribution network, wherein the scheduling and adjustment methods at least include one of the following: adjusting the energy storage charge and discharge power of the distribution network, and coordinately scheduling distributed resources of the distribution network, which solves the technical problems of low robust scheduling ability and poor stability in the related high-proportion photovoltaic distribution network. By obtaining the operation data and meteorological data of the distribution network, processing them using the adversarial game model to generate scheduling action data, and performing scheduling and adjustment of resources based on this data, the effect of significantly improving the operation stability and robustness of the high-proportion photovoltaic distribution network is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings that form a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0020] Figure 1 is a flowchart of a scheduling method for distribution network resources provided according to an embodiment of the present application;

[0021] Figure 2 is a schematic diagram of an optional scheduling method for distribution network resources provided according to an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of a scheduling device for distribution network resources provided according to an embodiment of the present application;

[0023] Figure 4It is a schematic diagram of an electronic device provided according to an embodiment of the present application. Detailed implementation manners

[0024] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. For example, an interface is provided between the present system and relevant users or institutions. Before obtaining relevant information, a request for obtaining needs to be sent to the aforementioned users or institutions through the interface, and after receiving the consent information feedback from the aforementioned users or institutions, the relevant information can be obtained.

[0028] It should be noted that the information collected in the present application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure, application and other processing of the relevant data all comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse to use.

[0029] The present invention will be described below in combination with preferred implementation steps.Figure 1 is a flowchart of a scheduling method for distribution network resources provided according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0030] Step S101, obtain the operation data and meteorological data of the distribution network within a preset time period. Among them, the operation data includes at least one of the following: photovoltaic data, load value data, and state of charge data.

[0031] Specifically, the photovoltaic data in the operation data can be historical photovoltaic processing data, the load value data can refer to historical load data, and the state of charge data can include distribution network topology data and energy storage system data. Among them, the distribution network topology data includes node information (number, type, geographical location, voltage level), line information (start and end nodes, line length, line parameters: resistance, reactance, susceptance), transformer information (start and end nodes, turns ratio, impedance), etc.; the historical load data can include the historical load curves of each load node, which can be typical daily load curves or multi-year historical load data; the historical photovoltaic processing data can include the power generation data of all photovoltaic power plants in the distribution network and the historical output curves of each photovoltaic power plant, such as the hourly power generation, light intensity curve, power plant location information, power plant type (such as monocrystalline silicon, polycrystalline silicon, thin-film photovoltaic, etc.), and the operation status of the power plant (such as maintenance, fault, etc.); the energy storage system data can include parameters such as the type, rated capacity, charge and discharge power, efficiency, and initial state of charge (SOC) of the energy storage system. The meteorological data can include meteorological data such as light intensity, temperature, wind speed, and humidity. For example, data such as sunlight irradiation angle, cloud cover, and air pollution index.

[0032] After obtaining the above data, it is necessary to clean the collected data, remove outliers and missing values, and process data noise to ensure data quality; the data can also be converted into a time series format suitable for machine learning, such as data interpolation and smoothing processing, to capture the continuity and dynamic characteristics of source-load data; the data can also be subjected to feature extraction and transformation, such as using wavelet transform or Fourier transform to process photovoltaic output and load data, and extract features reflecting volatility and trend.

[0033] Step S102, input the operation data and meteorological data into the adversarial game model, and output scheduling action data, where the scheduling action data represents the energy storage charge and discharge power adjustment strategy of the distribution network.

[0034] Specifically, the adversarial game model can refer to a model composed of a photovoltaic / load disturbance generator - energy storage optimization discriminator. The disturbance generator is constructed by a hybrid model of a long short-term memory network and a convolutional neural network (i.e., the LSTM-CNN model, where LSTM (Long Short-Term Memory) is the long short-term memory network and CNN (Convolutional Neural Network) is the convolutional neural network). It can generate challenging disturbance samples by analyzing the source-load fluctuation patterns in historical data, that is, more effectively simulate the spatio-temporal correlation disturbances of photovoltaic output and load, and then simulate the uncertain scenarios that may be encountered in actual operation to test and train the robustness of the energy storage optimization discriminator; the energy storage optimization discriminator uses a hybrid model of deep Q-learning and a multi-layer perceptron (i.e., DQN-MLP, where DQN: Deep Q-Networks, MLP: Multi-Layer Perceptron) to learn the current state of the distribution network and output the optimal energy storage charge and discharge power adjustment strategy according to the disturbance samples, so as to achieve efficient robust scheduling decisions.

[0035] It should be noted that before inputting the above operation data and meteorological data into the model, a two-way feature decoupling mechanism can be adopted to decompose the data into schedulable components (which can reflect trends and predictability) and uncontrollable fluctuation components (which can capture randomness and mutations), which helps the disturbance generator to focus more on simulating the uncontrollable fluctuation components, while the energy storage optimization discriminator can more effectively learn the robust scheduling strategy for these fluctuations. The decoupled features are further fused and used as the input of the discriminator, providing a more comprehensive state representation and helping the model to make more accurate decisions.

[0036] By processing the above data, the adversarial game model can obtain scheduling action data, that is, the energy storage charge and discharge power adjustment strategy, and then it can be applied to the distribution network simulation environment or the actual power grid to perform the charge and discharge operations of energy storage devices, providing strong decision support for the distributed resource scheduling of high-proportion renewable energy access to the distribution network.

[0037] Step S103, perform scheduling adjustments on the resources of the distribution network based on the scheduling action data, where the scheduling adjustment methods include at least one of the following: adjusting the energy storage charge and discharge power of the distribution network, and coordinately scheduling the distributed resources of the distribution network.

[0038] Specifically, after obtaining the scheduling action data, the resources of the distribution network can be scheduled and adjusted based on this scheduling action data to ensure the safe, stable, and economic operation of the distribution network in the face of source-load uncertainty. Among them, the distribution network can adjust the charging and discharging power of energy storage devices to cope with the uncertain disturbances simulated by the disturbance generator. For example, the charging and discharging power of each energy storage unit can be determined according to the scheduling action data. If the predicted load demand is higher than the supply, the energy stored in the energy storage system is released to fill the supply-demand gap. Conversely, if there is an oversupply, the excess energy can be stored for subsequent demand. While adjusting the charging and discharging power of the energy storage, the distribution of power among various energy storage units can be optimized to ensure the efficient use of resources; maintain the stability of the state of charge and avoid the situation of too low or too high SOC to ensure that the energy storage unit can respond at any time when needed and at the same time extend the service life of the equipment.

[0039] In addition, the distributed resources in the distribution network (such as distributed generation, demand-side response resources) can also be adjusted. For example, according to the photovoltaic output prediction and load demand, the power generation power of distributed power sources can be dynamically adjusted to ensure the balance between supply and demand: when the lighting conditions are good, renewable energy is preferentially utilized. When the lighting is insufficient, other distributed power sources can be coordinated through optimal scheduling to maintain the stability of the power grid; the load demand of users in the power grid can be adjusted through incentive or control mechanisms to avoid overload and improve the resource utilization efficiency; optimize the operating status of lines and transformers in the network, such as adjusting the tap position of transformers and controlling the opening and closing of lines, to optimize the power flow, reduce line losses, avoid overload or voltage over-limit, and improve the operating performance and economic benefits of the distribution network.

[0040] It should be noted that after the implementation of the scheduling, the operating status of the distribution network can be monitored, new operating data can be collected, and then the effectiveness of the scheduling strategy can be evaluated. Then, according to the feedback results, online learning and parameter tuning of the model can be carried out to improve the robustness and economy of the scheduling strategy in this way.

[0041] The scheduling method for distribution network resources provided by the embodiments of the present application obtains the operation data and meteorological data of the distribution network within a preset time period. The operation data at least includes one of the following: photovoltaic data, load value data, and state of charge data. The operation data and meteorological data are input into an adversarial game model, and scheduling action data is output, where the scheduling action data represents the energy storage charge and discharge power adjustment strategy of the distribution network. Based on the scheduling action data, the resources of the distribution network are scheduled and adjusted. The scheduling adjustment methods at least include one of the following: adjusting the energy storage charge and discharge power of the distribution network, and coordinating the scheduling of distributed resources of the distribution network. This solves the technical problems of low robust scheduling ability and poor stability in the related high-proportion photovoltaic distribution network. By obtaining the operation data and meteorological data of the distribution network, processing them using the adversarial game model, generating scheduling action data, and performing resource scheduling and adjustment based on this data, the stability and robustness of the high-proportion photovoltaic distribution network operation are significantly improved.

[0042] Optionally, in the scheduling method for distribution network resources provided by the embodiments of the present application, the adversarial game model includes a perturbation generator and an energy storage optimization discriminator. The perturbation generator includes a feature extraction layer, a temporal modeling layer, and a fully connected output layer. Outputting the scheduling action data by inputting the operation data and meteorological data into the adversarial game model includes: determining a state data matrix according to the operation data and meteorological data; extracting spatial features in the state data matrix by the feature extraction layer and outputting the spatial features, where the feature extraction layer includes N sub-feature extraction layers, and each sub-feature extraction layer is associated with a first weight matrix and a first bias parameter vector, and N is a positive integer; performing temporal modeling on the spatial features by the temporal modeling layer and outputting temporal features, where the temporal modeling layer includes Y sub-temporal modeling layers, and each sub-temporal modeling layer is associated with a second weight matrix and a second bias parameter vector, and Y is a positive integer; calculating the temporal features by the fully connected output layer according to a preset activation function, outputting perturbation parameters, and inputting the perturbation parameters into the energy storage optimization discriminator to output the scheduling action data.

[0043] Specifically, after obtaining the operation data and meteorological data, in order to determine the scheduling action data of the distribution network based on these data, they can be input into the adversarial game model, and the perturbation generator in the adversarial game model processes them. First, the input layer in the perturbation generator can sort the above data in a certain sequence and construct it into a state data matrix. For example, after receiving the preprocessed historical photovoltaic processing data historical load data L hist and meteorological data W hist they are constructed into a state data matrix in the way of "number of samples, time step, number of features", where the time step can be 24 hours.

[0044] Furthermore, a Convolutional Neural Network (CNN) is adopted as the feature extraction layer to extract spatial structure features from the state data matrix. That is, by sliding the convolutional kernel over the data, local patterns and spatial correlations in the source-load data are detected and learned. Among them, the feature extraction layer contains multiple sub-feature extraction layers (i.e., multiple CNNs), and each layer is associated with a convolutional kernel, which is composed of a first weight matrix and a first bias parameter vector. The stacking of multiple sub-feature extraction layers can learn spatial features at different scales and levels, providing rich representation information for subsequent time series modeling. For example, when the feature extraction layer is a three-layer stacked 1D-CNN layer, the first layer of CNN includes 64 convolutional kernels, the convolutional kernel size is 3, the stride is 1, and the activation function is the ReLU function; the second layer of CNN includes 64 convolutional kernels, the convolutional kernel size is 5, the stride is 1, the activation function is the ReLU function, followed by a batch normalization layer; the third layer of CNN includes 64 convolutional kernels, the convolutional kernel size is 7, the stride is 1, the activation function is the ReLU function, followed by a batch normalization layer and a max pooling layer (pooling size is 2). Each layer can be expressed as:

[0045]

[0046] Among them, are the first weight matrix and the first bias parameter vector of the i-th layer of CNN respectively, represents the i-th layer of CNN.

[0047] Furthermore, after the spatial features are extracted by the above-mentioned feature extraction layer, the time series modeling layer can perform time series modeling on the spatial features to output time series features. Among them, this layer is composed of a Long Short-Term Memory (LSTM) network, which can model the time series dependence of data. By identifying context information, it can make more accurate predictions and simulations when facing photovoltaic and load data with historical dependence. The time series modeling layer consists of multiple sub-time series modeling layers, and each layer has its corresponding second weight matrix and second bias parameter vector, which can capture the dynamic changes and long-term trends in the data based on these parameters. For example, when the time series modeling layer uses a two-layer LSTM network to perform time series modeling on the spatial features extracted by the CNN, the first layer of LSTM includes 128 LSTM units, the activation function is the tanh function, and the recurrent activation function is the sigmoid function; the second layer of LSTM includes 128 LSTM units, the activation function is the tanh function, and the recurrent activation function is the sigmoid function. Each layer can be expressed as:

[0048]

[0049] Among them, They are respectively the second weight matrix and the second bias parameter vector of the i-th layer of LSTM; They are respectively the hidden state and the cell state of the i-th layer of LSTM at time step t.

[0050] The fully connected output layer, as the last layer of the perturbation generator, can calculate the temporal features output by the temporal modeling layer according to a preset activation function, and finally output perturbation parameters (which can also refer to perturbation samples) that characterize the simulated perturbations of the photovoltaic output and the load value. Among them, these perturbations can refer to a sudden drop in light intensity, a sharp rise in temperature, or any other scenario that may cause source-load uncertainty, and can provide a series of challenging scenarios for the energy storage optimization discriminator, enabling it to learn to generate optimal scheduling decisions under different perturbation environments. The fully connected output layer can be expressed as:

[0051]

[0052] Wherein, They are respectively the weight matrix and the bias parameter vector of the i-th layer of the fully connected output layer, and δ is the perturbation parameter.

[0053] After the perturbation generator outputs the perturbation parameters, they can be input into the energy storage optimization discriminator, and then the scheduling action data is output. Among them, the energy storage optimization discriminator can receive the state data and perturbation parameters of the current distribution network, calculate through the DQN-MLP structure, and finally output the scheduling action data, such as outputting the energy storage charge and discharge power adjustment strategy and other resource scheduling instructions. In this embodiment, the energy storage optimization discriminator converts the operation data and meteorological data into a state data matrix, and performs multi-level feature extraction and temporal modeling through a deep learning model, and finally outputs the perturbation parameters and the scheduling action data. Even when the photovoltaic output and the load value fluctuate violently, it can maintain the power balance and voltage stability of the power grid and reduce the risk of line overload. In addition, the dynamic game between the perturbation generator and the optimization discriminator can balance the economy and security of the scheduling. By optimizing the scheduling strategy, the operation cost can be minimized, the utilization rate of distributed resources can be improved, and the maximum economic benefit can be pursued on the premise of ensuring the reliability of the system.

[0054] Optionally, in the scheduling method of distribution network resources provided in the embodiments of the present application, the energy storage optimization discriminator includes a hidden layer and an output layer. After inputting the perturbation parameter into the energy storage optimization discriminator, the output scheduling action data includes: concatenating the state data matrix and the perturbation parameter to obtain an input vector; calculating the input vector based on the network parameters of the hidden layer to obtain a hidden layer data matrix; obtaining K candidate scheduling action data within a preset time period, and the output layer calculates the action value function vectors of the K candidate scheduling action data according to the hidden layer data matrix, where K is a positive integer; screening the K candidate scheduling action data to obtain the candidate scheduling action data with the largest action value function vector, and determining the candidate scheduling action data with the largest action value function vector as the scheduling action data.

[0055] Specifically, after the perturbation generator in the adversarial game model outputs the perturbation parameter, since the state data matrix contains the real-time operating state of the distribution network, such as photovoltaic output, load value, energy storage state of charge (SOC), etc., and the perturbation parameter is the uncertain scenario simulated by the perturbation generator, including sudden changes in light intensity, sudden increase or decrease in user load, etc. At this time, the state data matrix s and the perturbation parameter δ can be concatenated to obtain the input vector x = [s, δ], so as to completely describe the current operating environment of the distribution network.

[0056] Subsequently, multi-layer processing is performed through a deep learning model (such as an energy storage optimization discriminator based on DQN) to generate the optimal scheduling action data. First, the above input vector is input into the hidden layer of the energy storage optimization discriminator, thereby outputting a hidden layer data matrix. The hidden layer can be composed of multiple fully connected neural networks, and each layer has its specific network parameters (for example, including a weight matrix and a bias vector). Then, linear and non-linear transformations are performed on the input vector based on these network parameters, and non-linearity is added through an activation function (such as the ReLU function) to extract high-level abstract features useful for decision-making, thereby realizing the conversion of the input complex state data into a hidden layer data matrix. For example, when using three fully connected layers as the hidden layer of the Q network, the first layer includes 512 neurons, and the activation function is the ReLU function; the second layer includes 512 neurons, and the activation function is the ReLU function; the third layer includes 512 neurons, and the activation function is the ReLU function, which can be expressed as:

[0057] h (1) =ReLU(W (1) x + b (1) );

[0058] h (2) =ReLU(W (2) h (1) + b (2) );

[0059] h (3) = ReLU(W (3) h (2) + b (3) );

[0060] where W (i) and b (i) are the weight matrix and bias vector of the i-th layer, respectively.

[0061] Furthermore, after the hidden layer outputs the hidden layer data matrix, the output layer can be used to process it, that is, based on the hidden layer data matrix, value evaluation is performed on a plurality of preset candidate scheduling action data. Each candidate scheduling action data contains various possible strategies for scheduling distributed resources in the distribution network, such as adjustment of energy storage charge and discharge power, increase and decrease of distributed generation, change of network topology, etc., and can be expressed as A = {P ESS,1 , P ESS,2 ,..., P ESS,n}, where P ESS,i is the charge and discharge power of the i-th energy storage, and n is the total number of energy storages. The output layer uses predefined network parameters to calculate each candidate action to evaluate the possible benefits or costs it may bring in the current state, that is, the action value function vector. For example, when the output layer is processed in the structure of a fully connected layer, the Q value of each action a (that is, each candidate scheduling action data) above can be output, and can be expressed as: q = W (4) h (3) + b (4) , where the activation function uses a linear function, q is the Q value vector; W (4) and b (4) are the weight matrix and bias vector of the output layer.

[0062] Since the calculation of the action value function vector is based on the principle of Deep Q-Networks (DQN), the potential value of each action is measured by the size of the Q value. An action with a high Q value means that executing this action in the current state can bring better long-term benefits. Therefore, based on the obtained action value function vector, the candidate scheduling action data can be screened, and the action corresponding to the maximum value in the action value function vector can be determined as the optimal scheduling action data. In this embodiment, by splicing the state data matrix and the perturbation parameter, and then calculating the action value function of the candidate scheduling action data, it is possible to select the most economical scheduling strategy on the basis of ensuring the stability of the power grid, minimize the operating cost, improve the utilization efficiency of distributed resources, avoid a large number of iterative calculations, and improve the real-time performance and efficiency of decision-making.

[0063] Optionally, in the power distribution network resource scheduling method provided in the embodiments of the present application, the adversarial game model is trained in the following manner: Obtain M types of historical operation data and historical meteorological data of the power distribution network in a historical time period, and obtain the historical state data of the power distribution network in the historical time period, where at least one of the M types of historical operation data of the power distribution network includes: power distribution network topology data, historical load data, historical photovoltaic processing data, and energy storage system data, and the historical state data is used for the state of the power distribution network during operation in the historical time period, and M is a positive integer; preprocess each type of historical operation data of the power distribution network respectively to obtain M types of processed operation data, and preprocess the historical meteorological data to obtain processed historical meteorological data; perform per-unit value conversion on the M types of processed operation data to obtain M types of converted operation data, and use the historical state data, the M types of processed operation data, and the processed meteorological data to form a sample set, where the sample set includes sample input data and sample output data, the sample input data is composed of the M types of processed operation data and the processed meteorological data, and the sample output data is composed of the historical state data; use the sample set to train a preset adversarial game model to obtain the adversarial game model, where the preset adversarial game model includes a preset perturbation generator and a preset energy storage optimization discriminator.

[0064] It should be noted that per-unit value processing refers to converting electrical quantity data (such as voltage, current, power) into per-unit values, thereby eliminating the influence of physical dimensions and enabling the data to have better comparability and generalization ability during model training. In order for the model to learn the internal laws of the uncertainty of the source and load in the operation of the power distribution network, so that when facing the uncertainty in actual operation, it can generate more robust and economical scheduling action data, it is necessary to train this model.

[0065] First, multiple data can be obtained from the historical operation records of the power distribution network. For example, historical state data, power distribution network topology data, historical load data, historical photovoltaic processing data, and energy storage system data in a historical time period (such as the previous month) can be obtained. Among them, the topology data is used to characterize the structure of the power grid, including nodes, lines, transformers, etc.; the load data and photovoltaic data are used to characterize the power consumption and power generation information of the time series; the energy storage data is used to characterize the charge and discharge status and performance of the energy storage device. At the same time, meteorological data corresponding to the power distribution network in the historical time period can also be collected, including light intensity, temperature, humidity, wind speed, etc. After obtaining the data, it needs to be preprocessed and per-unit value converted to eliminate the dimension influence, so that data of different magnitudes can be compared and processed on the same scale, which can simplify the model design and improve the training efficiency. Among them, when performing per-unit value processing, for electrical quantities such as voltage, current, power, and impedance, they can be converted into per-unit values: where, x pu is the per-unit value; xactual , x base are the actual value and the reference value respectively.

[0066] Furthermore, a sample set for training the adversarial game model can be constructed based on the preprocessed data. Each sample consists of input data and corresponding output data. The input data includes the processed operation data and the processed meteorological data, and the output data is the historical state data reflecting the real operation state of the distribution network under specific historical conditions. By using the historical data as the sample input and output, the model can learn the ability to predict the future state or response strategy from the current state.

[0067] Finally, the constructed sample set is used to train the preset adversarial game model, so that the perturbation generator learns the perturbation pattern under the source-load uncertainty and generates challenging perturbation samples to simulate the uncertain scenarios that may be encountered in actual operation; the energy storage optimization discriminator learns the optimal scheduling strategy when facing uncertainty based on the current power grid state and the perturbation samples provided by the perturbation generator, and then obtains the adversarial game model. It should be noted that in the training process, the adversarial training method in deep reinforcement learning can be adopted, that is, a variant of Wasserstein GAN (Generative Adversarial Network) is used during the training of the perturbation generator to optimize the generated perturbation samples to make these samples as close as possible to the real data distribution to confuse and challenge the energy storage optimization discriminator. During the training of the energy storage optimization discriminator, through the DQN (Deep Q-Network) algorithm, combined with the sliding time window and the online learning mechanism, the parameters of the discriminator are optimized so that it can generate robust scheduling action data based on the input state and perturbation to minimize the operation cost and ensure the safe and stable operation of the power grid. Through adversarial training in this embodiment, the perturbation generator can generate various possible source-load uncertainty scenarios, and the energy storage optimization discriminator learns a more robust and adaptive scheduling strategy, which can effectively cope with uncertainty in actual operation, ensure that the scheduling strategy can quickly respond to source-load changes, meet the needs of real-time scheduling, provide an adaptive scheduling strategy for the distribution network with high proportion of renewable energy access, and improve the flexibility and intelligent decision-making ability of the power grid.

[0068] Optionally, in the power distribution network resource scheduling method provided in the embodiments of the present application, constructing a sample set using historical state data, M types of processed operation data, and processed meteorological data includes: extracting processed historical load data and processed photovoltaic data from the M types of processed operation data to obtain first operation data; performing discrete wavelet transform on the first operation data using a wavelet basis function, where the wavelet basis function is used to capture the time-frequency characteristics of the first operation data, and the component vector includes a first component vector and a second component vector, the frequency of the first component vector is less than the frequency of the second component vector, the first component vector corresponds to a trend feature, and the second component vector corresponds to a volatility feature; obtaining a sparse autoencoder, and using the sparse autoencoder to extract the component features of the component vector to obtain schedulable component features and uncontrollable fluctuation component features; constructing sample input data according to the schedulable component features, uncontrollable fluctuation component features, processed historical load data, processed photovoltaic data, and processed meteorological data, and constructing sample output data according to the historical state data.

[0069] Specifically, in order to better analyze the trend and volatility characteristics in the source-load uncertainty and provide high-quality training data for subsequent model training, wavelet transform and a sparse autoencoder (SAE) can be used to perform in-depth feature extraction and decoupling on relevant data, and then construct sample input and output data.

[0070] First, since historical load data and photovoltaic data can reflect the operation state and uncertainty of the power distribution network, a wavelet basis function (such as the Daubechies 4 wavelet) can be used to perform discrete wavelet transform (DWT) to extract the time-frequency characteristics in the data. Through wavelet transform, historical load and photovoltaic output data can be decomposed into component vectors of different frequencies, where the first component vector (i.e., the low-frequency component C low ) represents the trend feature, reflecting the long-term trend and seasonal changes of the source-load data; the second component vector (i.e., the high-frequency component D high ) corresponds to the volatility feature, reflecting the short-term fluctuations and randomness in the source-load data, such as the instantaneous change in light intensity or the sudden increase in user load. Among them, the low-frequency and high-frequency components obtained by decomposing historical load data and photovoltaic data can be expressed as:

[0071] Further, a Sparse Autoencoder (SAE) is used to further extract features from the component vectors. As an unsupervised learning model, the Sparse Autoencoder can learn a compact representation of the input data and automatically discover the internal structure of the data. Two independent SAEs can be trained, with SAE low for the first component vector (trend feature), and SAE high for the second component vector (volatility feature). Then, the schedulable component features and uncontrollable fluctuation component features are obtained based on the SAE. Among them, the schedulable component features can capture the predictable trend changes in the source-load data, can be used to guide long-term planning and resource allocation, and can be expressed as: f schedule =SAE low (C low ), and this feature is the output of the SAE low encoder; the uncontrollable fluctuation component features can capture the random fluctuations in the data and can be expressed as: f 波动 =SAE high (D high ), and this feature is the output of the SAE high encoder.

[0072] It should be noted that each SAE can be composed of a three-layer fully connected network, which can include: an input layer (capable of accepting the low-frequency component and high-frequency component after wavelet transform), a hidden layer (encoder) (including 128 neurons, with the activation function being the sigmoid function), and an output layer (decoder) (the number of neurons is the same as that of the input layer, and the activation function is the linear function). When training the SAE, a loss function can be used for optimization. Among them, the loss function is jointly determined by the mean square error and the sparsity penalty term: where N S is the number of samples, x (i) is the i-th input sample (which can be C low or D high ), is the reconstructed output of the SAE, h is the number of neurons in the hidden layer, λ is the sparsity penalty term weight (set to 0.001), is the divergence, ρ is the sparsity parameter (which can be set to 0.05), is the average activation degree of the j-th hidden layer neuron (expressed as where is the activation value of the j-th hidden layer neuron when the i-th sample is input).

[0073] After the schedulable component features and uncontrollable fluctuation component features are output by the above sparse autoencoder, the schedulable component features, uncontrollable fluctuation component features, processed historical load data, processed photovoltaic data, and meteorological data can be integrated to construct sample input data; sample output data is constructed based on the historical state data. When training the preset adversarial game model, the historical state data is vector-converted to obtain a state vector, and the schedulable component features are concatenated with the state vector as the input of the preset energy storage optimization discriminator; the uncontrollable fluctuation component feature f 波动 is concatenated with the processed historical load data, processed photovoltaic data, and processed meteorological data as the input of the preset perturbation generator. By processing the obtained sample data in this embodiment, the trend and volatility characteristics in the source-load data can be analyzed, so as to more deeply understand the internal laws in the uncertainty and provide a solid data basis for the formulation of robust scheduling strategies. In addition, the decoupled schedulable and uncontrollable component features help the model distinguish the predictable part and the random part in the data, so that the perturbation generator and the energy storage optimization discriminator can learn to cope with uncertainty more effectively.

[0074] Optionally, in the scheduling method of the distribution network resources provided in the embodiment of the present application, training the preset adversarial game model with a sample set to obtain an adversarial game model includes: the preset perturbation generator outputs a perturbation distribution according to the sample input data in the sample set, samples a historical perturbation distribution from the historical state data, and determines an interpolation perturbation distribution according to the perturbation distribution and the historical perturbation distribution; obtaining a penalty parameter of the distribution network, and calculating a first loss function of the preset energy storage optimization discriminator according to the perturbation distribution, the interpolation perturbation distribution, the historical state data, and the penalty parameter to obtain a first loss function; obtaining a discount factor and network parameters, calculating a gradient target value according to M types of converted operation data, the discount factor, and network parameters in the sample set, and determining a second loss function of the preset energy storage optimization discriminator according to the discount factor to obtain a second loss function; calculating the sum of the first loss function and the second loss function to obtain a discriminator loss function; obtaining a preset optimizer, and using the preset optimizer and the gradient ascent algorithm to determine the minimum value of the discriminator loss function to obtain a first minimum loss function value; determining empirical parameters corresponding to the first minimum loss function value, determining discriminator network parameters according to the empirical parameters corresponding to the first minimum loss function value, and adjusting the model parameters of the preset energy storage optimization discriminator based on the discriminator network parameters to obtain the energy storage optimization discriminator in the adversarial game model.

[0075] It should be noted that when training the adversarial game model, a sliding time window mechanism and an improved Wasserstein GAN (WGAN-GP) algorithm can be used to construct an online adversarial training framework to achieve online adaptive update of scheduling action data. That is, the dynamic game between the perturbation generator and the energy storage optimization discriminator is completed by iteratively updating the model parameters to minimize the loss function. Therefore, when training the energy storage optimization discriminator in the adversarial game model, first set the sliding time window and sliding step size, and then perform a sliding time window loop on this basis. That is, first initialize the parameters of the energy storage optimization discriminator and the perturbation generator, initialize the experience replay buffer, then perform data update to obtain the historical data within the current time window, secondly perform feature decoupling, and finally execute the adversarial game loop. First, the preset perturbation generator generates a perturbation distribution by learning the sample input data in the sample set. These perturbation distributions can simulate the uncertainties of photovoltaic power output and load values at different time points. That is, the goal of the perturbation generator is to continuously evolve during the adversarial training process to generate perturbation scenarios that can challenge the decision-making ability of the energy storage optimization discriminator.

[0076] Based on the perturbation distribution output by the perturbation generator, the preset penalty parameter, and the real historical state data of the distribution network, the energy storage optimization discriminator calculates the first loss function J D (θ D ). This loss function can measure the decision-making quality of the energy storage optimization discriminator when facing perturbation scenarios and whether its decisions can effectively cope with these perturbations, while ensuring the economy and stability of the distribution network operation. This function can be expressed as:

[0077]

[0078] Among them, the above sample data is a batch of empirical samples sampled from the experience replay buffer, and p real is the real perturbation distribution that can be sampled from historical data; δ~p real is a batch of real perturbation samples sampled from the real perturbation distribution; p data is the real data distribution of the distribution network; s~p data is a batch of state samples sampled from the real data distribution; p fake is the perturbation distribution generated by the perturbation generator; δ~p fake is a perturbation sample sampled from the perturbation distribution generated by the perturbation generator; δ fake is a batch of perturbation samples generated by the perturbation generator; is the distribution evenly sampled between the real perturbation distribution and the generated perturbation distribution (that is, the interpolation perturbation distribution), λ gpis the weight of the gradient penalty term; R(s,δ,a) is the reward function; E is the expectation operation, and D(s,δ) is the discriminator's score for state s and perturbation δ.

[0079] Further, after obtaining the discount factor γ and the network parameters the gradient target value can be calculated based on the transformed operation data, the discount factor, and the network parameters: and the second loss function of the preset energy storage optimization discriminator is determined according to the discount factor. This loss function can reflect the performance of the energy storage optimization discriminator during long-term operation and whether its decision can achieve the balance between optimal economic benefits and grid safe operation in the long term, and can be expressed as: Finally, the discriminator loss function is obtained according to the sum of the first loss function and the second loss function: J total,D = J D (θ D ) + L DQN (θ D ).

[0080] Furthermore, use the preset optimizer (i.e., the Adam optimizer) and the gradient ascent method to minimize the total loss function. By continuously adjusting the network parameters of the discriminator to minimize the loss function value, the decision-making ability of the discriminator is iteratively optimized, so as to obtain the first minimum loss function value. Then, based on the empirical parameters corresponding to the first minimum loss function value, the discriminator network parameters are determined. Based on the discriminator network parameters, the model parameters of the preset energy storage optimization discriminator are adjusted to obtain the energy storage optimization discriminator in the adversarial game model, thereby improving its overall performance. In this embodiment, by calculating and minimizing the loss function, the energy storage optimization discriminator has stronger robustness and economy. The energy storage optimization discriminator can achieve the optimal balance between immediate scheduling decisions and long-term grid planning, providing strong technical support for immediate scheduling decisions.

[0081] Optionally, in the scheduling method of the distribution network resources provided in the embodiments of the present application, training the preset adversarial game model using the sample set to obtain the adversarial game model includes: obtaining the perturbation distribution output by the preset perturbation generator; calculating the perturbation score by the preset perturbation generator according to the sample set, determining the loss function of the preset perturbation generator based on the perturbation score and the perturbation distribution to obtain the generator loss function; obtaining the preset optimizer, using the preset optimizer and the gradient descent algorithm to determine the minimum value of the generator loss function to obtain the second minimum loss function value; determining the empirical parameters corresponding to the second minimum loss function value, determining the generator network parameters according to the empirical parameters corresponding to the second minimum loss function value, and adjusting the model parameters of the preset perturbation generator based on the generator network parameters to obtain the perturbation generator in the adversarial game model.

[0082] Specifically, the dynamic game between the perturbation generator and the energy storage optimization discriminator is achieved by calculating and minimizing their respective loss functions. The goal of the perturbation generator is to generate a perturbation distribution that can mimic real perturbations and is challenging, so as to promote the learning of the energy storage optimization discriminator and enable the latter to make more robust decisions in the face of various uncertainties. When training the perturbation generator, it is first preset that the perturbation generator can generate a perturbation distribution δ that simulates the uncertainties of photovoltaic power output and load values by learning the transformed operation data and historical meteorological data in the sample set. fake Then, the preset perturbation generator can calculate the perturbation score D(s,δ) according to the sample set, and then use the perturbation score and the above perturbation distribution to determine the loss function of the preset perturbation generator, obtaining the generator loss function: J G (θ G ) = -E s~pdata,pfake [D(s,δ)].

[0083] Furthermore, the quality of the generated perturbations is optimized by minimizing the Wasserstein distance from the real perturbation distribution, that is, using the preset optimizer and the gradient descent algorithm to determine the minimum value of the generator loss function, thereby optimizing the quality of perturbation generation, ensuring that the generated perturbations can better simulate the source-load uncertainties in the real distribution network, and at the same time challenging the decision boundary of the energy storage optimization discriminator. Finally, the model parameters of the perturbation generator are adjusted based on the network parameters obtained after minimizing the generator loss function to achieve higher-quality perturbation generation in the adversarial game model. In this embodiment, by minimizing the distance from the real perturbation distribution, a perturbation distribution closer to the uncertainty scenario in actual operation is generated, thereby optimizing the robustness and authenticity of perturbation generation, helping to enhance the decision-making ability of the discriminator, improving the overall performance of the distribution network distributed resource scheduling scheme, and ensuring that the distribution network can operate safely, stably and economically under the condition of high proportion of renewable energy access, while having good real-time performance and self-adaptability.

[0084] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0085] The embodiment of the present application also provides a method for scheduling distribution network resources. Figure 2 It is a schematic diagram of an optional method for scheduling distribution network resources provided by the embodiment of the present application, as Figure 2 shown, this method includes:

[0086] To ensure the safe and stable operation of the distribution network, firstly, historical operation data of the distribution network can be collected and preprocessed, including distribution network topology data, historical load data, historical photovoltaic processing data, meteorological data, and energy storage system data, and these data are cleaned to remove missing values, outliers, and noise. Then, all electrical quantities (such as voltage, current, and power) are converted into per-unit values for easy model processing. Finally, the dataset is divided into a training set, a validation set, and a test set, which are used for model training, parameter adjustment, and final performance evaluation respectively.

[0087] Furthermore, an adversarial model architecture is constructed, that is, a dynamic game model architecture of "photovoltaic / load disturbance generator - energy storage optimization discriminator" is constructed. The disturbance generator is used to simulate challenging photovoltaic power output and load disturbance samples; the energy storage optimization discriminator is used to generate robust energy storage scheduling strategies. To better train the above model architecture, bidirectional feature decoupling can be performed on the obtained sample data, that is, the photovoltaic power output and load data are decomposed using wavelet transform to extract trend features and volatility features. Subsequently, sparse autoencoders are used to process the low-frequency and high-frequency components respectively to extract schedulable component features and uncontrollable fluctuation component features, so as to effectively separate the fixed patterns and random disturbances in the source-load data, that is, the uncontrollable fluctuation component features and historical data are used as the input of the disturbance generator, and the schedulable component features are concatenated with the state vector as the input of the energy storage optimization discriminator.

[0088] Then, the above model architecture is trained by means of online adversarial training, that is, by setting a sliding time window to ensure that the model architecture can be trained based on the latest data. Then, by calculating the distance between the real disturbance distribution and the generated disturbance distribution, and the penalty term for the gradient, the decision-making ability of the discriminator is optimized; by minimizing the discriminator's score for the generated disturbance, the disturbance generation quality of the generator is optimized; the reward function is used to comprehensively consider the economic operation of the distribution network, voltage quality, health of the energy storage device, and robustness of the scheduling strategy to promote the energy storage optimization discriminator to generate economic and robust scheduling strategies. Finally, the Adam optimizer and the gradient ascent / descent algorithm are used to optimize the network parameters of the discriminator and the generator respectively, and minimize their respective loss functions.

[0089] After the above model architecture is trained, scheduling action data can be generated based on the current state of the distribution network, including photovoltaic and load values, and the current energy storage SOC state, that is, the generator is used to generate disturbance samples, and then the current state and disturbance samples are input into the energy storage optimization discriminator, and the action with the largest Q value is selected as the optimal scheduling action, so as to obtain the scheduling action data. Finally, the scheduling action data is applied to the distribution network simulation environment or the actual distribution network to execute the scheduling instructions, and the effectiveness of the model is evaluated according to the grid operation results, and the scheduling strategy is continuously optimized.

[0090] In this embodiment, an adversarial training is carried out online against the model architecture, and then scheduling actions are generated based on this architecture, improving the self - adaptability and real - time response ability of the model. Thus, the distribution network can effectively cope with the uncertainties of photovoltaic power output and load values, ensuring the safe and stable operation of the power grid under various disturbance conditions. At the same time, economic dispatch is realized, the operation cost is reduced, providing strong technical support for the safe operation of a distribution network with a high proportion of renewable energy access.

[0091] The embodiment of the present application also provides a scheduling device for distribution network resources. It should be noted that the scheduling device for distribution network resources in the embodiment of the present application can be used to execute the scheduling method for distribution network resources provided by the embodiment of the present application. The following introduces the scheduling device for distribution network resources provided by the embodiment of the present application.

[0092] Figure 3 is a schematic diagram of the scheduling device for distribution network resources provided by the embodiment of the present application, as Figure 3 shown, the device includes: an acquisition unit 30, an input unit 31, and an adjustment unit 32.

[0093] The acquisition unit 30 is configured to acquire the operation data and meteorological data of the distribution network within a preset time period, where the operation data includes at least one of the following: photovoltaic data, load value data, and state of charge data;

[0094] The input unit 31 is configured to input the operation data and meteorological data into the adversarial game model and output scheduling action data, where the scheduling action data represents the energy storage charge - discharge power adjustment strategy of the distribution network;

[0095] The adjustment unit 32 is configured to perform scheduling adjustment on the resources of the distribution network based on the scheduling action data, where the scheduling adjustment methods include at least one of the following: adjusting the energy storage charge - discharge power of the distribution network, and coordinately dispatching the distributed resources of the distribution network.

[0096] The dispatching device for distribution network resources provided by the embodiments of the present application obtains the operation data and meteorological data of the distribution network within a preset time period through an acquisition unit 30. Among them, the operation data includes at least one of the following: photovoltaic data, load value data, and state of charge data. An input unit 31 inputs the operation data and meteorological data into an adversarial game model and outputs dispatching action data, where the dispatching action data represents the energy storage charge and discharge power adjustment strategy of the distribution network. An adjustment unit 32 performs dispatching adjustment on the resources of the distribution network based on the dispatching action data. Among them, the dispatching adjustment methods include at least one of the following: adjusting the energy storage charge and discharge power of the distribution network, and coordinately dispatching the distributed resources of the distribution network, which solves the technical problems of low robust dispatching ability and poor stability in the related high-proportion photovoltaic distribution network. By obtaining the operation data and meteorological data of the distribution network, processing them using the adversarial game model, generating dispatching action data, and performing resource dispatching adjustment based on this data, the effect of significantly improving the operation stability and robustness of the high-proportion photovoltaic distribution network is achieved.

[0097] Optionally, in the dispatching device for distribution network resources provided by the embodiments of the present application, the input unit 31 includes: a first determination module for determining a state data matrix according to the operation data and meteorological data; a first extraction module for extracting the spatial features in the state data matrix by a feature extraction layer and outputting the spatial features, where the feature extraction layer includes N sub-feature extraction layers, and each sub-feature extraction layer is associated with a first weight matrix and a first bias parameter vector, and N is a positive integer; a modeling module for performing temporal modeling on the spatial features by a temporal modeling layer and outputting temporal features, where the temporal modeling layer includes Y sub-temporal modeling layers, and each sub-temporal modeling layer is associated with a second weight matrix and a second bias parameter vector, and Y is a positive integer; a first calculation module for calculating the temporal features by a fully connected output layer according to a preset activation function, outputting perturbation parameters, and inputting the perturbation parameters into an energy storage optimization discriminator to output dispatching action data.

[0098] Optionally, in the dispatching device for distribution network resources provided by the embodiments of the present application, the input unit 31 includes: a splicing module for splicing the state data matrix and the perturbation parameters to obtain an input vector; a second calculation module for calculating the input vector based on the network parameters of the hidden layer to obtain a hidden layer data matrix; a first acquisition module for acquiring K candidate dispatching action data within a preset time period, and calculating the action value function vectors of the K candidate dispatching action data by an output layer according to the hidden layer data matrix, where K is a positive integer; a screening module for screening the K candidate dispatching action data to obtain the candidate dispatching action data with the largest action value function vector, and determining the candidate dispatching action data with the largest action value function vector as the dispatching action data.

[0099] Optionally, in the dispatching device for distribution network resources provided in the embodiments of the present application, the input unit 31 includes: a second acquisition module, configured to acquire M types of historical operation data of the distribution network and historical meteorological data in a historical time period, and acquire historical state data of the distribution network in the historical time period, where at least one of the M types of historical operation data of the distribution network includes: distribution network topology data, historical load data, historical photovoltaic processing data, and energy storage system data, and the historical state data is used for the state of the distribution network during operation in the historical time period, and M is a positive integer; a processing module, configured to preprocess each type of historical operation data of the distribution network respectively to obtain M types of processed operation data, and preprocess the historical meteorological data to obtain processed historical meteorological data; a conversion module, configured to perform per-unit value conversion on the M types of processed operation data to obtain M types of converted operation data, and use the historical state data, the M types of processed operation data, and the processed meteorological data to form a sample set, where the sample set includes sample input data and sample output data, the sample input data is composed of the M types of processed operation data and the processed meteorological data, and the sample output data is composed of the historical state data; a training module, configured to use the sample set to train a preset adversarial game model to obtain an adversarial game model, where the preset adversarial game model includes a preset perturbation generator and a preset energy storage optimization discriminator.

[0100] Optionally, in the dispatching device for distribution network resources provided in the embodiments of the present application, the input unit 31 includes: a second extraction module, configured to extract the processed historical load data and the processed photovoltaic processing data from the M types of processed operation data to obtain first operation data; a transformation module, configured to perform discrete wavelet transform on the first operation data by using a wavelet basis function to obtain component vectors, where the wavelet basis function is used to capture the time-frequency characteristics in the first operation data, the component vectors include a first component vector and a second component vector, the frequency of the first component vector is less than the frequency of the second component vector, the first component vector corresponds to a trend feature, and the second component vector corresponds to a volatility feature; a third acquisition module, configured to acquire a sparse autoencoder, and use the sparse autoencoder to extract component features of the component vectors to obtain schedulable component features and uncontrollable fluctuation component features; a construction module, configured to construct sample input data according to the schedulable component features, the uncontrollable fluctuation component features, the processed historical load data, the processed photovoltaic processing data, and the processed meteorological data, and construct sample output data according to the historical state data.

[0101] Optionally, in the dispatching device for distribution network resources provided in the embodiments of the present application, the input unit 31 includes: an output module, configured to output a disturbance distribution by the preset disturbance generator according to the sample input data in the sample set, sample a historical disturbance distribution from the historical state data, and determine an interpolation disturbance distribution according to the disturbance distribution and the historical disturbance distribution; a fourth acquisition module, configured to acquire a penalty parameter of the distribution network, and calculate a first loss function of the preset energy storage optimization discriminator according to the disturbance distribution, the interpolation disturbance distribution, the historical state data, and the penalty parameter, to obtain the first loss function; a fifth acquisition module, configured to acquire a discount factor and network parameters, calculate a gradient target value according to the M types of converted operation data, the discount factor, and the network parameters in the sample set, and determine a second loss function of the preset energy storage optimization discriminator according to the discount factor, to obtain the second loss function; a third calculation module, configured to calculate the sum of the first loss function and the second loss function, to obtain a discriminator loss function; a sixth acquisition module, configured to acquire a preset optimizer, and use the preset optimizer and the gradient ascent algorithm to determine the minimum value of the discriminator loss function, to obtain a first minimum loss function value; a second determination module, configured to determine the empirical parameters corresponding to the first minimum loss function value, determine the discriminator network parameters according to the empirical parameters corresponding to the first minimum loss function value, and perform model parameter adjustment on the preset energy storage optimization discriminator based on the discriminator network parameters, to obtain the energy storage optimization discriminator in the adversarial game model.

[0102] Optionally, in the dispatching device for distribution network resources provided in the embodiments of the present application, the input unit 31 includes: a seventh acquisition module, configured to acquire the disturbance distribution output by the preset disturbance generator; a fourth calculation module, configured to calculate a disturbance score by the preset disturbance generator according to the sample set, and determine a loss function of the preset disturbance generator based on the disturbance score and the disturbance distribution, to obtain a generator loss function; an eighth acquisition module, configured to acquire a preset optimizer, and use the preset optimizer and the gradient descent algorithm to determine the minimum value of the generator loss function, to obtain a second minimum loss function value; a third determination module, configured to determine the empirical parameters corresponding to the second minimum loss function value, determine the generator network parameters according to the empirical parameters corresponding to the second minimum loss function value, and perform model parameter adjustment on the preset disturbance generator based on the generator network parameters, to obtain the disturbance generator in the adversarial game model.

[0103] The above-mentioned dispatching device for distribution network resources includes a processor and a memory. The above-mentioned acquisition unit 30, input unit 31, adjustment unit 32, etc. are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above program units stored in the memory.

[0104] The processor contains a kernel, which retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the technical problems of low robust scheduling ability and poor stability in high-proportion photovoltaic distribution networks in related technologies can be solved.

[0105] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory includes at least one memory chip.

[0106] An embodiment of the present invention provides a computer storage medium for storing a program, where when the program runs, it controls the device where the computer storage medium is located to execute a scheduling method for distribution network resources.

[0107] Figure 4 It is a schematic diagram of an electronic device provided according to an embodiment of the present application, as Figure 4 shown, an embodiment of the present invention provides an electronic device. The electronic device 40 includes a processor, a memory, and a program stored on the memory and executable on the processor. The processor is used to run computer-readable instructions, where when the computer-readable instructions run, they execute a scheduling method for distribution network resources. The devices herein can be servers, PCs, PADs, mobile phones, etc.

[0108] The present application also provides a computer program product, including a computer program, which when executed by a processor implements the steps of a scheduling method for distribution network resources in various embodiments of the present application.

[0109] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0110] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementation in the processFigure 1 one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks

[0111] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks

[0112] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the functions in the process Figure 1 one or more processes and / or blocks Figure 1 steps for the functions specified in one or more blocks

[0113] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory

[0114] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media

[0115] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves

[0116] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0117] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for dispatching distribution network resources, characterized in that: include: Acquiring operation data and meteorological data of the distribution network within a preset time period, wherein the operation data includes at least one of the following: photovoltaic data, load value data, and state of charge data; Inputting the operation data and the meteorological data into a confrontation game model, and outputting scheduling action data, wherein the scheduling action data represents the energy storage charging and discharging power adjustment strategy of the distribution network; The resources of the distribution network are dispatched and adjusted based on the dispatch action data, wherein the dispatch adjustment method includes at least one of the following: adjusting the energy storage charging and discharging power of the distribution network, and coordinating the dispatch of distributed resources of the distribution network.

2. The method according to claim 1, characterized in that The adversarial game model includes a disturbance generator and an energy storage optimization discriminator. The disturbance generator includes a feature extraction layer, a time series modeling layer, and a fully connected output layer. The operation data and the meteorological data are input into the adversarial game model, and the output scheduling action data includes: Determine a state data matrix according to the operation data and the meteorological data; The feature extraction layer extracts spatial features in the state data matrix and outputs spatial features, wherein the feature extraction layer includes N layers of sub-feature extraction layers, each sub-feature extraction layer is associated with a first weight matrix and a first bias parameter vector, and N is a positive integer; The temporal modeling layer performs temporal modeling on the spatial features and outputs temporal features, wherein the temporal modeling layer includes Y sub-temporal modeling layers, each sub-temporal modeling layer is associated with a second weight matrix and a second bias parameter vector, and Y is a positive integer; The fully connected output layer calculates the timing features according to a preset activation function, outputs disturbance parameters, inputs the disturbance parameters into the energy storage optimization discriminator, and outputs the scheduling action data.

3. The method according to claim 2, characterized in that The energy storage optimization discriminator includes a hidden layer and an output layer. The disturbance parameter is input into the energy storage optimization discriminator, and the scheduling action data is outputted, including: Concatenating the state data matrix and the disturbance parameter to obtain an input vector; Calculate the input vector based on the network parameters of the hidden layer to obtain a hidden layer data matrix; Acquire K candidate scheduling action data within the preset time period, and calculate the K candidate scheduling action data by the output layer according to the hidden layer data matrix to obtain the action value function vector of the K candidate scheduling action data, wherein K is a positive integer; The K candidate scheduling action data are screened to obtain the candidate scheduling action data with the largest action value function vector, and the candidate scheduling action data with the largest action value function vector is determined as the scheduling action data.

4. The method according to claim 1, characterized in that: The adversarial game model is trained in the following way: Obtain historical operation data and historical meteorological data of M types of distribution networks in a historical time period, and obtain historical status data of the distribution network in the historical time period, wherein the M types of distribution network historical operation data include at least one of the following: distribution network topology data, historical load data, historical photovoltaic processing data, and energy storage system data, and the historical status data is used for the status of the distribution network when it was running in the historical time period, and M is a positive integer; Preprocessing each type of distribution network historical operation data to obtain M types of processed operation data, and preprocessing the historical meteorological data to obtain processed historical meteorological data; Performing per-unit conversion on the M-type processed operation data to obtain the M-type converted operation data, and using the historical state data, the M-type processed operation data and the processed meteorological data to form a sample set, wherein the sample set includes sample input data and sample output data, the M-type processed operation data and the processed meteorological data constitute the sample input data, and the historical state data constitutes the sample output data; The preset adversarial game model is trained using the sample set to obtain the adversarial game model, wherein the preset adversarial game model includes a preset disturbance generator and a preset energy storage optimization discriminator.

5. The method according to claim 4, characterized in that The sample set is formed by using the historical status data, the M-type processed operation data and the processed meteorological data, including: Extracting processed historical load data and processed photovoltaic processing data from the M-type processed operation data to obtain first operation data; Performing discrete wavelet transform on the first operation data using a wavelet basis function to obtain a component vector, wherein the wavelet basis function is used to capture the time-frequency characteristics of the first operation data, the component vector includes a first component vector and a second component vector, the frequency of the first component vector is less than the frequency of the second component vector, the first component vector corresponds to a trend feature, and the second component vector corresponds to a volatility feature; Obtain a sparse autoencoder, and use the sparse autoencoder to extract component features of the component vector to obtain schedulable component features and uncontrollable fluctuation component features; The sample input data is constructed based on the characteristics of the dispatchable component, the characteristics of the uncontrollable fluctuating component, the processed historical load data, the processed photovoltaic processing data, and the processed meteorological data, and the sample output data is constructed based on the historical status data.

6. The method according to claim 4, characterized in that The preset adversarial game model is trained by using the sample set to obtain the adversarial game model, including: The preset disturbance generator outputs a disturbance distribution according to the sample input data in the sample set, samples the historical state data to obtain a historical disturbance distribution, and determines an interpolated disturbance distribution according to the disturbance distribution and the historical disturbance distribution; Obtaining a penalty parameter of the distribution network, and calculating a first loss function of the preset energy storage optimization discriminator according to the disturbance distribution, the interpolated disturbance distribution, the historical state data, and the penalty parameter to obtain a first loss function; Obtaining a discount factor and a network parameter, calculating a gradient target value according to the M types of converted operating data in the sample set, the discount factor and the network parameter, and determining a second loss function of the preset energy storage optimization discriminator according to the discount factor to obtain a second loss function; Calculate the sum of the first loss function and the second loss function to obtain a discriminator loss function; Obtain a preset optimizer, and use the preset optimizer and a gradient ascent algorithm to determine the minimum value of the discriminator loss function to obtain a first minimum loss function value; Determine the empirical parameters corresponding to the first minimum loss function value, determine the discriminator network parameters according to the empirical parameters corresponding to the first minimum loss function value, adjust the model parameters of the preset energy storage optimization discriminator based on the discriminator network parameters, and obtain the energy storage optimization discriminator in the adversarial game model.

7. The method according to claim 4, characterized in that The preset adversarial game model is trained by using the sample set to obtain the adversarial game model, including: Obtaining a disturbance distribution output by the preset disturbance generator; The preset disturbance generator calculates a disturbance score according to the sample set, and determines a loss function of the preset disturbance generator based on the disturbance score and the disturbance distribution to obtain a generator loss function; Obtain a preset optimizer, and use the preset optimizer and a gradient descent algorithm to determine the minimum value of the generator loss function to obtain a second minimum loss function value; Determine the empirical parameters corresponding to the second minimum loss function value, determine the generator network parameters according to the empirical parameters corresponding to the second minimum loss function value, adjust the model parameters of the preset disturbance generator based on the generator network parameters, and obtain the disturbance generator in the adversarial game model.

8. A dispatching device for distribution network resources, characterized in that: include: An acquisition unit, used to acquire operation data and meteorological data of the distribution network within a preset time period, wherein the operation data includes at least one of the following: photovoltaic data, load value data, and state of charge data; An input unit, used to input the operation data and the meteorological data into a confrontation game model, and output dispatch action data, wherein the dispatch action data represents the energy storage charging and discharging power adjustment strategy of the distribution network; An adjustment unit is used to perform scheduling adjustment on the resources of the distribution network based on the scheduling action data, wherein the scheduling adjustment method includes at least one of the following: adjusting the energy storage charging and discharging power of the distribution network, and coordinating the scheduling of distributed resources of the distribution network.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for dispatching distribution network resources described in any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the scheduling method for distribution network resources as described in any one of claims 1 to 7.