Power resource updating method, device, computer equipment and storage medium

By using convex hull improvement model and reinforcement learning model in power resource update, combined with power resource update model, the problem that SCUC and SCED problems in the prior art are difficult to solve multiple times in a short time, and the efficiency and timeliness of power resource updates are achieved.

CN117435602BActive Publication Date: 2025-05-06CHINA SOUTHERN POWER GRID COMPANY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311284793.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-07
Publication Date
2025-05-06
Estimated Expiration
2043-10-07

AI Technical Summary

Technical Problem

The existing SCUC and SCED problems are difficult to solve multiple times in a short time, resulting in complex and difficult solutions for robust update models, which cannot meet the timeliness requirements of power resource updates.

Method used

By acquiring the initial data packets, convex hull improvements are performed, and combined with the pre-trained reinforcement learning model and the power resource update model, the approximate optimal results for each period of the current power resource update requirements are obtained, thereby realizing power resource updates.

Benefits of technology

It improves the efficiency of power resource updates, reduces the time and computing resources required for model operations, and enhances the accuracy and stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117435602B_ABST
    Figure CN117435602B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device, computer equipment, storage medium and computer program product for updating electric power resources. The method comprises: obtaining an initial data packet for the update demand of the current electric power resources; the initial data packet contains the update constraint information, update target information and update parameter information of the current electric power resources; inputting the initial data packet into the convex hull improvement model for data improvement to obtain an improved data packet; obtaining the initial electric power resource data of the current electric power resources, inputting the initial electric power resource data and the improved data packet into a reinforcement learning model with an initial policy network to obtain the approximate optimal results of each time period for the update demand of the current electric power resources; inputting the approximate optimal results of each time period, the initial electric power resource data and the improved data packet into a pre-trained electric power resource update model to obtain the electric power resource update results for the update demand of the current electric power resources. The use of this method can improve the efficiency of electric power resource update calculations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart grid technology, and in particular to a method, device, computer equipment, storage medium and computer program product for updating electric power resources. Background Art

[0002] Currently, there are two main types of solutions to power update models, namely robust update and random update. Robust update is usually performed under the most conservative conditions; while random update takes multiple solutions to the scenario as its core idea to obtain the expected value.

[0003] However, the core idea of ​​random update is to solve the scenario multiple times to obtain the expected value, but the existing SCUC (security constrained unit commitment) or SCED (security constrained economic dispatch) problems are difficult to solve multiple times in a short period of time; at this stage, SCUC and SCED problems are mixed integer linear programming problems. If complex robust update sets are considered, the robust equivalence model will be very complex and difficult to solve, which does not meet the timeliness requirements of power resource updates. Summary of the invention

[0004] Based on this, it is necessary to provide a power resource updating method, device, computer equipment, computer readable storage medium and computer program product that can improve the power resource updating efficiency in response to the above technical problems.

[0005] In a first aspect, the present application provides a method for updating electric power resources, comprising:

[0006] Acquire an initial data packet for the update requirements of the current power resources; the initial data packet includes update constraint information, update target information and update parameter information of the current power resources;

[0007] Inputting the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet;

[0008] Acquire initial power resource data of the current power resource, input the initial power resource data and the improved data packet into a pre-trained reinforcement learning model, and obtain approximate optimal results for each time period of the update demand of the current power resource; the pre-trained reinforcement learning model has an initial strategy network;

[0009] The approximate optimal results of each time period, the initial power resource data and the improved data packet are input into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource.

[0010] In one embodiment, before inputting the initial power resource data and the improved data packet into a pre-trained reinforcement learning model to obtain the approximate optimal result for each time period of the update demand for the current power resource, the method further includes:

[0011] Get updated samples of historical power resources;

[0012] Inputting the historical power resource update samples into the imitation learning model to obtain historical update results and a historical strategy network;

[0013] The initial reinforcement learning model is trained according to the historical update results and the historical strategy network to obtain the pre-trained reinforcement learning model with the initial strategy network.

[0014] In one embodiment, before inputting the approximate optimal results of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource, the method further includes:

[0015] Determining the feasibility of the approximate optimal result of each of the time periods with respect to the update demand of the current power resource;

[0016] Adopting the feasibility repair model, the infeasible approximate optimal result is repaired to obtain the repaired approximate optimal result;

[0017] The step of inputting the approximate optimal result of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource includes:

[0018] The repaired approximate optimal result, the feasible approximate optimal result, the initial power resource data and the improved data packet are input into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource.

[0019] In one embodiment, the approximate optimal result of each time period, the initial power resource data and the improved data packet are input into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource, including:

[0020] In the power resource update model, the improved data packet is preprocessed to obtain a preprocessed improved data packet; the preprocessing at least includes redundant constraint identification processing, variable reduction processing, variable boundary tightening processing, variable parameter updating processing and variable boundary detection processing;

[0021] According to the pre-processed improved data packet, the approximate optimal result of each time period and the initial power resource data, a power resource update result for the update demand of the current power resource is obtained.

[0022] In one embodiment, obtaining the power resource update result for the update demand of the current power resource according to the preprocessed improved data packet, the approximate optimal result of each time period and the initial power resource data includes:

[0023] Using a linear relaxation solution model to process the preprocessed improved data packet, the approximate optimal results of each time period and the initial power resource data, to obtain a relaxed lower bound of the power resource update result;

[0024] The relaxed lower bound is used as an initial lower bound, and a cutting plane model is used to process the preprocessed improved data packet, the approximate optimal result of each time period, and the initial power resource data to obtain a result range of the power resource update result;

[0025] Using heuristic models to obtain suboptimal and local optimal results from the range of results;

[0026] The suboptimal result and the local optimal result are input into a parallel branch and bound model to obtain a global optimal result as a power resource update result for the update demand of the current power resource.

[0027] In one embodiment, before inputting the approximate optimal results of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource, the method further includes:

[0028] Get updated samples of historical power resources;

[0029] The historical power resource update samples are used as training samples, and a black box optimization model is used to adjust the parameters of the initial power resource update model to obtain a trained power resource update model.

[0030] In a second aspect, the present application also provides a power resource updating device, comprising:

[0031] A data acquisition module, used to acquire an initial data packet for the update requirements of the current power resources; the initial data packet includes update constraint information, update target information and update parameter information of the current power resources;

[0032] A data improvement module, used for inputting the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet;

[0033] A reinforcement learning module, used for obtaining initial power resource data of the current power resource, inputting the initial power resource data and the improved data packet into a pre-trained reinforcement learning model, and obtaining approximate optimal results for each time period of the update demand of the current power resource; the pre-trained reinforcement learning model has an initial strategy network;

[0034] The resource update module is used to input the approximate optimal results of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource.

[0035] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0036] Acquire an initial data packet for the update requirements of the current power resources; the initial data packet includes update constraint information, update target information and update parameter information of the current power resources;

[0037] Inputting the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet;

[0038] Acquire initial power resource data of the current power resource, input the initial power resource data and the improved data packet into a pre-trained reinforcement learning model, and obtain approximate optimal results for each time period of the update demand of the current power resource; the pre-trained reinforcement learning model has an initial strategy network;

[0039] The approximate optimal results of each time period, the initial power resource data and the improved data packet are input into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource.

[0040] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0041] Acquire an initial data packet for the update requirements of the current power resources; the initial data packet includes update constraint information, update target information and update parameter information of the current power resources;

[0042] Inputting the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet;

[0043] Acquire initial power resource data of the current power resource, input the initial power resource data and the improved data packet into a pre-trained reinforcement learning model, and obtain approximate optimal results for each time period of the update demand of the current power resource; the pre-trained reinforcement learning model has an initial strategy network;

[0044] The approximate optimal results of each time period, the initial power resource data and the improved data packet are input into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource.

[0045] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0046] Acquire an initial data packet for the update requirements of the current power resources; the initial data packet includes update constraint information, update target information and update parameter information of the current power resources;

[0047] Inputting the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet;

[0048] Acquire initial power resource data of the current power resource, input the initial power resource data and the improved data packet into a pre-trained reinforcement learning model, and obtain approximate optimal results for each time period of the update demand of the current power resource; the pre-trained reinforcement learning model has an initial strategy network;

[0049] The approximate optimal results of each time period, the initial power resource data and the improved data packet are input into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource.

[0050] The above-mentioned power resource updating method, device, computer equipment, storage medium and computer program product first obtain an initial data packet for the update demand of the current power resource, wherein the initial data packet contains the update constraint information, update target information and update parameter information of the current power resource; then, the initial data packet is input into the convex hull improvement model for data improvement to obtain an improved data packet. The convex hull improvement model can be used to improve the mathematical representation of the data in the initial data packet to more accurately reflect the update constraint information, update target information and update parameter information of the current power resource, thereby reducing the search space of the update operation and improving the model operation efficiency during the subsequent power resource update; then, the initial power resource data of the current power resource is obtained, and the initial power resource data of the power resource and the improved data packet are input into the pre-trained reinforcement learning model to obtain the data packet for the current power resource. The approximate optimal results of each time period for the power resource update demand are obtained, wherein the pre-trained reinforcement learning model has an initial policy network, and the initial policy network can reduce the ineffective exploration cost of the reinforcement learning model when the power resources are updated, thereby improving the efficiency of the reinforcement learning model in obtaining the approximate optimal results; finally, the approximate optimal results of each time period, the initial power resource data of the power resources, and the improved data packet are input into the pre-trained power resource update model to obtain the power resource update results for the current power resource update demand. The approximate optimal results can be used as the starting point of the power resource update model, which helps to reduce the computational burden of the model, and can significantly reduce the time and computing resources required for the model operation. In addition, the approximate optimal results can also be used as the initial solution in the power resource update model, which significantly improves the model operation efficiency, as well as improves the accuracy and stability of the model. In the above method, before using the power resource update model, the reinforcement learning model is first used to obtain the approximate optimal result for the current power resource update demand, and the approximate optimal result is used as the initial solution of the power resource update model, thereby improving the solution operation efficiency of the power resource update model; and the reinforcement learning model has an initial strategy network, which can reduce the invalid exploration cost of the reinforcement learning model when the power resource is updated, thereby improving the efficiency of the reinforcement learning model in obtaining the approximate optimal result; in addition, the input to the reinforcement learning model and the power resource update model is an improved data packet that has been improved by the convex hull improvement model, and the improved data packet can more accurately reflect the update constraint information, update target information and update parameter information of the current power resources, reduce the search space of the update operation, and further improve the solution operation efficiency of the reinforcement learning model and the power resource update model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0052] Figure 1 is a flow chart of a method for updating electric power resources in one embodiment;

[0053] Figure 2 is a schematic diagram of a training process of a reinforcement learning model in one embodiment;

[0054] Figure 3 A schematic diagram of a process flow of a feasibility repair model in an embodiment;

[0055] Figure 4 is a flow chart of a heuristic model in one embodiment;

[0056] Figure 5 A schematic diagram of a training process of an electric power resource update model in one embodiment;

[0057] Figure 6 is a flow chart of a method for updating electric power resources in another embodiment;

[0058] Figure 7 is a structural block diagram of a power resource updating device in one embodiment;

[0059] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0061] In one embodiment, Figure 1 As shown, a method for updating power resources is provided. This embodiment uses the method applied to a terminal as an example for explanation. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smart phones, and tablet computers, and the server can be implemented as an independent server or a server cluster consisting of multiple servers.

[0062] In this embodiment, the method includes the following steps:

[0063] Step S101, obtaining an initial data packet for the update requirement of the current power resources.

[0064] The initial data packet includes update constraint information, update target information and update parameter information of the current power resources.

[0065] Exemplarily, the current update demand of power resources may be a demand such as SCUC (security constrained unit commitment) or SCED (security constrained economic dispatch). The update template information may be to minimize the operating cost, specifically to minimize the operating cost = unit operating cost + unit startup cost + power plant abandoned water penalty function + unit optimization unit operating cost + interconnection line unit operating cost + line constraint penalty function + section constraint penalty function. The update constraint information may include basic constraint information and extended constraint information. Among them, the basic constraint information may include unit constraints, system constraints and network constraints; unit constraints at least include output upper and lower limit constraints, ramp rate constraints, minimum continuous start and stop time constraints and maximum start and stop times constraints; system constraints at least include load balance constraints, system positive and negative reserve constraints and system rotating reserve constraints; network constraints at least include line flow constraints and section transmission polar line constraints. Extended constraints at least include island system load balance constraints, interconnection line refined modeling constraints, unit group constraints, zoned utilization constraints, cascade hydropower constraints and gas unit constraints. The update parameter information corresponds to the update constraint information and the update target information, and the final power resource update result is an update adjustment of the update parameters.

[0066] Step S102, inputting the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet.

[0067] Exemplarily, when the terminal adopts the convex hull improvement model, it utilizes the information in the initial data packet and combines the convex hull modeling technology to generate a more accurate and optimized improved data packet. Convex hull modeling technology is applied to process the initial data packet. The convex hull is a mathematical tool used to describe the smallest convex polygon containing all points in a multidimensional space. In this embodiment, the convex hull is used to determine the update space of possible power resource configurations. The convex hull improvement model adjusts the content of the initial data packet according to the update constraint information and the update target information, which may specifically include linear programming, integer programming or other optimization methods to ensure that the generated improved data packet better meets the requirements of the power system. The initial data packet is preset by the user and can be a mathematical expression that describes the update target and update constraint. The convex hull improvement model improves the mathematical expression to obtain an improved data packet.

[0068] Step S103, obtaining initial power resource data of the current power resource, inputting the initial power resource data and the improved data packet into a pre-trained reinforcement learning model, and obtaining approximate optimal results for each time period for the update demand of the current power resource.

[0069] Among them, the pre-trained reinforcement learning model has an initial policy network.

[0070] For example, when the terminal adopts a reinforcement learning model, each decision of the reinforcement learning model will calculate the update result for the current power resources based on the input data and the current state (observation value). Reinforcement learning optimizes its own policy function through continuous exploration and verification, so it requires a lot of resources for trial and error, which also makes the cost of reinforcement learning training very high. The initial policy network enables reinforcement learning to explore, verify and optimize the existing initial policy, so as to obtain the optimal sequential policy network more efficiently.

[0071] For example, the SCUC problem is a multi-period planning problem. Since there are correlations between multi-period planning, such as unit climbing constraints, unit minimum start-stop constraints, etc., it is also a sequential optimal decision-making problem. Reinforcement learning is used to convert the SCUC problem into a sequential decision-making method, and attempts are made to approximate the optimal solution to the unit start-stop integer variables or partial integer variables at each moment, thereby completing the unit start-stop planning at multiple moments, reducing the solution complexity of the original SCUC problem, and achieving an accelerated solution effect.

[0072] Step S104, inputting the approximate optimal result of each time period, the initial power resource data and the improved data packet into the pre-trained power resource update model to obtain the power resource update result for the current power resource update demand.

[0073] Exemplarily, the power resource update model can be a mixed integer linear programming model, specifically a MindOpt model (a mathematical programming solver suite). Reinforcement learning models usually obtain solutions to continuous parameters rather than integer solutions. The terminal obtains continuous solutions to certain integer type update parameters in step S103. Therefore, the terminal obtains approximate optimal results for each time period for the current power resource update demand. The power resource update model is subsequently required to perform calculations to obtain power resource update results that meet the update parameters. The power resource update model adopts a mixed integer linear programming model, which can solve update requirements that have both integer type update parameters and continuous type update parameters. In addition, when the terminal adopts the power resource update model, using the approximate optimal result as input can effectively improve the solution operation efficiency of the power resource update model.

[0074] In the above-mentioned power resource updating method, first, an initial data packet for the update demand of the current power resource is obtained, wherein the initial data packet contains the update constraint information, update target information and update parameter information of the current power resource; then, the initial data packet is input into the convex hull improvement model for data improvement to obtain an improved data packet. The convex hull improvement model can be used to improve the mathematical representation of the data in the initial data packet to more accurately reflect the update constraint information, update target information and update parameter information of the current power resource, thereby reducing the search space of the update operation and improving the model operation efficiency during the subsequent power resource update; then, the initial power resource data of the current power resource is obtained, and the initial power resource data of the power resource and the improved data packet are input into the pre-trained reinforcement learning model to obtain the various update requirements for the current power resource. The approximate optimal result of each time period is obtained, wherein the pre-trained reinforcement learning model has an initial policy network, and the initial policy network can reduce the ineffective exploration cost of the reinforcement learning model when the power resources are updated, thereby improving the efficiency of the reinforcement learning model in obtaining the approximate optimal result; finally, the approximate optimal result of each time period, the initial power resource data of the power resources and the improved data packet are input into the pre-trained power resource update model to obtain the power resource update result for the current power resource update demand, and the approximate optimal result can be used as the starting point of the power resource update model, which helps to reduce the computational burden of the model, and can significantly reduce the time and computing resources required for the model operation, and the approximate optimal result can also be used as the initial solution in the power resource update model, which significantly improves the model operation efficiency, as well as improves the accuracy and stability of the model. In the above method, before using the power resource update model, the reinforcement learning model is first used to obtain the approximate optimal result for the current power resource update demand, and the approximate optimal result is used as the initial solution of the power resource update model, thereby improving the solution operation efficiency of the power resource update model; and the reinforcement learning model has an initial strategy network, which can reduce the invalid exploration cost of the reinforcement learning model when the power resource is updated, thereby improving the efficiency of the reinforcement learning model in obtaining the approximate optimal result; in addition, the input to the reinforcement learning model and the power resource update model is an improved data packet that has been improved by the convex hull improvement model, and the improved data packet can more accurately reflect the update constraint information, update target information and update parameter information of the current power resources, reduce the search space of the update operation, and further improve the solution operation efficiency of the reinforcement learning model and the power resource update model.

[0075] In an exemplary embodiment, before the above step S103 inputs the initial power resource data and the improved data packet into the pre-trained reinforcement learning model to obtain the approximate optimal results for each time period of the current power resource update demand, it also includes: obtaining historical power resource update samples; inputting the historical power resource update samples into the imitation learning model to obtain historical update results and a historical strategy network; and training the initial reinforcement learning model according to the historical update results and the historical strategy network to obtain a pre-trained reinforcement learning model with an initial strategy network.

[0076] Exemplarily, when the terminal adopts the imitation learning model, it will first solve the historical power resource update samples through the global optimization solution algorithm to obtain the corresponding update results, and complete the construction of the training sample data. The training sample data will contain features and labels, where the feature is the observation space in reinforcement learning, which includes the power resource update results made before the current moment and the observation values ​​of the power grid at the current moment, and the label is the action space in reinforcement learning, which includes the power resource update results at the current moment. Using the above sample data, the expert strategy can be modeled through supervised learning, so as to obtain a policy network with certain expert knowledge and initialize the policy network of reinforcement learning, reducing the ineffective exploration cost. Then, in the training stage of the reinforcement learning model, the intelligent agent in the reinforcement learning will first interact with the simulation environment and accumulate data (trajectory). During the interaction, the reinforcement learning agent will determine the action at the current moment according to the current state (observation value), and perform a certain probability perturbation on this result to exploit the learned knowledge and explore the unknown possibilities. The simulation environment, i.e., the single-moment optimization solver, will solve the remaining update variables according to the given power resource update results at the current moment, including the update results of ungiven integer variables and continuous variables. The optimization solver will return information such as whether the update demand can be solved in the current situation, the time taken to solve, and the target value (reward), thereby completing an interaction process and accumulating interaction data (observation, action, reward). When the trajectory data set is continuously collected and updated, the reinforcement learning model can be trained synchronously.

[0077] In a specific example, if Figure 2The figure below is a diagram of the training process of a deep reinforcement learning algorithm model using the Actor-Critic architecture. The Actor is the agent policy network mentioned above. During the training process, it adjusts the parameters of the policy network according to the overall reward formed by different trajectories, so that the policy network can output actions that make the overall trajectory reward high. In addition, reinforcement learning generally uses Q-function or Value-function to represent the average overall reward that can be obtained after the current moment. In the Actor-Critic architecture, the Q-function or Value-function is generally approximated by optimizing the Critic network parameters, and this information is used to update the Actor's network parameters.

[0078] In this embodiment, the historical update results and the historical strategy network are obtained through the imitation learning model, and then the initial reinforcement learning model is trained according to the historical update results and the historical strategy network, so that the reinforcement learning model is trained to obtain the initial strategy network, and subsequently the approximate optimal results for the update requirements of the current power resources can be obtained more efficiently.

[0079] In an exemplary embodiment, before the above step S104 inputs the approximate optimal results of each time period, the initial power resource data and the improved data packet into the pre-trained power resource update model to obtain the power resource update results for the current power resource update requirements, it also includes: determining the feasibility of the approximate optimal results of each time period for the current power resource update requirements; using a feasibility repair model to perform feasibility repair on the approximate optimal results that are not feasible, and obtaining the repaired approximate optimal results.

[0080] Furthermore, the above-mentioned step S104 inputs the approximate optimal results of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain the power resource update results for the current power resource update needs, and also includes: inputting the repaired approximate optimal results, the feasible approximate optimal results, the initial power resource data and the improved data packet into the pre-trained power resource update model to obtain the power resource update results for the current power resource update needs.

[0081] Exemplarily, before inputting the approximate optimal result into the power resource update model, the terminal may also first determine the feasibility of each approximate optimal result for the current power resource update demand, such as whether it is an integer result. The terminal performs feasibility repair on the approximate optimal result that is not feasible to obtain the repaired approximate optimal result. Then the terminal inputs the repaired approximate optimal result, the approximate optimal result that is initially feasible, the initial power resource data, and the improved data packet into the pre-trained power resource update model to obtain the power resource update result for the current power resource update demand. In a specific example, Figure 3 Shown is a schematic diagram of a feasibility repair model.

[0082] In this embodiment, the feasibility repair model is used to repair the approximate optimal result to reduce misleading. The repaired approximate optimal result and the feasible approximate optimal result are input into the power resource update model together, providing a more comprehensive selection to help find a better power resource update result, which can reduce the complexity of searching in the power resource update model and improve the operation efficiency.

[0083] In an exemplary embodiment, the above-mentioned step S104 inputs the approximate optimal results of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain the power resource update results for the current power resource update requirements, and also includes: in the power resource update model, preprocessing the improved data packet to obtain the preprocessed improved data packet; the preprocessing at least includes redundant constraint identification processing, variable reduction processing, variable boundary tightening processing, variable parameter update processing and variable boundary detection processing; according to the preprocessed improved data packet, the approximate optimal results of each time period and the initial power resource data, the power resource update results for the current power resource update requirements are obtained.

[0084] Exemplarily, before the terminal uses the power resource update model to perform resource update operations, it is also necessary to pre-process the improved data packet to improve the model's solution efficiency and result quality. Among them, the redundant constraint identification process is to identify and delete redundant constraints in the improved data packet. Redundant constraints are unnecessary restrictions in the optimization problem. Their existence may make the problem more complicated, but will not affect the optimal result. The variable reduction process is to identify and remove unnecessary variables that may not have a significant impact on the optimal result of the problem. The tightening variable boundary process is to tighten the upper and lower bounds of the variables in the demand by analyzing the information in the data packet to more accurately reflect the actual situation. The updating variable parameter process is to update the variable parameters in the model according to the current power resource update demand and the update parameter information in the improved data packet. The variable boundary detection process is to detect possible variable boundaries to better constrain demand.

[0085] In this embodiment, the improved data packet is preprocessed to improve the solution efficiency and result quality of the model. Specifically, by removing redundant constraints, the scale of the demand problem is reduced, the complexity of the calculation is reduced, which helps to improve the solution speed and reduce the demand for computing resources. Through variable reduction processing, the scale of the demand problem is reduced, the complexity of the calculation is reduced, and the number of decision variables is reduced, which helps to find a solution faster. By tightening the variable boundary processing, the accuracy of the model is improved, making the demand problem closer to the actual situation, helping to find a more reasonable solution and reducing errors. By updating the variable parameter processing, it is ensured that the variable parameters in the model are consistent with the actual situation, which improves the feasibility and solution efficiency of the demand problem. Through variable boundary detection processing, the constraints of the demand problem are increased, making the demand problem easier to solve.

[0086] In an exemplary embodiment, the above-mentioned method of obtaining the power resource update result for the current power resource update demand based on the preprocessed improved data packet, the approximate optimal result of each time period and the initial power resource data also includes: using a linear relaxation solution model to process the preprocessed improved data packet, the approximate optimal result of each time period and the initial power resource data to obtain a relaxed lower bound of the power resource update result; using the relaxed lower bound as the initial lower bound, and using a cutting plane model to process the preprocessed improved data packet, the approximate optimal result of each time period and the initial power resource data to obtain a result range of the power resource update result; using a heuristic model to obtain suboptimal results and local optimal results from the result range; inputting the suboptimal results and the local optimal results into a parallel branch and bound model to obtain a global optimal result as the power resource update result for the current power resource update demand.

[0087] Exemplarily, the terminal adopts the power resource update model, including a linear relaxation solution model processing step, a cutting plane model processing step, a heuristic model processing step and a parallel branch and bound model processing step. Among them, the linear relaxation model is a method for solving integer programming problems. It converts the integer programming problem into a linear programming problem by relaxing the constraints of the integer variables, and then solves the linear programming problem to obtain a relaxed lower bound. The cutting plane model is a method for enhancing the solution of integer programming problems. It improves the upper bound of the problem by adding additional constraints (cutting planes) according to the current solution. The constraints are usually obtained based on the linear relaxation model. The heuristic model is an algorithm based on experience and rules, which is used to quickly find suboptimal results and local optimal results in the search space. It does not guarantee to find the global optimal result, but it can usually find a feasible solution in a shorter time. The parallel branch and bound model is an advanced search algorithm for solving integer programming problems. It divides the problem into multiple sub-problems and searches the results of these sub-problems at the same time in order to find the global optimal result more quickly. When the relaxation of a sub-problem fails to obtain a result that meets the integer requirements, the sub-problem will be divided into two or more sub-problems. The entire solution trajectory constitutes a tree structure with the original problem as the root node, and the remaining sub-problems are branch nodes. A large number of sub-problems that need to be solved will still be generated during the solution process. In order to make full use of the parallel capabilities of computing devices, the traditional serial method is improved to solving multiple sub-problems in parallel at the same time, which can greatly speed up the algorithm process.

[0088] In a specific example, the cutting plane model can be a Gomory fractional cut model, a MIR (Mixed Integer Rounding) cut model, a cover cut model, or a clique cut model. The heuristic model can be Rounding, RENS (Rounding-based, Enumeration and Normalization Scheme), and Diving, where the Diving process diagram is as follows: Figure 4 shown.

[0089] In this embodiment, the update processing efficiency of the power resource update model is effectively improved by processing through the linear relaxation solution model, the cutting plane model, the heuristic model and the parallel branch and bound model in the power resource update model.

[0090] In an exemplary embodiment, before the above step 104 inputs the approximate optimal results of each time period, the initial power resource data and the improved data packet into the pre-trained power resource update model to obtain the power resource update results for the current power resource update requirements, it also includes: obtaining historical power resource update samples; using the historical power resource update samples as training samples, and using a black box optimization model to adjust the parameters of the initial power resource update model to obtain a trained power resource update model.

[0091] Exemplarily, first, the terminal needs to obtain historical power resource update samples, which contain information such as power resource demand at different time points, actual power resource status, and corresponding update results. Then, the terminal organizes the historical power resource update samples and converts them into the form of a training data set, including input features and target labels. Obtain a black box optimization model, which is usually used to find the optimal hyperparameter configuration to optimize the performance of the power resource update model. Define the search space of hyperparameters, that is, which hyperparameters need to be optimized and their value ranges. Finally, run the black box optimization model to find the best hyperparameter configuration.

[0092] In a specific example, if Figure 5 As shown in the figure, it is a schematic diagram of training a mixed integer programming model (i.e., MILP solver). Specifically, it includes: ① Taking the historical SCUC example file as input. ② Defining the search space (possible values, parameter types) of the hyperparameters of the solver to be adjusted. ③ Setting the configuration parameters of the parameter adjustment task (parameter adjustment target, total parameter adjustment duration, etc.). ④ The parameter adjustment algorithm continuously iterates to generate candidate hyperparameter combinations and evaluates their performance. ⑤ After receiving the evaluation value, the parameter adjustment algorithm updates the proxy model of the parameter optimization objective function. ⑥ Output the optimal hyperparameters when the parameter adjustment is completed.

[0093] In this embodiment, by using a black box optimization model to adjust the parameters of the power resource update model, the optimal hyperparameter configuration can be automatically found to improve the performance and accuracy of the power resource update model.

[0094] In another exemplary embodiment, Figure 6 As shown, a method for updating power resources is provided, the method comprising the following steps:

[0095] Step S601, obtaining an initial data packet for the update requirement of the current power resources.

[0096] The initial data packet includes update constraint information, update target information and update parameter information of the current power resources.

[0097] Step S602, inputting the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet.

[0098] Step S603, obtaining historical power resource update samples, inputting the historical power resource update samples into the imitation learning model, and obtaining historical update results and a historical strategy network.

[0099] Step S604: train the initial reinforcement learning model according to the historical update results and the historical policy network to obtain a pre-trained reinforcement learning model with the initial policy network.

[0100] Step S605: Use the historical power resource update samples as training samples, and use the black box optimization model to adjust the parameters of the initial power resource update model to obtain a trained power resource update model.

[0101] Step S606, obtaining initial power resource data of the current power resource, inputting the initial power resource data and the improved data packet into a pre-trained reinforcement learning model, and obtaining approximate optimal results for each time period for the update demand of the current power resource.

[0102] Step S607, determining the feasibility of the approximate optimal result of each time period for the update demand of the current power resources, and using a feasibility repair model to repair the approximate optimal result that is not feasible, to obtain the repaired approximate optimal result.

[0103] Step S608, inputting the repaired approximate optimal result, the feasible approximate optimal result, the initial power resource data and the improved data packet into the pre-trained power resource update model.

[0104] Step S609: preprocessing the improved data packet in the power resource update model to obtain a preprocessed improved data packet.

[0105] The preprocessing includes at least redundant constraint identification processing, variable reduction processing, variable boundary tightening processing, variable parameter updating processing and variable boundary detection processing.

[0106] Step S610, using a linear relaxation solution model to process the pre-processed improved data packet, the approximate optimal results of each time period and the initial power resource data, to obtain a relaxed lower bound of the power resource update result.

[0107] Step S611, taking the relaxed lower bound as the initial lower bound, and using the cutting plane model to process the preprocessed improved data packet, the approximate optimal results of each time period, and the initial power resource data, to obtain a result range of the power resource update result.

[0108] Step S612, using a heuristic model to obtain suboptimal results and local optimal results from the result range, and inputting the suboptimal results and local optimal results into a parallel branch and bound model to obtain a global optimal result as a power resource update result for the current power resource update demand.

[0109] In this embodiment, before using the power resource update model, the reinforcement learning model is first used to obtain the approximate optimal result for the update demand of the current power resources, and the approximate optimal result is used as the initial solution of the power resource update model, thereby improving the efficiency of the solution operation of the power resource update model; and the reinforcement learning model is an initial strategy network, which can reduce the invalid exploration cost of the reinforcement learning model when the power resources are updated, thereby improving the efficiency of the reinforcement learning model in obtaining the approximate optimal result; in addition, the input of the reinforcement learning model and the power resource update model is an improved data packet that has been improved by the convex hull improvement model, and the improved data packet can more accurately reflect the update constraint information, update target information and update parameter information of the current power resources, reduce the search space of the update operation, and further improve the solution operation efficiency of the reinforcement learning model and the power resource update model. At the same time, before the power resource update model is used, a variety of preprocessing methods are used to further improve the solution operation efficiency of the power resource update model.

[0110] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0111] Based on the same inventive concept, the embodiment of the present application also provides a power resource updating device for implementing the power resource updating method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more power resource updating device embodiments provided below can refer to the limitations of the power resource updating method above, and will not be repeated here.

[0112] In an exemplary embodiment, Figure 7As shown, a power resource updating device is provided, comprising: a data acquisition module 701, a data improvement module 702, a reinforcement learning module 703 and a resource updating module 704, wherein:

[0113] The data acquisition module 701 is used to acquire an initial data packet for the update requirements of the current power resources; the initial data packet includes update constraint information, update target information and update parameter information of the current power resources;

[0114] A data improvement module 702 is used to input the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet;

[0115] The reinforcement learning module 703 is used to obtain the initial power resource data of the current power resource, input the initial power resource data and the improved data packet into the pre-trained reinforcement learning model, and obtain the approximate optimal results of each time period for the update demand of the current power resource; the pre-trained reinforcement learning model has an initial strategy network;

[0116] The resource update module 704 is used to input the approximate optimal results of each time period, the initial power resource data and the improved data packet into the pre-trained power resource update model to obtain the power resource update results for the current power resource update requirements.

[0117] In one embodiment, the above-mentioned power resource updating device also includes a model training module, which is used to obtain historical power resource update samples; input the historical power resource update samples into the imitation learning model to obtain historical update results and a historical strategy network; according to the historical update results and the historical strategy network, the initial reinforcement learning model is trained to obtain a pre-trained reinforcement learning model with an initial strategy network.

[0118] In one embodiment, the above-mentioned model training module is also used to obtain historical power resource update samples; the historical power resource update samples are used as training samples, and the black box optimization model is used to adjust the parameters of the initial power resource update model to obtain a trained power resource update model.

[0119] In one embodiment, the above-mentioned power resource updating device also includes a result repair module, which is used to determine the feasibility of the approximate optimal result of each time period for the current power resource update demand; using a feasibility repair model, the approximate optimal result that is not feasible is repaired for feasibility to obtain the repaired approximate optimal result.

[0120] In one embodiment, the resource update module 704 is also used to input the repaired approximate optimal result, the feasible approximate optimal result, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain the power resource update result for the current power resource update requirements.

[0121] In one embodiment, the resource update module 704 is also used to preprocess the improved data packet in the power resource update model to obtain a preprocessed improved data packet; the preprocessing includes at least redundant constraint identification processing, variable reduction processing, variable boundary tightening processing, variable parameter update processing and variable boundary detection processing; based on the preprocessed improved data packet, the approximate optimal results of each time period and the initial power resource data, the power resource update result for the current power resource update demand is obtained.

[0122] In one embodiment, the resource update module 704 is also used to use a linear relaxation solution model to process the preprocessed improved data packet, the approximate optimal results of each time period and the initial power resource data to obtain a relaxed lower bound of the power resource update result; using the relaxed lower bound as the initial lower bound, using a cutting plane model to process the preprocessed improved data packet, the approximate optimal results of each time period and the initial power resource data to obtain a result range of the power resource update result; using a heuristic model to obtain a suboptimal result and a local optimal result from the result range; inputting the suboptimal result and the local optimal result into a parallel branch and bound model to obtain a global optimal result as the power resource update result for the current power resource update demand.

[0123] Each module in the above-mentioned power resource updating device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0124] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store pre-trained model data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for updating electric power resources is implemented.

[0125] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0126] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0127] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0128] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0129] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0130] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0131] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0132] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for updating electric power resources, characterized in that: The method comprises: Acquire an initial data packet for the update requirements of the current power resources; the initial data packet includes update constraint information, update target information and update parameter information of the current power resources; Inputting the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet; Acquire initial power resource data of the current power resource, input the initial power resource data and the improved data packet into a pre-trained reinforcement learning model, and obtain approximate optimal results for each time period of the update demand of the current power resource; the pre-trained reinforcement learning model has an initial strategy network; Inputting the approximate optimal results of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource; The step of inputting the approximate optimal result of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource includes: In the power resource update model, the improved data packet is preprocessed to obtain a preprocessed improved data packet; the preprocessing at least includes redundant constraint identification processing, variable reduction processing, variable boundary tightening processing, variable parameter updating processing and variable boundary detection processing; Using a linear relaxation solution model to process the preprocessed improved data packet, the approximate optimal results of each time period and the initial power resource data, to obtain a relaxed lower bound of the power resource update result; The relaxed lower bound is used as an initial lower bound, and a cutting plane model is used to process the preprocessed improved data packet, the approximate optimal result of each time period, and the initial power resource data to obtain a result range of the power resource update result; Using heuristic models to obtain suboptimal and local optimal results from the range of results; The suboptimal result and the local optimal result are input into a parallel branch and bound model to obtain a global optimal result as a power resource update result for the update demand of the current power resource.

2. The method according to claim 1, characterized in that Before inputting the initial power resource data and the improved data packet into a pre-trained reinforcement learning model to obtain approximate optimal results for each time period of the update demand for the current power resource, the method further includes: Get updated samples of historical power resources; Inputting the historical power resource update samples into the imitation learning model to obtain historical update results and a historical strategy network; The initial reinforcement learning model is trained according to the historical update results and the historical strategy network to obtain the pre-trained reinforcement learning model with the initial strategy network.

3. The method according to claim 1, characterized in that Before inputting the approximate optimal results of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource, the method further includes: Determining the feasibility of the approximate optimal result of each of the time periods with respect to the update demand of the current power resource; Adopting the feasibility repair model, the infeasible approximate optimal result is repaired to obtain the repaired approximate optimal result; The step of inputting the approximate optimal result of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource includes: The repaired approximate optimal result, the feasible approximate optimal result, the initial power resource data and the improved data packet are input into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource.

4. The method according to claim 1, characterized in that: Before inputting the approximate optimal results of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource, the method further includes: Get updated samples of historical power resources; The historical power resource update samples are used as training samples, and a black box optimization model is used to adjust the parameters of the initial power resource update model to obtain a trained power resource update model.

5. A power resource updating device, characterized in that: The device comprises: A data acquisition module, used to acquire an initial data packet for the update requirements of the current power resources; the initial data packet includes update constraint information, update target information and update parameter information of the current power resources; A data improvement module, used for inputting the initial data packet into the convex hull improvement model to perform data improvement to obtain an improved data packet; A reinforcement learning module, used for obtaining initial power resource data of the current power resource, inputting the initial power resource data and the improved data packet into a pre-trained reinforcement learning model, and obtaining approximate optimal results for each time period of the update demand of the current power resource; the pre-trained reinforcement learning model has an initial strategy network; A resource update module, used for inputting the approximate optimal results of each time period, the initial power resource data and the improved data packet into a pre-trained power resource update model to obtain a power resource update result for the update demand of the current power resource; The resource update module is further used to preprocess the improved data packet in the power resource update model to obtain a preprocessed improved data packet; the preprocessing at least includes redundant constraint identification processing, variable reduction processing, variable boundary tightening processing, variable parameter update processing and variable boundary detection processing; Using a linear relaxation solution model to process the preprocessed improved data packet, the approximate optimal results of each time period and the initial power resource data, to obtain a relaxed lower bound of the power resource update result; The relaxed lower bound is used as an initial lower bound, and a cutting plane model is used to process the preprocessed improved data packet, the approximate optimal result of each time period, and the initial power resource data to obtain a result range of the power resource update result; Using heuristic models to obtain suboptimal and local optimal results from the range of results; The suboptimal result and the local optimal result are input into a parallel branch and bound model to obtain a global optimal result as a power resource update result for the update demand of the current power resource.

6. The device according to claim 5, characterized in that The device also includes: The model training module is used to obtain historical power resource update samples; input the historical power resource update samples into the imitation learning model to obtain historical update results and a historical strategy network; and train the initial reinforcement learning model according to the historical update results and the historical strategy network to obtain the pre-trained reinforcement learning model with the initial strategy network.

7. The device according to claim 5, characterized in that The device also includes: The model training module is used to obtain historical power resource update samples; the historical power resource update samples are used as training samples, and the black box optimization model is used to adjust the parameters of the initial power resource update model to obtain a trained power resource update model.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Electromagnetic frequency spectrum visualization analysis method based on spatio-temporal integrated digital earth

    CN114003981A

  • Distributed power grid transient and steady state operation method and device based on digital twinborn map

    CN116404760A